Skip to content

Open nowPosted 17 hours ago

Senior QA Engineer - AI

Workable (global search)107,808 open roles

Where
India
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior QA Engineer - AIWorkable (global search) · India
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Workable (global search)'s own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.0% of postings close within 7 days. Measured by our own scanner across the market. Workable (global search) postings stay open a median of 2 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 8.0%7 days
  4. 15.0%14 days
  5. 34.2%30 days
This job: posted 17 hours ago

Workable (global search) median: 2 days open

The posting

About the Company

At Delta, we are reimagining and rebuilding the financial system. Join our team to make a positive impact on the future of finance.

🎯 Mission Driven: Re-imagine and rebuild the future of finance.

💡 Most innovative cryptocurrency derivatives exchange. With a daily traded volume of ~$ 10 billion, and increasing. Delta is bigger than all the Indian crypto exchanges combined.

📈 Offer the widest range of derivative products and have been serving traders all over the globe since 2018 and growing fast.

💪🏻 The founding team is comprised of IIT and ISB graduates. Business co-founders have previously worked with Citibank, UBS and GIC; and our tech co-founder is a serial entrepreneur who previously co-founded TinyOwl and Housing.com.

💰 Funded by top crypto funds (Sino Global Capital, CoinFund, Gumi Cryptos) and crypto projects (Aave and Kyber Network).

Senior QA Engineer - AI

🎯 Role Overview

Delta operates three AI products used by live traders: a customer support chatbot, an API Copilot that generates and executes trading scripts, and an MCP server. This role owns the quality of those products and is responsible for measuring it objectively.

Testing probabilistic systems differs fundamentally from testing deterministic ones. The same input produces different output on every run, an incorrect answer can be entirely fluent, and a prompt change in one flow can degrade another without any visible signal. The core of this role is establishing what correct means for these systems, building the datasets and scoring that measure it, and operating the release gate that prevents a regression from reaching customers.

This is an emerging discipline with no established playbook. We are defining the methodology as we build it, and the role carries a high degree of autonomy and ownership.

🔬 What Sets This Role Apart

Most QA roles that mention AI mean using AI to do testing faster: generating test cases from a PRD, healing flaky selectors, exploring an app with an agent. Those are useful and we do them.

This role is the other thing. The system under test is itself an AI, and the hard problem is deciding whether its output is correct when the same question produces a different answer every time and a wrong answer reads as convincingly as a right one. That means building evaluation datasets, defining what correct looks like, scoring against it, and defending a number that decides whether a release ships.

If you have spent time on the second problem, this role is built for you.

🛠 Key Responsibilities

  • Build and maintain golden datasets for our AI products, trace mined from production or generated, and verified against a documented source of truth.
  • Design layered scoring for AI outputs: deterministic rules where behaviour can be asserted, and model-graded evaluation where it cannot.
  • Manage evaluation runs end to end: schedule and execute them against release candidates, compare results against prior baselines, and triage failures into product defects, incorrect expectations, and platform issues.
  • Own and operate the release gate that determines whether a prompt or model change is approved for production.
  • Identify failure modes ahead of customers, including hallucination, incorrect tool selection, loss of context across conversation turns, and PII exposure.
  • Investigate production traces to distinguish retrieval failures from generation failures from tool failures.
  • Convert findings into actionable engineering evidence and into permanent regression coverage.
  • Define and report quality metrics for AI surfaces, and drive improvement against them.

✅ Requirements

Non-negotiable

  • At least 1 year owning quality for AI products in a lead or primary-owner capacity. This includes chatbots and conversational assistants, code or content generation products, and agentic systems. You should have worked directly with prompts, tool and function schemas, and model behaviour, rather than testing around them.
  • Demonstrated hands-on experience building or operating an evaluation harness for an AI system, whether in-house or using a framework such as DeepEval, RAGAS or Promptfoo. You should be able to describe what the harness measured, the results it produced, and the decisions those results informed.

Also required

  • 4-6 years of QA / SDET experience.
  • Working proficiency in Python. Our evaluation platform, trace analysis and internal tooling are Python-based.
  • Ability to analyse traces and tool calls, using observability tooling such as Opik, Langfuse or LangSmith or an equivalent, to determine root cause rather than reporting the symptom.
  • Strong API testing experience. The majority of the surface under test is API-level.
  • Sufficient engineering ability to build your own tooling and automation.
  • Clear written communication. Findings must be documented in a form engineering can act on directly.

Bonus

  • Experience with trading or exchange platforms, including familiarity with futures and options, crypto derivatives, margin, or settlement flows. Domain knowledge can be picked up on the job, but arriving with it shortens the ramp considerably.

🚀 Why Join Us?

  • Play a pivotal role in shaping the regulatory landscape for digital assets and Web3 in India.
  • Work directly with founders and senior leadership on high-impact strategic initiatives.
  • Be part of a mission-driven, fast-growing organisation at the forefront of financial innovation.
  • Competitive compensation, leadership exposure, and significant growth opportunities.
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Workable (global search)'s own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Workable (global search)'s form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Workable (global search)'s answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.