Skip to content

Open nowPosted 91 days ago

ML Model Evaluation Engineer

Triomics16 open roles

Where
India Office
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowML Model Evaluation EngineerTriomics · India Office
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Triomics's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market. Triomics postings stay open a median of 13 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.4%3 days
  3. 7.8%7 days
  4. 14.3%14 days
  5. 33.7%30 days
This job: posted 91 days ago

Triomics median: 13 days open

The posting

GROWTH PATH

This is an individual contributor role with strong ownership expectations. High performers may be considered for workstream lead or functional lead responsibilities after approximately 12 months, based on demonstrated ownership, delivery, technical judgment, mentoring, cross-functional influence, and ability to reduce dependency on the Director of ML.

ABOUT THE ROLE

We are looking for an ML Evaluation Engineer to own model quality, regression testing, release validation, and production impact analysis for clinical AI systems. This role sits between applied ML, clinical data, MLOps, and production operations.

Your job is to ensure that every model or workflow release is measurable, stable, and not degrading important clinical behavior. You will maintain evaluation datasets, create hidden test sets, run regression checks, analyze production issues, and produce release-readiness reports.

WHAT YOU WILL DO

- Build and maintain evaluation frameworks for clinical NLP, LLM, RAG, information extraction, and structured abstraction systems.

- Create and manage hidden test datasets that are not directly visible to model developers, reducing overfitting risk.

- Define release metrics, regression thresholds, slice-based evaluation, failure-mode tracking, and release/blocker criteria.

- Compare model versions and identify performance degradation across clinical segments, document types, clients, data sources, labels, and edge cases.

- Work with Clinical AI Data Specialists to design gold sets, hidden test sets, adjudication workflows, and label quality checks.

- Work with Research Engineers to understand model changes, expected behavior, and evaluation risks without compromising test-set independence.

- Work with MLOps/Data Engineering to monitor production behavior, triage bugs, analyze incident impact, and prioritize fixes.

- Create release-readiness reports before production deployment.

- Build dashboards, scripts, and automated checks for evaluation, monitoring, regression testing, and model comparison.

- Prioritize model bugs based on clinical severity, user impact, frequency, regression risk, and operational urgency.

WHAT WE EXPECT

- 3–6+ years of experience in ML engineering, data science, model evaluation, ML QA, applied NLP evaluation, or data-heavy quality engineering.

- Strong Python and data analysis skills.

- Strong understanding of precision/recall/F1, calibration, confidence thresholds, dataset splits, leakage, overfitting, statistical testing, and error analysis.

- Experience building evaluation pipelines, benchmark suites, test harnesses, dashboards, or regression frameworks.

- Ability to work with imperfect labels, annotation disagreement, clinical ambiguity, and hidden evaluation sets.

- Strong independence and judgment; ability to challenge releases when evidence is weak.

- Clear written communication for release reports, incident analysis, and quality decisions.

NICE TO HAVE

- Experience with LLM evaluation, RAG evaluation, extraction evaluation, clinical NLP, or healthcare ML.

- Experience with model monitoring, production incident analysis, data drift, or observability.

- Experience with MLflow, Weights & Biases, Evidently, Great Expectations, DeepEval, Ragas, pytest, Airflow, Prefect, or similar tools.

- Clinical or biomedical NLP exposure.

SUCCESS IN 6 MONTHS

- Establishes a repeatable release validation process.

- Maintains hidden evaluation datasets and prevents overfitting to test data.

- Produces release reports that leadership, ML, and engineering can trust.

- Catches meaningful regressions before release.

- Provides reliable impact analysis for production issues and helps prioritize fixes.

About Triomics

Triomics is building the agentic AI layer for oncology EHRs. Cancer hospitals spend billions on highly trained staff manually reading unstructured patient records - pathology reports, clinical notes, genomic panels - to power workflows like trial matching, registry curation, visit prep, and quality reporting. We replace that manual work with task-driven AI agents that sit inside the EMR and process records at scale, in real time.

Our platform is trusted by leading cancer centers including Memorial Sloan Kettering, Mount Sinai, and Yale Cancer Center. We have grown 10x in the last year and process millions of oncology medical documents monthly.

Our investors include Battery Ventures, Lightspeed, General Catalyst, Nexus Venture Partners, and Y Combinator.

Why Join Triomics

- Impact at scale. The systems your teams build directly power AI workflows that accelerate cancer research and improve patient outcomes.

- Cutting-edge problems. Hard, data-intensive systems at the intersection of AI, healthcare, and scale - in a highly regulated industry where reliability is non-negotiable.

- World-class team. Work alongside top talent across AI, engineering, and product, with best-in-industry compensation.

- Culture that ships. Fast-paced, ownership-driven, with company-sponsored workations.

Perks & Benefits

- Lunch provided at the office - one less daily decision.

- Flexible working hours - we care about output, not clock-ins.

- Comprehensive health insurance for you and your family.

- Zomato meal benefits for early starts and late nights.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Triomics's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Triomics's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Triomics's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.