Skip to content

Open nowPosted 2 hours agoWe saw it 74 min after it went up

Software Engineer L5/L6 — Model Evaluations & Data Curation (MEDC)

Jobgether4,076 open roles

Where
US
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSoftware Engineer L5/L6 — Model Evaluations & Data Curation (MEDC)Jobgether · US
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Jobgether's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market. Jobgether postings stay open a median of 6 days.

Share of postings closed within
  1. 1.7%1 day
  2. 3.6%3 days
  3. 8.1%7 days
  4. 15.1%14 days
  5. 34.0%30 days
This job: posted 2 hours ago

Jobgether median: 6 days open

The posting

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Software Engineer L5/L6 — Model Evaluations & Data Curation (MEDC) based in the United States.

The Software Engineer will build foundational infrastructure that accelerates how AI teams create, evaluate, and improve high-quality datasets. You will transform ad-hoc data curation workflows into reusable, scalable platforms that researchers and engineers can rely on. The role combines strong software engineering with hands-on work in LLM-driven data generation, evaluation, sampling, and quality control. You will work closely with researchers and data scientists to understand how curation decisions influence model behavior and performance. Your work will help make datasets more discoverable, reproducible, versioned, and production-ready across multiple AI initiatives. The environment is highly collaborative and technically ambitious, with significant autonomy and opportunities to influence engineering practices. At the L6 level, the role additionally calls for technical leadership and the ability to establish direction across data and evaluation infrastructure.

Accountabilities

  • Design and build reusable data curation infrastructure, including shared libraries, components, and workflows that replace fragmented notebook-based processes.
  • Develop scalable LLM-powered pipelines that transform raw catalog, metadata, and other data sources into training and evaluation datasets such as question-answer pairs and synthetic scenarios.
  • Implement large-scale batch inference workflows while balancing data quality, computational efficiency, token usage, and cost.
  • Develop sampling strategies that optimize coverage, diversity, difficulty, and representation across relevant content and member segments.
  • Create data-quality and filtering systems using techniques such as LLM-as-judge scoring, evaluation-model-based ranking, deduplication, validation, and other quality controls.
  • Partner closely with researchers to design experiments that measure how data curation choices affect downstream model behavior and performance.
  • Establish curated datasets as discoverable, reusable artifacts with clear versioning, lineage, documentation, and reproducibility.
  • Drive adoption of standardized data curation practices across engineering, research, and modeling teams.
  • At the L6 level, provide technical leadership across data and evaluation infrastructure and help define technical direction for multi-engineer initiatives.
  • Strong software engineering expertise in Python, including experience developing reusable infrastructure, libraries, frameworks, or platforms used by other engineers and researchers.
  • Hands-on experience building LLM-driven data generation or transformation pipelines, including synthetic data generation, structured outputs, or large-scale batch inference.
  • Practical experience with data quality techniques such as sampling, filtering, deduplication, validation, and model-based quality scoring, including LLM-as-judge approaches.
  • Strong modeling intuition and an understanding of how dataset composition and curation decisions can influence model behavior and performance.
  • Experience designing experiments or evaluation approaches to measure the impact of data and modeling decisions.
  • Experience with distributed data processing technologies such as Spark, Ray, or comparable frameworks.
  • Excellent collaboration and communication skills, particularly when partnering with researchers, data scientists, and platform engineering teams.
  • For L6 roles, demonstrated experience with LLM evaluation systems is required.
  • For L6 roles, demonstrated technical leadership across data or evaluation infrastructure, including setting technical direction for multi-engineer initiatives, is required.
  • Experience with dataset versioning, lineage, artifact management, experiment tracking, or model registries is highly valued.
  • Experience optimizing large-scale LLM inference for cost, throughput, or operational efficiency is a plus.
  • Familiarity with human annotation workflows and methods for calibrating LLM judges against human ratings is beneficial.
  • Experience with pipeline orchestration frameworks such as Metaflow, Airflow, or similar tools is advantageous.
  • Background in recommendation systems, personalization, search, content catalogs, or metadata-driven applications is a plus.
  • Annual compensation range of $600,000–$1,066,000, with the range varying based on location and individual market factors.
  • Compensation is structured primarily around annual salary, with the flexibility to determine the desired balance between salary and stock options each year.
  • Comprehensive health insurance plans and mental health support.
  • 401(k) retirement plan with employer matching.
  • Stock option program.
  • Health Savings Accounts and Flexible Spending Accounts.
  • Family-forming benefits.
  • Life and serious injury benefits.
  • Disability programs.
  • Paid leave of absence programs.
  • Flexible paid time off for full-time salaried employees.
  • Remote work opportunity within the United States.
  • Opportunity to work on high-impact AI infrastructure spanning foundation models, evaluation, and data curation.
  • Collaborative environment with substantial technical autonomy and opportunities for senior-level technical leadership.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Jobgether's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Jobgether's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Jobgether's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.