Skip to content

Open nowPosted 9 hours ago

Post-Training Applied Researcher

Baseten110 open roles

Pay
$200,000 – $275,000 a year
Where
San Francisco
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowPost-Training Applied ResearcherBaseten · San Francisco
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Baseten's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market. Baseten postings stay open a median of 39 days.

Share of postings closed within
  1. 1.8%1 day
  2. 3.6%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.0%30 days
This job: posted 9 hours ago

Baseten median: 39 days open

The posting

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F https://www.baseten.co/blog/announcing-our-series-f/, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.

THE ROLE

This role sits at the applied end of our post-training research efforts. You will work directly with stakeholders from the world’s fastest-growing AI companies to post-train open-source models that outperform frontier closed models on their specialised tasks. Your day-to-day is finding creative ways to extract signal from complex, domain-specific datasets and building the reward functions, environments, eval harnesses, and training pipelines that turn that signal into better models. The models you train ship to production and reach millions of users.

We are looking for people with hands-on LLM fine-tuning and RL experience. Researchers who are excited by the prospect of shipping models into production, who can translate a customer's domain-specific requirements into an effective training curriculum, and who know when to be rigorous and when to iterate fast.

RECENT RESEARCH

- Dense, on-policy or both? https://www.baseten.co/research/dense-on-policy-or-both/

- Repeated kv cache for long-running agents https://www.baseten.co/research/repeated-kv-cache-for-long-running-agents/

- Distillation without the dark – replicating black-box on-policy distillation on Baseten https://www.baseten.co/research/distillation-without-the-dark/

RESPONSIBILITIES

- Design and run post-training pipelines: SFT, GRPO, DPO, RLVR, reward function engineering, and synthetic data generation.

- Build task-specific training environments and evals tailored to customer domains like healthcare, code generation, and legal, spanning multi-turn tool use, sandboxed execution, and agentic workflows.

- Work directly with customers to translate production data into training signal, designing reward loops from real usage patterns and handling distribution shift.

- Run and analyze training experiments end-to-end: diagnose reward hacking, importance sampling drift, and advantage estimation instabilities.

- Publish findings at top venues and contribute to Baseten's open-source training libraries.

QUALIFICATIONS

- Hands-on experience training LLMs with reinforcement learning — demonstrated understanding of GRPO or PPO beyond recipe-level reproduction, including group advantage computation, clipped objectives, and KL penalty design

- Strong intuition for reward engineering: the ability to distinguish between a reward that trains effectively and one that will exploit at scale

- Experience building multi-turn agent environments with tool use, not limited to single-turn question-answering setups

- Comfort working across the full pipeline from dataset construction through training, evaluation, and deployment

- Experience with production ML systems. Preference for candidates who have closed a training–inference loop where production data feeds back into model improvement

PREFERRED QUALIFICATIONS

- Experience with RL training frameworks

- Publications at NeurIPS, ICML, ICLR, focused on RL for LLMs, reward modeling, or alignment

BENEFITS

- Competitive compensation, including meaningful equity

- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents

- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)

- Paid parental leave

- Fertility and family-building stipend through Carrot

- (U.S. only) Company-facilitated 401(k)

- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

Apply now to embark on a rewarding journey in shaping the future of AI! If you are a motivated individual with a passion for machine learning and a desire to be part of a collaborative and forward-thinking team, we would love to hear from you.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Baseten's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Baseten's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Baseten's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.