Skip to content

Open nowPosted 66 days ago

Member of Technical Staff — Agent Post-Training

moonlake7 open roles

Where
San Francisco, CA
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowMember of Technical Staff — Agent Post-Trainingmoonlake · San Francisco, CA
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on moonlake's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.6%1 day
  2. 3.4%3 days
  3. 7.8%7 days
  4. 14.3%14 days
  5. 33.7%30 days
This job: posted 66 days ago

The posting

Introducing Moonlake, AI for creating world simulations.

ABOUT MOONLAKE

Moonlake is building the frontier of interactive world models: systems that generate, simulate, and reason over 3D environments for robotics, embodied AI, and interactive applications.

We develop the infrastructure that enables intelligent systems to learn, evaluate, and interact within realistic virtual environments before operating in the physical world.

Our work sits at the intersection of:

- Robotics

- Embodied AI

- Interactive 3D Worlds

- World Models

- Simulation Infrastructure

- Physical AI

Moonlake is building the next generation of AI infrastructure for interactive digital worlds. Our mission is to enable anyone to create, simulate, and interact with rich environments using natural language and multimodal inputs, turning simple ideas into worlds with structure, physics, and intelligent behavior.

Our team has raised $50M in seed funding from NVIDIA Ventures, Threshold Ventures, AIX Ventures, and notable angels including Naval Ravikant and Jeff Dean to build the foundational layer for the future of AI—powering everything from robotics training and simulation to digital twins and interactive environments.

We are looking for exceptional engineers to help build the simulation systems that will power the next generation of robotics and embodied intelligence.

THE ROLE

We are hiring a Member of Technical Staff to lead reinforcement learning infrastructure and model post-training.

You will work closely with Qi and the research team to improve large vision-language and code-generating agents through fine-tuning, reinforcement learning, trajectory data, and scalable evaluation.

Moonlake already has deep expertise in 3D and world-building. This role adds the model-training experience needed to systematically improve agent performance and prepare the company for larger-scale RL across both digital and physical environments.

We are looking for a full-stack researcher and engineer who understands the complete training system and can make strong judgments about when training is necessary, which methods are likely to work, and what not to pursue.

WHAT YOU’LL DO

- Build RL and post-training pipelines for multimodal, vision-language, and code-generating agents

- Develop infrastructure for supervised fine-tuning, preference optimization, reward modeling, and reinforcement learning

- Create systems for collecting, filtering, replaying, and learning from agent trajectories

- Design rewards, verifiers, and evaluations for long-horizon agent tasks

- Improve agents’ ability to plan, write and execute code, use tools, recover from errors, and complete complex workflows

- Scale distributed training and high-throughput rollout generation across multi-GPU environments

- Improve training reliability, reproducibility, observability, and cost efficiency

- Help define Moonlake’s long-term strategy for agent, robotics, and embodied-model training

WHAT WE’RE LOOKING FOR

- Real-world experience training large language, vision-language, multimodal, or code models

- Strong experience in reinforcement learning, post-training, or large-scale fine-tuning

- Experience building distributed training or high-throughput inference systems

- Familiarity with supervised fine-tuning, preference optimization, reward modeling, and agentic RL

- Experience with code-generation agents, long-horizon evaluation, or tool-using systems

- Strong Python skills and experience with PyTorch, JAX, or similar frameworks

- Ability to work across data, models, environments, rewards, evaluation, and infrastructure

- Strong research judgment and a bias toward building reliable systems

PREFERRED EXPERIENCE

- Experience at a frontier AI lab or organization operating large-scale training systems

- Experience with code-model post-training or autonomous coding agents

- Experience with multimodal models, robotics, simulation, or embodied AI

- Experience designing verifiable rewards or outcome-based training systems

- Experience scaling RL workloads across large GPU clusters

WHAT SUCCESS LOOKS LIKE

Within your first year, you will have:

- Built Moonlake’s core post-training and RL infrastructure

- Created scalable systems for learning from agent trajectories

- Delivered measurable improvements in agent quality and task completion

- Helped the team determine which problems require training and which do not

- Established a reliable foundation for larger-scale agent and embodied-model training

WHY THIS ROLE MATTERS

Moonlake’s agents must do more than generate content. They must understand complex requests, reason across vision and language, write and execute code, operate tools, build interactive worlds, and recover from mistakes.

This role will build the training systems that allow those agents to continuously improve.

We are committed to being an on-site, in-person team currently based in San Francisco.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against moonlake's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on moonlake's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    moonlake's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.