Skip to content

Open nowPosted 55 days ago

Founding AI Researcher, RL

goaly7 open roles

Where
Palo Alto, CA, USA
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowFounding AI Researcher, RLgoaly · Palo Alto, CA, USA
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on goaly's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.7%1 day
  2. 3.5%3 days
  3. 7.8%7 days
  4. 14.6%14 days
  5. 34.1%30 days
This job: posted 55 days ago

The posting

ABOUT US

We’re building toward a world where every company can become its own AI lab. Goaly is a stealth AI startup founded by ex-Meta MSL engineers and researchers. Our mission is to dramatically lower the cost, time, and talent barriers to building proprietary AI — and make each generation of models faster and cheaper to build than the last.

Backed by leading AI investors and endorsed by frontier AI researchers and builders, we’re looking for exceptional new grads who want to work on hard, foundational AI systems problems with outsized ownership from day one.

ABOUT THE ROLE

You will own the experimental loop that turns a capable base model into a useful agent. You will design tasks and environments, prepare training and evaluation data, run reinforcement-learning and related post-training experiments, diagnose model behavior, and convert results into better recipes and production models.

This is a research-engineering role. The best candidates are equally comfortable forming hypotheses, writing high-quality code, operating training pipelines, and investigating why a model or metric moved. You will work closely with RL systems, training, inference, product, and domain experts; when infrastructure slows the science, you will help improve the infrastructure rather than treating it as someone else's problem.

WHAT YOU'LL DO

- Design and run post-training experiments for agentic capabilities, including tool use, coding, reasoning, planning, long-horizon task completion, and recovery from failure.

- Prepare high-quality training and evaluation data: define task distributions, curate and filter examples, control contamination, balance difficulty, and build reproducible data-generation pipelines.

- Build realistic RL environments and task harnesses with clear interfaces, reliable resets, isolated execution, useful telemetry, and reward signals that are hard to game.

- Develop evaluations that measure both capability and reliability. Create regression suites, behavioral slices, error taxonomies, and dashboards that connect aggregate metrics to concrete model failures.

- Iterate on training recipes, including supervised warm starts, sampling strategies, reward design, verifiers, curricula, optimization choices, and reinforcement fine-tuning methods.

- Analyze trajectories and model behavior to find reward hacking, shortcut learning, mode collapse, distribution gaps, and other failure modes; turn those findings into targeted experiments.

- Improve the research workflow through better experiment configuration, rollout inspection, reproducibility, checkpoint evaluation, and automated comparison of runs.

- Partner with systems engineers to debug cross-layer problems in rollout inference, environment execution, distributed training, and data movement.

- Translate successful research ideas into stable, repeatable pipelines and help set the team's longer-term post-training roadmap.

YOU MAY BE A GOOD FIT IF YOU HAVE

- Strong Python and software-engineering skills, including the ability to turn ambiguous research ideas into reliable experimental systems.

- Hands-on experience training, fine-tuning, or evaluating modern language models, or an exceptional record in a closely related ML research area.

- Solid understanding of deep learning and optimization, plus enough reinforcement-learning intuition to reason about policies, rewards, sampling, credit assignment, and evaluation bias.

- Excellent experimental judgment: you define controls, inspect data, validate metrics, keep results reproducible, and distinguish a real improvement from noise or leakage.

- Ability to debug across model behavior, data, code, and distributed infrastructure without losing sight of the user-facing capability being improved.

- Clear written and verbal communication and a track record of productive collaboration across research and engineering.

STRONG PLUSES

- Experience with RLHF, reinforcement fine-tuning, preference optimization, reward or verifier modeling, or large-scale online sampling.

- Experience building agent environments, secure sandboxes, coding benchmarks, tool-use tasks, or long-horizon evaluations.

- Familiarity with PyTorch or JAX and distributed ML systems; experience with frameworks such as FSDP, Megatron, DeepSpeed, Ray, VeRL, or related stacks.

- A record of influential research, open-source contributions, technically ambitious independent projects, or production model launches.

HOW WE WORK

- Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.

- High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.

- Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.

- Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.

- Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.

- Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.

Location, visa sponsorship & benefits

- Hybrid in Palo Alto: 4+ days/week in office.

- Visa sponsorship: H-1B and OPT/CPT support available, with immigration counsel.

- Meals & perks: Complimentary lunch, dinner, snacks, and drinks.

A note on qualifications. We value exceptional ability over perfect keyword matches. If the work excites you and you can show strong technical ability, learning speed, or ownership, we encourage you to apply.

EQUAL OPPORTUNITY

We are an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. We provide reasonable accommodations for candidates who need them during the hiring process.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against goaly's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on goaly's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    goaly's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.