Founding Research Engineer
Applying for this one?
We write the CV against this exact posting — its wording, its requirements — not a template with your name in it.
$25, one-time. No subscription.
THE MISSION
Nolla is building AI doctors so that top-quality healthcare is accessible to everyone. Nolla Derm is the #1 medical skincare treatment app in the U.S. App Store. We've treated thousands of acne patients in the U.S., scanned 1% of Norway's population for skin cancer, and our in-house clinical models are state of the art on clinical benchmarks. We just launched NollaMD, our urgent care app, and we're building dedicated specialty apps for conditions like women’s health. We’ve raised $6.5M from General Catalyst and other strategic investors.
THE ROLE
We have built our own clinical benchmarks, post-trained on them, and produced models that beat frontier performance on clinical tasks. We are in a unique position of both delivering care to real patients and training the models that deliver it. You’ll work on our evals and harnesses, build the data loop with clinicians that feeds them, and lead where they go next.
You'll work directly with the founding team and with our post-training partners, and your work reaches patients immediately. This is a hands-on role: you will write code most days, design studies, and be the author of record on what we publish.
WHAT YOU'LL BUILD
EVALS AND BENCHMARKS
- Design and maintain the benchmarks we use to judge clinical accuracy, safety, and documentation quality
- Extend them to our agentic system: tool use, long-horizon tasks that span many visits, memory of a patient's history, uploaded records, labs, and escalation behavior
GRADERS AND HARNESSES
- Build the grading stack: deterministic checks, rubric graders, and LLM judges calibrated against clinician ratings
- Improve the production harness the product runs on
DATA LOOP WITH CLINICIANS
- Turn real, consented cases into eval cases and training examples, with clinician review, de-identification, and provenance built in
- Run the clinician review workflows that produce rubrics, labels, and feedback at scale
RESEARCH
- Lead our clinical evaluations from protocol through publication, and publish our benchmarks for the field
- Collaborate with post-training on reward design and data curation, and run experiments where it helps
WHAT YOU BRING
- You have built eval infrastructure for LLM systems that other people depended on
- You have opinions about contamination, difficulty calibration, judge bias, and reward hacking
- You can write a paper. Authorship on empirical ML work, ideally with a human comparison or a benchmark release
- You are excited to work with some of the best physicians in the country and turn their judgment into model performance
- You've worked on small teams and built things from zero
BONUS POINTS
- Clinical AI evaluation experience: rubric-based health benchmarks, simulated-patient studies, agentic clinical benchmarks
- Post-training experience (RLHF, DPO, GRPO or similar)
- Experience with multimodal models
- Familiarity with eval and environment frameworks such as Inspect, Verifiers, or Harbor
ROLE LOGISTICS, COMPENSATION & BENEFITS
- Role Type: Engineering
- Salary: $180,000–$240,000, based on experience
- Job Type: Full-time
- Work Setup: In-person, New York City
- Equity: Meaningful equity, commensurate with experience
- Health Insurance: Medical, dental, and vision
- HSA/FSA: Eligible
- Time Off: Flexible, unlimited vacation
- Additional Perks: Meal stipends, team retreats
- Work Hours: Flexible but demanding. We're building something that matters
- Growth: Founding team members step into expanded roles as we scale
Seen 7 hours ago · within 12 minutes of the employer posting it.
Original posting on Nolla Health's site ↗
Posting text belongs to the employer. Removal requests: contact us.
Nearby
Live postings like this one
Same employer first, then the same role elsewhere.
- 1h ago
- 8d ago
- 8d ago
One job at a time
One posting. One CV. $25.
Pick the job you actually want and we write for it.