The posting
ABOUT THE ROLE
Join an AI/ML team focused on making production agents more reliable by detecting failures, evaluating potential fixes, and helping teams build systems they can trust. This role owns the intelligence behind what gets flagged, how confidently it is identified, and whether a proposed fix works.
WHAT YOU'LL DO
- Detect subtle agent failures, including incorrect responses that appear compliant, omissions, and patterns that emerge across many traces.
- Turn indirect user signals, such as rephrasing, abandonment, and retries, into evidence that an agent has failed.
- Build evaluations that help measure detection quality and improve performance even when labeled ground truth is unavailable.
- Make patch generation trustworthy by reproducing failures, verifying fixes, and avoiding low-confidence changes.
- Improve the cost and quality tradeoffs of model-based evaluation, including when to use a smaller model or no model.
- Review real production traces regularly to identify problems and guide improvements.
WHAT WE'RE LOOKING FOR
- Experience can range from new graduate to approximately 10 years; strong ability to learn and improve is valued.
- A track record of shipping products or systems that people actually use.
- Hands-on experience with LLMs in production, including evaluations, LLM-as-judge approaches, and embeddings.
- Good research judgment and the ability to distinguish meaningful improvements from noise, including when ground truth is limited.
- Experience operating agents in production and learning from their failures is a plus.
COMPENSATION & BENEFITS
Salary range is $120,000 to $200,000 USD annually. Visa sponsorship is not available.
LOCATION
This is an on-site role based in San Francisco, United States.



