The posting
ABOUT THE ROLE
Lead the data quality team at an early-stage AI company, building systems to measure and improve training data for AI agents. Your work will shape how teams evaluate data across reinforcement learning environments, synthetic datasets, benchmarks, and domain-specific workflows.
WHAT YOU'LL DO
- Set data quality strategy and build systems to evaluate thousands of tasks across varied data sources and workflows.
- Develop scalable synthetic data validation methods, including failure-mode analysis, task mutation checks, and trajectory auditing.
- Design metrics, experiments, quality standards, and processes for evaluating agent outputs.
- Work with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation.
- Turn research insights into production tools, dashboards, validation pipelines, and feedback loops.
- Mentor research engineers and strengthen technical rigor across the team.
WHAT WE'RE LOOKING FOR
- At least 5 years of research or engineering experience building AI/ML data evaluation or quality systems.
- Advanced proficiency in Python, Docker, and Linux.
- Experience leading technical projects or teams from problem definition through implementation and iteration.
- Strong understanding of AI evaluations, post-training, and the qualities that make agent training data realistic, learnable, diverse, reliable, and useful.
- Experience validating synthetic data at scale and translating domain expertise into scalable review or generation workflows.
- Clear written communication and comfort working independently in an early-stage environment.
COMPENSATION & BENEFITS
Annual salary range: $130,000 to $225,000. Visa sponsorship is available.
LOCATION
On-site in the United States. Singapore is also listed as a work location.



