The posting
ABOUT THE ROLE
Join a small, technical research and engineering team building rigorous benchmarks for evaluating AI agents on realistic, domain-specific workflows. You will own benchmark design and implementation, helping ensure evaluation results are reliable and useful to research and industry teams.
WHAT YOU'LL DO
- Design, implement, and maintain benchmarks for evaluating AI agents on domain-specific tasks.
- Collaborate with subject-matter experts to turn real workflows into realistic tasks and evaluation criteria.
- Build reliable infrastructure to run models and agents against evaluation tasks at scale.
- Develop metrics and analyses to assess benchmark difficulty, reliability, and failure modes.
- Validate how benchmark results relate to real-world performance and evaluation needs.
- Write clear technical documentation and reports for research and engineering audiences.
WHAT WE'RE LOOKING FOR
- Two to four years of experience in software engineering, machine learning engineering, or research, including at least two years focused on AI benchmarks, evaluations, or agent environments.
- Hands-on experience designing, implementing, and operating benchmarks or evaluation infrastructure for AI agents or large language models.
- Proficiency with Python, Docker, and Linux environments.
- Experience working with subject-matter experts to model workflows across technical or business domains and define evaluation criteria.
- Experience developing metrics, statistical analyses, or validation studies for benchmark quality and real-world relevance.
- Strong technical writing, attention to detail, and ability to work independently in an early-stage environment.
- Experience with reinforcement learning pipelines, data generation, or agent evaluation is useful. Published work or technical writing on AI evaluation is also valued.
COMPENSATION & BENEFITS
Annual salary range: $100,000 to $170,000 USD. Visa sponsorship is available.
LOCATION
On-site in Singapore, Singapore.



