Skip to content

Research Engineer, Benchmarks

Clera

Singapore$150,000 – $250,000 a year

Applying for this one?

We write the CV against this exact posting — its wording, its requirements — not a template with your name in it.

Get my CV for this job

$25, one-time. No subscription.

ABOUT THE ROLE

This is a core technical role on a small, high-caliber team building rigorous benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You will own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand real-world agent performance. The work is critical to the credibility and impact of the company's benchmark platform.

WHAT YOU'LL DO

- Design, implement, and maintain the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.

- Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks.

- Build reliable infrastructure to run models and agents against benchmark tasks at scale.

- Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.

- Validate that benchmark performance correlates with real-world evaluations and customer expectations.

- Write clear technical documentation and benchmark reports for research and engineering audiences.

WHAT WE'RE LOOKING FOR

- 2 to 4 years of experience in research engineering or machine learning engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments.

- Strong proficiency in Python, Docker, and Linux for building research or production infrastructure.

- Hands-on experience designing and running benchmarks or evaluation environments for AI agents or large language models.

- Experience developing metrics and validation studies to assess benchmark difficulty, reliability, and real-world correlation.

- Experience collaborating with domain experts to turn workflows into concrete evaluation criteria.

- Strong technical writing skills; published papers or blog posts on AI benchmarking, model evaluation, or failure modes are a plus.

- Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a plus.

- Background at a frontier AI lab, research institution, or on a widely used public benchmark project is a plus.

- Comfort working independently in fast-paced, early-stage environments with unstructured problem spaces.

- Sharp attention to detail and the ability to reason from first principles about task design, scoring, and edge cases.

COMPENSATION & BENEFITS

Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.

LOCATION

On-site in Singapore.

Seen 25 hours ago · Clera postings close after a median of 0 days.

Original posting on Clera's site ↗

Posting text belongs to the employer. Removal requests: contact us.

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

One job at a time

One posting. One CV. $25.

Pick the job you actually want and we write for it.

Get my CV for this job