Skip to content

Research Engineer, Benchmarks

Clera

Singapore$150,000 – $250,000 a year

Applying for this one?

We write the CV against this exact posting — its wording, its requirements — not a template with your name in it.

Get my CV for this job

$25, one-time. No subscription.

ABOUT THE ROLE

This is a hands-on research engineering role focused on designing and owning high-quality benchmarks that evaluate frontier AI agents on realistic, domain-specific workflows. You will sit within a small, highly technical team and play a critical part in ensuring evaluations are rigorous, credible, and trusted by leading AI labs and customers.

WHAT YOU'LL DO

- Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.

- Partner with subject-matter experts to define realistic workflows and translate them into benchmark tasks and evaluation criteria.

- Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale.

- Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.

- Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.

- Write clear technical documentation and benchmark reports for research and engineering audiences.

WHAT WE'RE LOOKING FOR

- 2 to 4 years of experience in software engineering, ML engineering, or research roles, with a focused track record in AI benchmarks or evaluation infrastructure.

- Strong proficiency in Python, Docker, and Linux environments.

- Demonstrated experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models.

- Experience building infrastructure to reliably run AI models or agents against benchmark or evaluation tasks.

- Ability to analyze and model workflows across diverse technical or business domains to support task design.

- Sharp attention to detail with a habit of spotting subtle inconsistencies and edge cases.

- Comfort reasoning from first principles about task design, scoring, and failure modes.

- Strong written communication skills; experience producing technical documentation or benchmark reports.

- Ability to thrive in unstructured problem spaces at an early-stage startup.

- Bonus: experience with reinforcement learning pipelines, data generation, or RL agent evaluation; published work on AI benchmarking or model evaluation.

COMPENSATION & BENEFITS

Salary range: USD 150,000 to 250,000 annually. Visa sponsorship is available.

LOCATION

On-site in Singapore.

Seen 8 hours ago · Clera postings close after a median of 0 days.

Original posting on Clera's site ↗

Posting text belongs to the employer. Removal requests: contact us.

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

One job at a time

One posting. One CV. $25.

Pick the job you actually want and we write for it.

Get my CV for this job