Skip to content

Open nowPosted 10 hours ago

Evaluations Team Lead

Fundamental21 open roles

Where
Europe
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowEvaluations Team LeadFundamental · Europe
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Fundamental's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.2% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.9%1 day
  2. 3.8%3 days
  3. 8.2%7 days
  4. 15.2%14 days
  5. 34.1%30 days
This job: posted 10 hours ago

The posting

ABOUT FUNDAMENTAL

Fundamental is an AI research lab pioneering the future of enterprise decision-making. Our flagship model, NEXUS is the world's most powerful Large Tabular Model (LTM) - purpose-built for the structured records that contain trillions of dollars in business value. With $275m in funding from leading investors and trusted by Fortune 100 companies, Fundamental is giving businesses the Power to Predict.

At Fundamental, you'll work on unprecedented technical challenges in foundation model development and build technology that transforms how the world's largest companies make decisions. This is your opportunity to be part of a category-defining company from the ground-up. Join the team defining the future of enterprise AI.

ABOUT THE ROLE

NEXUS is already in production, and we are working to substantially improve its predictive quality, latency, and cost efficiency. You will lead the team that owns how NEXUS is measured, building the shared evaluation platform that engineering, research, and Applied AI all rely on: consistent benchmarks, curated datasets, and standards for metrics, data splits, leakage prevention, and benchmark contamination that hold up under scrutiny.

Research needs to trust results before committing compute to it, engineering needs regressions caught before they reach production, and when a customer's own data science team benchmarks NEXUS against their own models, your platform is what Applied AI relies upon. You will benchmark NEXUS against competing approaches, using fair tuning budgets, data access, and latency measurement protocols, and turn what you find into research priorities and release recommendations.

This is a player-coach role: you will hire and manage a small team while staying hands-on with the code, the experiment design, and the methodology yourself. There is no evaluation function to inherit here - what you build becomes the standard the rest of the company measures NEXUS against.

KEY RESPONSIBILITIES

- Build a shared evaluation platform for engineering, research, and Applied AI. Continuously add and maintain models and curated datasets so teams can run benchmarks and investigate results independently.

- Define evaluation standards for metrics, data splits, leakage prevention, calibration, uncertainty, and benchmark contamination.

- Build reproducible pipelines with versioned inputs and artifacts, and integrate regression checks into research and release workflows.

- Benchmark NEXUS against competing approaches using fair tuning budgets, data access, compute, and latency measurement protocols.

- Support Applied AI’s customer POC evaluations with tooling, methodological guidance, and analysis.

- Measure predictive quality, latency, and cost across deployment configurations, task types, and dataset characteristics.

- Turn findings into research priorities, release recommendations, and evidence-backed customer improvement plans.

- Hire and develop the team, set priorities, and stay hands-on with code and experimental design.

MUST HAVE

- Experience owning evaluation for tabular ML systems used in production or consequential customer decisions.

- Strong statistical judgment: choosing metrics and validation schemes, estimating uncertainty, comparing models across datasets, and accounting for repeated experimentation.

- Practical experience finding leakage in preprocessing, feature construction, joins, temporal dependencies, and related entities across splits.

- Strong Python and SQL skills, familiarity with scikit-learn and gradient-boosted trees, and experience building reliable ML tooling or platforms used by other teams.

- Experience designing fair model comparisons, including hyperparameter search, resource budgets, and end-to-end latency measurement.

- Prior people management experience, including hiring, technical coaching, and performance feedback, while remaining technically involved.

- Clear written and spoken communication with researchers, engineers, and customer data scientists, including the willingness to challenge claims the evidence does not support.

NICE TO HAVE

- Experience evaluating tabular foundation models or AutoML systems.

- Experience measuring how model optimisations affect predictive quality and inference performance.

- Experience with relational or multi-table data, and Snowflake or Databricks environments.

- Experience evaluating automated or agent-driven ML workflows, including failures that aggregate metrics can hide.

BENEFITS

- Competitive compensation with salary and equity

- Comprehensive health coverage for you and your dependents

- Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys

- Relocation support for employees moving to join the team in one of our office locations

- A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Fundamental's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Fundamental's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Fundamental's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.