Skip to content

ML Research Engineer

whitecircle

ParisHybrid

Applying for this one?

We write the CV against this exact posting — its wording, its requirements — not a template with your name in it.

Get my CV for this job

$25, one-time. No subscription.

TL;DR: We are looking for several ML Engineers to train, post-train, and evaluate the LLMs at the core of our platform. This is hands-on modern model training work: large-scale data pipelines, SFT/RLHF/DPO-style alignment, reward models, distributed multi-GPU training, and evaluation.

About us

White Circle https://whitecircle.ai/ is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies – simple natural-language rules that define what an AI model should and shouldn’t do. We automatically test, enforce, and continuously improve these policies at scale.

- We’ve raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others

- We process over 100M+ API calls every month

- We fine-tune and train our own LLMs so they run faster and cheaper than any open or proprietary model

We’re a small, highly focused team. If you want to work deeply on hard problems, see your work ship to production quickly, and influence how AI safety is actually built – you’re the one we need.

What you’ll do

- Turn petabytes of unstructured text into a structured, explorable view (topics, clusters, segments, trends, anomalies): iterate from “unknown unknowns” to stable definitions we can track.

- Build scalable representation pipelines: sampling strategies, preprocessing/normalization, embeddings at scale, indexing, and retrieval to make the corpus searchable and analyzable.

- Use LLMs pragmatically: labeling/classification, weak supervision, data enrichment, summarization, and automated diagnostics of inbound volumes (with cost/quality controls).

- Deliver insights that change decisions: translate findings into product and operational actions (what data we have, what’s missing, where quality breaks, what to prioritize next).

- Ship self-serve analytics: datasets, data models, and lightweight tools/dashboards so the team can explore and answer questions without ad-hoc requests.

- Partner closely with engineering/research: align pipelines with production constraints (latency/cost/privacy), and integrate outputs into workflows.

You'll fit right in if you

- Have strong Python + SQL with an engineering mindset: you can build reliable pipelines, not just notebooks.

- Have solid applied NLP/ML experience on real-world text: embeddings, clustering, topic modeling, semantic search, classification; you understand failure modes and how to debug them.

- Are comfortable at scale: distributed processing, large-scale storage-querying, and performance-cost tradeoffs.

- Know how to evaluate fuzzy problems: offline/online metrics, human-in-the-loop labelling, inter-annotator agreement, drift monitoring, and reproducibility.

- Have prior work with safety/moderation datasets, policy/rule systems, or high-volume logging/observability

A big plus

- A public builder footprint: open-source models, datasets, or training frameworks on HuggingFace/GitHub, benchmarks, papers (workshop or main conference), or technical posts with real usage

- Experience training models at a frontier or near-frontier lab, or leading open-source model releases with documented adoption

- Experience with RL methods for LLMs beyond standard RLHF: online RL, GRPO-style methods, or novel alignment approaches

- Experience with moderation, safety, or classification models at scale

- Multilingual model training experience

Compensation & benefits

- Competitive compensation, including equity

- Flexible time off

- Office in central London/Paris with flexible hybrid setup

- Relocation support if you’re moving to Paris, available after your probationary period

- Premium private health insurance

- Mental health support, including coverage for therapy when you need it

- Lunch and dinner covered when you work from the office

- Learning and development support for courses, conferences, and opportunities to grow your skills

- All the hardware, subscriptions, tools, and services you need

- Team off-sites twice a year: we’ve recently been to the Alps, Saint-Tropez, and Marbella

Process

1. Intro call with Talent Team

2. Test assignment

3. Technical interview with Head of Applied Research

4. Final conversation with our CEO

Seen 7 hours ago.

Original posting on whitecircle's site ↗

Posting text belongs to the employer. Removal requests: contact us.

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

One job at a time

One posting. One CV. $25.

Pick the job you actually want and we write for it.

Get my CV for this job