Skip to content

Open nowPosted 5 days agoWe saw it 50 min after it went up

Research Scientist, Applied White-Box Methods

FAR.AI16 open roles

Pay
$150,000 – $250,000 a year
Where
Berkeley Office
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowResearch Scientist, Applied White-Box MethodsFAR.AI · Berkeley Office
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on FAR.AI's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market. FAR.AI postings stay open a median of 6 days.

Share of postings closed within
  1. 1.7%1 day
  2. 3.5%3 days
  3. 7.8%7 days
  4. 14.6%14 days
  5. 34.1%30 days
This job: posted 5 days ago

FAR.AI median: 6 days open

The posting

ABOUT US

FAR.AI http://FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.

Since our founding in July 2022, we've grown to 50+ staff https://www.far.ai/about/team, published 40+ academic papers https://scholar.google.com/citations?user=FVJ24k8AAAAJ, and convened leading AI safety events https://far.ai/events/. Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML including a Best Paper Honorable Mention in 2026 https://icml.cc/virtual/2026/oral/71065, and ICLR, and features in the Financial Times https://www.ft.com/content/175e5314-a7f7-4741-a786-273219f433a1, Nature News https://www.nature.com/articles/d41586-024-02218-7, Wired Magazine https://www.wired.com/story/jailbreaking-ai-models-google-anthropic-openai-spacexai/ and MIT Technology Review https://www.technologyreview.com/2020/02/28/905615/reinforcement-learning-adversarial-attack-gaming-ai-deepmind-alphazero-selfdriving-cars/. We conduct pre-deployment testing on behalf of frontier developers such as OpenAI, independent evaluations for governments including the EU AI Office https://www.far.ai/news/far-ai-selected-to-lead-eu-ai-act-cbrn-risk-consortium, and publish the AI Security Leaderboard https://leaderboard.far.ai/ based on our red-teaming expertise. We help steer and grow the AI safety field through developing https://arxiv.org/abs/2405.06624 research https://arxiv.org/abs/2506.20702 roadmaps https://www.researchgate.net/publication/396910034_Open_Technical_Problems_in_Open-Weight_AI_Model_Risk_Management with renowned researchers such as Yoshua Bengio; running FAR.Labs https://www.far.ai/programs/far-labs, an AI safety-focused co-working space in Berkeley housing 40+ members; and supporting the community through targeted grants https://www.far.ai/programs/grantmaking to technical researchers.

ABOUT THE APPLIED WHITE-BOX METHODS TEAM

The Applied White-Box Methods team develops, evaluates, and demonstrates methods that leverage model internals to improve the safety of AI systems. We work on diverse AI safety applications of white-box methods, from white-box control to evaluation awareness to shaping training dynamics to improve alignment.

Black-box methods, such as chain-of-thought monitoring, work well for now, but we are quickly entering a world where black-box interventions and monitoring are insufficient. Interpretability research is still often early-stage, curiosity-driven work without realistic evaluations on applications that matter. The team bridges the gap between exploration and deployment by stress-testing white-box methods on real-world (e.g. long-context agentic coding) tasks and using this feedback loop to enable the deployment of better white-box methods at frontier scale.

Our scope includes any method that uses model internals to understand, predict, or intervene on model behavior, not only what is conventionally called interpretability. Methods of current interest include natural language autoencoders and other activation explainers, activation oracles, steering, patching, and other activation-level interventions, influence functions and data attribution, and singular learning theory. We evaluate our methods against real baselines – strong black-box methods and activation probes – to be able to make an honest case that the methods are worth implementing, or conclude that simpler methods work better for now.

AI Research Automation. AI will soon automate most of the hill-climbing in the research process. We anticipate this by focusing our effort and judgement on defining realistic evaluations with Goodhart-resistant metrics. With the evaluation framework properly built, we can pour vast amounts of AI labor into method development and iteration without overfitting. We believe this is the way to scale white-box research into the age of RSI.

Realistic, large-scale models. Unrealistic models lead to unconfident conclusions about which methods do and do not work. We leverage FAR's shared compute and infrastructure to work on organisms that come out of pipelines a frontier lab could plausibly have run: for example, reward seekers trained by RL in broken environments, with other contaminated data mixed in to induce other misalignments.

Realistic, large-scale evaluations. We will primarily study long-context agentic coding as this is where most of the risk currently lies. In addition to the classic AI control sabotage settings, we will also study reward hacking, sandbagging, and research tampering, which are some of the key failure modes that matter during RSI.

Practical monitors and interventions. Deployment of new methods has real costs for AI developers. Part of our focus on real-world applications is ensuring that methods are simple and efficient enough to deploy at frontier scale.

ABOUT THE ROLE

As a Research Scientist on the team, you will take ownership of and accelerate the team's research agenda, publish findings broadly, and engage with the AI alignment community. You are encouraged to propose new directions within the team's agenda. You are welcome and encouraged to attend relevant conferences and other community events, and leverage FAR.AI http://FAR.AI’s existing comprehensive infrastructure for events convening and government relations as you see fit. Beyond FAR.AI http://FAR.AI, you can work with national AI safety institutes, frontier model developers, and top academics.

ABOUT YOU

Two backgrounds are particularly well suited to this role: researchers with some experience in more fundamental interpretability who want to evaluate and refine those methods in realistic settings, including agentic coding, long contexts, and realistic threat models; and researchers from evaluations, AI control, red-teaming, or reinforcement learning who have begun working with model internals and want to deepen that work. In both cases, prior hands-on experience with white-box methods and a demonstrated interest in their practical application are preferred. (Note that this role is not well suited to researchers whose primary interest is interpretability as a foundational science.)

If you are new to AI Safety research, that is also ok! But please be prepared to explain how your previous research experience (e.g. in applied ML) could be leveraged for our work, and how you are engaging with the field of technical AI safety research today.

We are hiring across a range of levels of experience. You may already have some of the following, though not necessarily all.

- Hands-on experience applying at least one white-box method to a real model (activation explainers, SAEs, steering, attribution, influence functions, probes, or similar), and an informed view of its limitations.

- A track record in AI safety: a paper, a fellowship project, or substantive public writing. Alumni of programs such as MATS, Astra, Anthropic Fellows, SPAR, Pivotal, LASR or similar programs are especially encouraged to apply.

- Experience with evaluations, AI control, red-teaming, reinforcement learning, or post-training of LLMs. Experience running large training runs is a plus.

- The ability to communicate novel methods and results clearly to technical and non-technical audiences.

- A PhD or several years of research experience in computer science, machine learning, physics, statistics, or a related field.

- Previous experience in applied ML for other fields: e.g. biology, chemistry, materials science, robotics, etc.

BENEFITS*

- Health Insurance - 94% of Insurance premium paid by Organization commencing within 1 month after your start date

- Retirement - 401(k) plan with up to 2% match

- PTO - 25 days Paid Time Off per year, accrued weekly and up to 10 days of paid sick leave per year

- Paid Leave - Paid Bereavement, Family, Medical and Pregnancy Disability Leave

- WFH Stipend & Equipment - Work computer and stipend provided for eligible employees

- Catered Meals (Berkeley Office Only) - Catered lunches and dinners on workdays at our office

*(Available only to full-time employees located in the US)

LOGISTICS

- If based in the USA, you will be an employee of FAR.AI http://FAR.AI, a 501(c)(3) research non-profit. Outside the USA, you will be an employee of an EoR organization on behalf of FAR.AI http://FAR.AI.

- Location: Berkeley, CA. This role is in-person; we strongly prefer candidates who are in the Bay Area or willing to relocate, and we sponsor visas for in-person employees. We will consider remote arrangements for exceptional candidates.

- Hours: full-time (40 hours/week).

- Compensation: $150,000–$250,000/year depending on experience and location, plus work-related travel and equipment expenses. Catered lunch and dinner at our Berkeley office.

- Application process: interviews with members of our technical staff, followed by a paid work trial of up to one week. If you are not available for a work trial we may be able to find alternative ways of assessing fit.

If you have any questions about the role, feel free to contact us at [email protected]. Otherwise, if you don't have questions, the best way to ensure a proper review of your skills and qualifications is by applying directly via the application form. Please don't email us to share your resume (it won't have any impact on our decision). Thank you!

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against FAR.AI's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on FAR.AI's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    FAR.AI's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.