Skip to content

Open nowPosted 210 days ago

Senior Distributed Systems Engineer

ifm-us48 open roles

Pay
$200,000 – $400,000 a year
Where
Sunnyvale, CA
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Distributed Systems Engineerifm-us · Sunnyvale, CA
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on ifm-us's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.7% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.4%1 day
  2. 3.5%3 days
  3. 7.7%7 days
  4. 13.4%14 days
  5. 34.5%30 days
This job: posted 210 days ago

The posting

About the Institute of Foundation Models

The Institute of Foundation Models (IFM) designs and operates ultra-scale GPU supercomputing systems to train next-generation foundation models. We believe performance, fault tolerance, and scalability are co-designed across model architecture, communication systems, runtime, and hardware topology.

This role sits at the core of that effort — driving communication performance, distributed reliability, and cross-layer optimization for large-scale training workloads.

The Mission

We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including hybrid parallelism and Mixture-of-Experts (MoE) workloads.

This is not a network operations role. This is a systems-level engineering position focused on performance engineering, distributed debugging, and communication-runtime co-design.

· Design and optimize expert-parallel and hybrid-parallel communication patterns

· Drive high-performance hierarchical collectives for MoE workloads

· Co-design runtime orchestration with communication topology awareness

· Reduce tail latency and improve determinism across thousands of GPUs

· Architect fault-tolerant distributed execution under real-world cluster failures

Core Technical Scope

· Communication-compute overlap and topology-aware collective optimization

· Deep debugging of NCCL, RDMA, and custom communication layers

· Hybrid expert parallel strategies in modern large-scale MoE systems

· Elastic and resilient distributed job orchestration concepts

· Congestion analysis and routing optimization across InfiniBand/RoCE fabrics

· Microbenchmarking and performance modeling for communication-heavy workloads

Expected Technical Depth

· Hybrid expert parallel communication for Mixture-of-Experts training

· Scaling behavior under network pressure

· Distributed orchestration for elastic, large-scale training

· Fault detection and recovery in distributed GPU workloads

· Cross-layer bottlenecks: GPU ↔ NIC ↔ PCIe ↔ NVSwitch ↔ Fabric ↔ Scheduler

Required Background

· Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth)

· Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA

· Deep familiarity with NCCL and/or UCX internals

· Strong systems programming ability (C/C++, Rust, or Go)

· Strong familiarity with modern model training frameworks such as PyTorch

· Ability to troubleshoot and profile training performance issues related to communication bottlenecks

· Ability to translate research ideas into production-grade optimizations

· Experience debugging distributed hangs, desynchronization, and performance regressions

What We Mean by "Hardcore"

· You can explain why an communication degrades at scale and how to fix it

· You have improved real cluster throughput via communication redesign

· You can trace a distributed hang across ranks and identify the root cause

· You are comfortable working at the boundary between hardware and runtime

Application Requirements

· Include a link to your GitHub (required)

· Provide links to relevant distributed systems, HPC, or large-scale training projects

· Include a list of publications and/or public technical reports (if applicable)

· Describe the hardest distributed debugging problem you solved

· Include measurable performance improvements you have delivered

Academic Qualifications

Master’s, or Bachelor’s + 1 year of relevant experience.

Visa Sponsorship

This position is eligible for visa sponsorship.

Benefits Include

*Comprehensive medical, dental, and vision benefits

*Bonus

*401K Plan

*Generous paid time off, sick leave and holidays

*Paid Parental Leave

*Employee Assistance Program

*Life insurance and disability

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against ifm-us's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on ifm-us's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    ifm-us's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.