Skip to content

Open nowPosted 79 days ago

AI Training Infrastructure Engineer

Designworks Talent37 open roles

Where
Bellevue
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowAI Training Infrastructure EngineerDesignworks Talent · Bellevue
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Designworks Talent's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market. Designworks Talent postings stay open a median of 9 days.

Share of postings closed within
  1. 1.8%1 day
  2. 3.5%3 days
  3. 8.1%7 days
  4. 15.1%14 days
  5. 33.9%30 days
This job: posted 79 days ago

Designworks Talent median: 9 days open

The posting

AI TRAINING INFRASTRUCTURE ENGINEER

Location: Hybrid | Bellevue, WA (downtown) Titles: Senior and Staff (multiple roles available)

BUILD THE TRAINING INFRASTRUCTURE POWERING NEXT-GENERATION AI MODELS

ABOUT THE OPPORTUNITY

A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.

Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads.

We're seeking AI Training Infrastructure Engineers to build and scale the distributed systems that power large-scale AI model training. This team focuses on reliability, efficiency, and operational excellence across GPU clusters, enabling researchers and engineers to train and deploy advanced AI models at scale.

THE OPPORTUNITY

This is a foundational engineering role focused on building the infrastructure layer behind large-scale AI training workloads. You'll work on distributed training systems, GPU clusters, model pipelines, and the tooling required to make AI development more reliable, efficient, and scalable.

You'll collaborate closely with infrastructure, orchestration, performance, and machine learning teams to solve complex challenges around distributed computing, fault tolerance, training efficiency, and production readiness.

This opportunity is ideal for engineers who enjoy building highly scalable systems and working at the intersection of AI research, infrastructure engineering, and distributed computing.

WHAT YOU'LL DO

- Build and scale distributed training infrastructure supporting large AI models across large GPU clusters.

- Design and improve systems that increase training reliability, efficiency, and resource utilization.

- Develop solutions for fault tolerance, checkpointing, recovery, and large-scale training operations.

- Integrate AI models into production training pipelines in partnership with platform, orchestration, and performance engineering teams.

- Diagnose and resolve issues impacting training throughput, stability, reliability, and cost efficiency.

- Build tools and automation that improve the developer experience for AI researchers and engineers.

- Establish best practices for training infrastructure, operational processes, and platform reliability.

- Contribute to the evolution of the AI infrastructure platform as an early member of the engineering team.

WHAT WE'RE LOOKING FOR

- Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure.

- Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems.

- Strong understanding of the reliability, scalability, and efficiency challenges associated with multi-node GPU training.

- Experience integrating training systems with production machine learning pipelines.

- Strong programming skills and experience working with complex distributed systems.

- Ability to independently own technically challenging projects in a fast-moving engineering environment.

- Comfortable operating with high ownership and limited process overhead.

PREFERRED QUALIFICATIONS

- Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies.

- Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows.

- Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment.

- Experience optimizing GPU utilization, training performance, or distributed system reliability.

- Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms.

COMPENSATION

- Competitive base pay for Bellevue market

- Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance

- U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays

LOCATION

- Hybrid role based in downtown Bellevue, WA.

- Approximately three days per week in the office.

- Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.

- U.S. work authorization is required. Visa sponsorship is not currently available.

WHY JOIN?

- Build the infrastructure powering the next generation of AI models and applications.

- Work directly on distributed training systems, GPU clusters, and large-scale AI platforms.

- Solve some of the industry's most challenging problems around AI scalability, reliability, and efficiency.

- Join early enough to influence architecture, tooling, and engineering practices.

- Collaborate with a highly experienced team building critical AI infrastructure from the ground up.

- Enjoy the ownership and technical impact of a startup environment backed by significant long-term investment.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Designworks Talent's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Designworks Talent's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Designworks Talent's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.