Skip to content

Open nowPosted 6 days ago

Senior Inference Engineer

Designworks Talent37 open roles

Where
Bellevue
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Inference EngineerDesignworks Talent · Bellevue
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Designworks Talent's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market. Designworks Talent postings stay open a median of 9 days.

Share of postings closed within
  1. 1.8%1 day
  2. 3.5%3 days
  3. 8.1%7 days
  4. 15.1%14 days
  5. 33.9%30 days
This job: posted 6 days ago

Designworks Talent median: 9 days open

The posting

INFERENCE ENGINEER

Location: Hybrid | Bellevue, WA (downtown) Titles: Senior and Staff (multiple roles available)

BUILD THE INFERENCE PLATFORM POWERING NEXT-GENERATION AI APPLICATIONS

ABOUT THE OPPORTUNITY

A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads—including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.

Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads.

We're seeking Inference Engineers to build and operate the model-serving systems behind a next-generation AI inference platform. This team focuses on delivering high-throughput, low-latency, reliable inference experiences that enable customers to consume advanced AI capabilities through production-scale APIs.

THE OPPORTUNITY

This is a foundational engineering role focused on building the systems that bring AI models from research environments into reliable production services. You'll work on the infrastructure layer responsible for serving large models efficiently, optimizing performance, and ensuring reliability as usage scales.

You'll collaborate closely with GPU performance, AI training infrastructure, platform engineering, and operations teams to solve complex challenges around model serving, latency optimization, resource efficiency, and production reliability.

This opportunity is ideal for engineers who enjoy working at the intersection of distributed systems, machine learning infrastructure, GPU computing, and large-scale production systems.

WHAT YOU'LL DO

- Build and operate production-grade model-serving and inference systems supporting high-throughput, low-latency AI workloads.

- Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads.

- Design systems that maximize GPU utilization while maintaining predictable performance and reliability.

- Improve the scalability and operational maturity of inference platforms as customer demand grows.

- Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving.

- Develop monitoring, alerting, and operational practices to maintain reliable inference services.

- Investigate and resolve performance, reliability, and capacity challenges across inference workloads.

- Contribute to architecture decisions and engineering standards as the platform evolves.

WHAT WE'RE LOOKING FOR

- Experience building and operating production machine learning inference or model-serving systems at scale.

- Strong understanding of the performance trade-offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency.

- Experience designing reliable distributed systems or production infrastructure.

- Understanding of GPU-backed AI workloads and the challenges of scaling inference systems.

- Strong engineering fundamentals and the ability to independently own complex technical problems.

- Comfortable working in a fast-moving environment where systems and processes are being built from the ground up.

PREFERRED QUALIFICATIONS

- Experience with modern inference-serving frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or similar technologies.

- Experience optimizing LLM inference workloads or large-scale AI serving platforms.

- Background operating API-based AI products or high-volume production services.

- Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms.

- Familiarity with model optimization techniques such as quantization, batching, caching, or performance tuning.

- Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure organization.

COMPENSATION

- Competitive base pay for Bellevue market

- Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance

- U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.

LOCATION

- Hybrid role based in the Bellevue, WA area.

- Approximately three days per week in the office.

- Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.

- U.S. work authorization is required. Visa sponsorship is not currently available.

WHY JOIN?

- Build the inference platform powering the next generation of AI applications.

- Work directly on large-scale model serving, GPU optimization, and production AI systems.

- Solve complex challenges around latency, throughput, reliability, and cost efficiency.

- Join early enough to influence architecture, tooling, and engineering practices.

- Collaborate with a highly experienced team building critical AI infrastructure from the ground up.

- Enjoy the ownership and technical impact of a startup environment backed by significant long-term investment.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Designworks Talent's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Designworks Talent's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Designworks Talent's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.