Skip to content

Open nowPosted today

Software Engineer, Production Inference (Distributed Inference) — Thinking Machines Lab

Lightspeed portfolio6,214 open roles

Pay
$350,000 – $500,000 a year
Where
San Francisco, California, United States; San Francisco
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSoftware Engineer, Production Inference (Distributed Inference) — Thinking Machines LabLightspeed portfolio · San Francisco, California, United States; San Francisco
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Lightspeed portfolio's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.3% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.9%1 day
  2. 3.9%3 days
  3. 8.3%7 days
  4. 15.3%14 days
  5. 34.2%30 days
This job: posted today

The posting

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer to build and scale the distributed production inference systems that serve Inkling, Inkling-Small, and Tinker in production. You'll own the systems that turn trained models into fast, reliable, cost-efficient services — from request routing and batching to multi-node serving and GPU utilization at scale.

This is a systems-heavy, production-first role. You'll work closely with research and infrastructure teams to translate rapidly evolving model architectures into serving systems that meet real-world latency, throughput, and reliability requirements, and you'll be on the front line when production inference systems need to scale, recover, or improve.

What You'll Do

  • Design, build, and operate distributed infrastructure for large-scale model serving, including request routing, load balancing, batching, and multi-node coordination
  • Optimize inference latency and throughput in production, including work on KV cache management, continuous batching, speculative decoding, and quantization
  • Build and maintain high-concurrency serving systems with strong uptime, low tail latency, and deep observability
  • Benchmark, tune, and extend inference engines to support new model architectures as they move from research into production
  • Partner with research and infrastructure teams to translate emerging model designs into production-ready serving systems
  • Build tooling for tracing, debugging, and resolving issues across the serving stack, from orchestration down to GPU kernels
  • Participate in on-call rotation to support production inference systems

Skills & Qualifications

  • 3+ years of experience building and operating distributed systems in production
  • Strong systems programming skills in Python, C++, Rust, or similar languages
  • Experience with production infrastructure at scale: reliability, observability, and performance under real-world load
  • Solid understanding of networking, concurrency, and distributed systems fundamentals

Preferred Qualifications

  • Experience with LLM inference engines such as vLLM, SGLang, or TensorRT-LLM
  • Familiarity with GPU programming (CUDA) or low-level performance optimization
  • Experience with model parallelism, tensor/pipeline parallelism, or other distributed inference techniques
  • Track record of operating large-scale production systems with strict latency and uptime requirements
  • Experience with Kubernetes or similar orchestration systems for GPU workloads
  • Contributions to open-source ML systems or inference infrastructure projects

Logistics

  • Location: This role is based in San Francisco, CA.
  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $500,000 USD (placeholder — verify against current internal bands before publishing).
  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Lightspeed portfolio's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Lightspeed portfolio's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Lightspeed portfolio's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.