Skip to content

Open nowPosted 5 days ago

Senior AI Inference Engineer

Workable (global search)107,990 open roles

Where
United States
Work mode
Remote
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior AI Inference EngineerWorkable (global search) · United States
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Workable (global search)'s own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.9% of postings close within 7 days. Measured by our own scanner across the market. Workable (global search) postings stay open a median of 6 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 7.9%7 days
  4. 14.9%14 days
  5. 34.0%30 days
This job: posted 5 days ago

Workable (global search) median: 6 days open

The posting

About StackYak

StackYak is building the infrastructure layer for AI.

We are an early-stage, funded company building software that brings compute, GPU infrastructure, networking, and inference together into one product. The opportunity is large, the market is moving quickly, and we are building for real production workloads from the start.

This is not internal IT. This is not a slow-moving infrastructure team maintaining someone else's platform. The infrastructure is the product.

We are a small, senior team with very little bureaucracy. This is a founding role in its discipline. You will be the first person here whose primary responsibility is the systems layer beneath the product, and the shape it takes will largely be yours to decide.

Treat this document as a starting point rather than a boundary. The people who do well here take ground early and are not asked to give it back.

We move quickly. We do not have months for someone to learn the fundamentals of their discipline. You should already be very good at what you do, be able to ramp into adjacent areas quickly, and be comfortable operating without perfect requirements or neatly defined boundaries.

The Role

We need someone who can take a model, a set of GPUs, and a production requirement and determine how that model should actually run.

Hand you a frontier open-weight model — dense or mixture-of-experts, tens to hundreds of billions of parameters — along with the hardware actually available and the latency a customer actually needs, and you should be able to make the calls that follow: precision, quantization, tensor and pipeline and expert parallelism, memory strategy, batching, concurrency, topology, and serving runtime. Then defend them, measure them, and operate them.

Doing that well is the baseline. What makes the role interesting is the question sitting underneath it.

Serving expertise is scarce, slow to acquire, and currently lives in the heads of a small number of people who have done it enough times to have the instinct. We do not accept that it has to stay that way. A great deal of what those people know is reasoning, and reasoning can be written down, tested, and eventually carried out by something other than a person at two in the morning. How far that can be pushed is genuinely open, and you would be one of the people finding out.

This is not a research position and it is not an architecture-only role. You will benchmark, deploy, debug, tune, automate, and operate real inference systems — and then carry them in production.

What You Will Own

Not tasks. Outcomes, and the authority that comes with them.

  • How models run. Placement, parallelism, precision, and memory strategy across single-GPU, multi-GPU, and multi-node — and the reasoning behind each choice, in a form someone other than you can audit.
  • The serving runtimes. vLLM, SGLang, TensorRT-LLM, and whatever displaces them: which one we use where, and what we do when the one we picked turns out to be wrong.
  • What we are willing to promise. Concurrency limits, latency and TTFT targets, throughput under load. You own the evidence behind a number before it becomes a commitment to a customer.
  • Benchmarking as an institution. Methodology, harness, reproducibility — and the standing to say that a result does not mean what someone wants it to mean.
  • Model lifecycle in production. Loading, startup, health, upgrades, capacity, and failure handling, all of it happening while customers are connected.
  • Multi-tenant versus dedicated serving, and where the line between them actually falls.
  • Inference incidents, including the ones where the serving runtime turns out to be innocent and the cause is GPU, driver, host, or network.
  • Turning all of the above into software, so that it stops being institutional memory with a single point of failure.

What Success Looks Like

You can be given a model, a hardware target, and a workload profile and quickly produce a deployment that is sensible, measurable, reproducible, and ready to operate.

Then you help us make that expertise programmatic.

You will help us answer questions such as:

  • What hardware can this model run on?
  • What is the right precision or quantization strategy?
  • How should we split it across GPUs or nodes?
  • What serving runtime should we use?
  • How much concurrency can we safely offer?
  • How should we trade latency against throughput for this particular workload?
  • How does StackYak make those decisions automatically instead of needing an expert every time?

Requirements

What We Need

The bar is what you have already done, not what you could learn. You should have done most of this:

  • Run large language models in production, under real load, with someone depending on them.
  • Served a model across multiple GPUs, and across multiple nodes.
  • Sized a model against hardware and been right about whether it would fit.
  • Chosen a quantization and precision strategy with a real quality and cost consequence attached to the choice.
  • Tuned batching and concurrency past the point where the easy wins ran out.
  • Measured latency, TTFT, throughput, utilization, and cost — and defended the numbers to someone motivated to disbelieve them.
  • Worked in NVIDIA and/or AMD inference environments.
  • Written Python that other people run in production: tooling and product, not scripts.
  • Debugged Linux and systems problems below the container boundary.
  • Found the root cause of an inference failure that was not in the serving runtime.

You Will Be Especially Strong If

  • You have run inference infrastructure at an AI company, inference provider, GPU cloud, hyperscaler, or serious internal AI platform.
  • You have worked across multiple GPU generations and vendors, and can say concretely how that changed your decisions.
  • You understand distributed inference and the networking implications of multi-node serving.
  • You have benchmarked and compared serving runtimes rather than treating one framework as the answer to every problem.
  • You can explain why a deployment is configured the way it is instead of repeating a vendor recipe.
  • You have automated deployment decisions, or built schedulers, placement systems, capacity planners, or similar infrastructure.
  • You build and test things outside your assigned roadmap because you want to know how they actually work.

This Is Probably Not For You If

  • Your background is primarily model training or ML research, with little production serving behind it.
  • You have put a model behind an API but have never had to reason about GPU memory, parallelism, topology, and concurrency.
  • You treat vLLM defaults as an architecture.
  • You are interested in benchmarks but not in operating the systems they describe.
  • You want to optimize kernels all day and have little interest in reliability or productization.
  • You would rather stay the person who knows how to configure it than encode that knowledge into software that makes you unnecessary.
  • You prefer writing recommendations to implementing them.
  • You need narrowly defined ownership, or a long runway before taking responsibility.

Benefits

How We Work

  • Small, senior team with direct access to the founders.
  • Strong opinions, loosely held.
  • Everyone is expected to participate in technical decisions.
  • Everyone shares responsibility for production and on-call.
  • We value people who can move between design, implementation, debugging, and operations.
  • We care much more about what you have built and operated than degrees, certifications, or academic credentials.
  • We expect people to leave ego at the door, argue the technical case, make a decision, and then execute.
  • We are remote and distributed across time zones. Whether a role is an employment or a contract engagement depends on where you are, and we work that out at offer.
  • Hiring here is a few real conversations with the people you would actually work with, not a recruiter screen followed by a panel of strangers.
  • We are hiring across inference, infrastructure, and networking. The boundaries between the three are blurry on purpose. If you sit between two of them, say so.
  • This is an early-stage startup. The pace is high, the problems are hard, and the scope will change as we grow.

Compensation

Competitive compensation plus meaningful equity. Exact structure will depend on location, engagement model, and experience.

A Note For Agencies

We are not using external recruiters or agencies for this role, and we will not be persuaded otherwise by an email. We do not want your spam. We will not read the CVs you send, we will not reply to your follow-up, and no candidate you put in front of us creates a fee obligation of any kind. Do not contact us.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Workable (global search)'s own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Workable (global search)'s form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Workable (global search)'s answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.