Skip to content

Open nowPosted today

Inference Optimization Engineer — Hedra

a16z portfolio7,203 open roles

Pay
$95,000 – $165,000 a year
Where
San Francisco, California, United States; San Francisco
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowInference Optimization Engineer — Hedraa16z portfolio · San Francisco, California, United States; San Francisco
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on a16z portfolio's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.8%1 day
  2. 3.5%3 days
  3. 8.1%7 days
  4. 15.1%14 days
  5. 33.9%30 days
This job: posted today

The posting

About Hedra

Hedra is the platform, models, and infrastructure for visual intelligence.

We build models and systems that push the frontier of visual intelligence, along with the infrastructure required to make those models fast, efficient, reliable, and accessible at scale.

We’re a small, highly technical team in San Francisco, backed by a16z and other leading investors. Researchers and engineers at Hedra work closely across boundaries, own problems end to end, and have significant influence over both what we build and how we build it.

The Role

We’re looking for an Inference Optimization Engineer to work alongside our research team on making state-of-the-art visual models fast and efficient at inference time.

You’ll work at the boundary between research and systems, taking new model architectures and figuring out how to run them efficiently on modern hardware. That means understanding where time and memory are being spent, identifying opportunities for algorithmic and systems-level improvements, and implementing optimizations across model architecture, inference algorithms, runtimes, kernels, and distributed execution.

The problems rarely live neatly within one layer of the stack. Depending on what you find, you might modify how a model executes, develop a new inference technique, write a custom GPU kernel, rethink memory movement, or change how work is distributed across accelerators.

We care more about technical depth, curiosity, and demonstrated ability than years of experience. We’re open to experienced ML systems engineers as well as exceptional early-career engineers or researchers who have already gone unusually deep on model performance, GPU systems, or efficient inference.

What You’ll Do

  • Work directly with research scientists and engineers to make new visual models fast and efficient at inference time.
  • Profile model architectures and workloads to understand bottlenecks across compute, memory, communication, and model execution.
  • Develop and implement new approaches to improving inference latency, throughput, memory efficiency, and GPU utilization.
  • Explore algorithmic optimizations including quantization, sparsity, caching, compilation, attention optimizations, and alternative execution strategies.
  • Build or optimize GPU kernels using CUDA, Triton, or similar technologies when existing implementations leave performance on the table.
  • Optimize model execution across single-GPU, multi-GPU, and multi-node environments.
  • Reason about the interaction between model architecture and hardware, and work with researchers when architectural changes can unlock meaningful performance improvements.
  • Investigate communication, memory movement, parallelism, and distributed execution strategies for large visual models.
  • Build rigorous benchmarking, profiling, and performance-regression infrastructure to understand performance and evaluate new optimization ideas.
  • Evaluate new inference runtimes, compilers, frameworks, optimization techniques, and accelerator hardware.
  • Stay close to advances in efficient inference, GPU programming, model architectures, compilers, and ML systems research, and rapidly test promising ideas.
  • Help turn research breakthroughs into models that can be deployed and served efficiently at scale.

What We’re Looking For

  • Deep technical ability in efficient ML inference, ML systems, GPU computing, or adjacent research, demonstrated through research, production engineering, open-source contributions, or unusually ambitious independent work.
  • Strong understanding of how modern deep learning models execute on hardware, including the relationship between compute, memory, communication, and performance.
  • Experience profiling ML workloads, identifying bottlenecks, forming hypotheses, and driving measurable performance improvements.
  • Strong programming fundamentals in Python, C++, or another systems-oriented language.
  • Experience with some combination of PyTorch, CUDA, Triton, TensorRT, vLLM, SGLang, or comparable technologies.
  • Ability to reason across abstraction layers rather than treating model architecture, framework, runtime, kernel, and hardware boundaries as fixed.
  • Strong intuition for performance tradeoffs across latency, throughput, memory, numerical precision, model quality, and complexity.
  • Curiosity about how models work internally and a willingness to modify or rethink existing approaches when the performance problem calls for it.
  • Comfort working on ambiguous problems where the bottleneck, and sometimes even the right question, is not known in advance.
  • Ability to communicate technical ideas clearly and collaborate closely with research scientists and engineers.

We don’t expect every candidate to have experience across the entire stack. Exceptional depth in one or more relevant areas, combined with the ability and curiosity to reason across the others, matters more to us than checking every box.

Nice to Have

  • Experience optimizing large generative, multimodal, vision, or video models.
  • CUDA, Triton, CUTLASS, or other GPU kernel development.
  • Deep knowledge of GPU architecture, memory hierarchy, and hardware-aware optimization.
  • Experience with attention optimization, kernel fusion, memory-efficient execution, or custom operators.
  • Model compilation or graph optimization experience.
  • Quantization, sparsity, caching, speculative execution, or other efficient inference techniques.
  • Experience optimizing diffusion, autoregressive, transformer, or other large generative architectures.
  • Multi-GPU or multi-node model execution, including tensor, pipeline, sequence, or other forms of parallelism.
  • Experience optimizing communication or data movement between accelerators.
  • Experience with profiling tools such as Nsight Systems or Nsight Compute.
  • Contributions to ML systems, inference runtimes, compilers, GPU libraries, or performance-focused open-source projects.
  • Research or publications in efficient ML, ML systems, GPU computing, compilers, or related areas.

Benefits:

  • Competitive compensation and equity
  • 401k
  • Healthcare (Silver PPO Medical, Vision, Dental)
  • Lunch and snacks at the office

This role is based in San Francisco, and we work together in person five days a week.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against a16z portfolio's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on a16z portfolio's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    a16z portfolio's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.