Skip to content

Open nowFirst seen 2 hours ago

Software Engineer, AI Kernels & Performance Optimization — MTIA Software

Meta1,071 open roles

Where
Menlo Park, CA
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSoftware Engineer, AI Kernels & Performance Optimization — MTIA SoftwareMeta · Menlo Park, CA
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Meta's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market. Meta postings stay open a median of 35 days.

Share of postings closed within
  1. 1.7%1 day
  2. 3.6%3 days
  3. 8.1%7 days
  4. 15.0%14 days
  5. 33.9%30 days
This job: first seen 2 hours ago

Meta median: 35 days open

The posting

Meta designs and deploys its own AI systems. MTIA — the Meta Training and Inference Accelerator — is Meta's family of in-house AI accelerator ASICs, running recommendation and ranking workloads in production across Meta's data centers today and expanding into generative AI inference and training as successive silicon generations land (see [https://bit.ly/metamtia](https://bit.ly/metamtia)).

The MTIA Software team is part of the **AI & Compute Foundation (ACF)** organization within Meta Infrastructure. Because the hardware is ours, the software is ours too: we build the entire stack a chip vendor would normally supply — compiler and LLVM toolchain, runtime, kernel authoring frameworks and libraries, developer tooling, and deep PyTorch integration — and we co-design it with the silicon teams generation over generation.

Within that stack, the AI kernel and optimization software development team drives the layer where architecture meets arithmetic. Our mission is performance *and* programmability at scale: hit roofline enablements on the workloads that matter, and make kernel authoring accessible enough that the whole organization can close coverage gaps without funneling every problem through a handful of experts. We do this by shipping high-performance kernel libraries with broad PyTorch operator coverage, by building the C++ and Python kernel authoring frameworks and DSL surfaces that others build on, and by writing production kernels against new architectures long before first silicon — turning hardware proposals into measured roofline evidence while the design can still change.

We are hiring an experienced kernel and performance engineer to take on this role at a senior level. You will own the performance of workloads that serve billions of people, from the innermost loop of a fused attention kernel to the numeric decisions that determine whether a model converges. You will read hardware specifications and RTL-adjacent documentation as easily as you read code, and you will be expected to say clearly when the hardware — not the software — is the problem. Your findings will change what gets built next.

This is a hands-on engineering role with wide latitude. The problems aren't incremental.

## What you'll work on

- **Roofline-level kernels.** GEMM and attention variants, normalization, collectives, elementwise and reduction fusions, sparse and quantized paths — implemented against novel architectural features (matrix engines, on-chip reduction fabrics, software-managed memory hierarchies) and tuned until the remaining gap to the machine's limit is explainable in a sentence. - **Numerics under precision constraints.** Low-precision formats (FP8, MX-style block-scaled types, integer quantization) where the difference between a correct scale choice and a plausible one is several decibels of signal, and where the fix has to work on silicon that has already been taped out. - **Kernel authoring frameworks.** Templateized, composable C++ kernel SDKs in the spirit of CUTLASS, Python DSLs in the spirit of Triton and CuTe, and the compiler-facing interfaces that let automated codegen reach performance that used to require a specialist. - **Pre-silicon and bring-up.** Kernels on simulators and emulators, validating architectural features and rooflines before tapeout, then first-light bring-up on real parts. - **Software mitigations for hardware realities.** Every chip ships with something you wish were different. Finding the workaround that recovers most of the lost performance — and generalizing it so nobody rediscovers it — is core to the job.

Responsibilities

  • Design, implement, and optimize high-performance compute and communication kernels for MTIA accelerators, taking ownership from architectural analysis through production deployment
  • Profile and root-cause performance across the full stack — instruction scheduling, memory hierarchy and DMA behavior, on-chip interconnect, multi-device collectives — and drive the fixes to the right layer, whether that is the kernel, the compiler, the runtime, or the hardware
  • Build and extend kernel authoring frameworks, templates, and libraries so that other engineers can reach high performance without deep architectural expertise; raise the ceiling and lower the floor at the same time
  • Deliver and maintain broad kernel coverage for PyTorch operators across recommendation, ranking, and generative AI workloads, in both eager and compiled execution paths
  • Partner with silicon architecture and design teams on hardware/software co-design: quantify the value of proposed features with real kernels, characterize rooflines pre-silicon, and advocate for the changes the software stack actually needs
  • Work with compiler, runtime, framework, and product-facing teams to land end-to-end wins on production models rather than isolated microbenchmark improvements
  • Investigate numerics and precision trade-offs, and design software mitigations that recover performance or accuracy lost to hardware limitations
  • Set technical direction for a kernel domain, write the design documents that align cross-functional partners, and mentor engineers on performance methodology and accelerator programming

Minimum Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • Bachelor's degree in Computer Science, Computer Engineering, a related technical field, or equivalent practical experience
  • 6+ years of professional experience in high-performance computing, accelerator kernel development, compiler backends, or systems performance engineering
  • Proficiency in C++ and Python, including low-level systems programming, templates and generic programming, and comfort reading and writing performance-critical code
  • Demonstrated experience writing and optimizing kernels for a parallel architecture — GPU (CUDA, ROCm/HIP, SYCL/OpenCL), TPU or other AI ASICs, or SIMD/vector CPU targets
  • Working knowledge of computer architecture as it applies to performance: memory hierarchies and bandwidth, latency hiding, occupancy and scheduling, vectorization, and synchronization
  • A rigorous, measurement-driven approach to performance: the ability to build a roofline or analytical model, profile against it, and explain the residual gap

Preferred Qualifications

  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Deep familiarity with transformer and attention kernel design: FlashAttention-class algorithms, KV-cache management, paged and chunked attention, linear and state-space attention variants, MoE routing and expert dispatch
  • 8+ years of experience in accelerator software, HPC, or ML systems performance (or equivalent with an advanced degree)
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Experience with compiler and codegen technologies relevant to kernels: MLIR, LLVM, TVM, XLA, Halide, or polyhedral scheduling
  • Experience with low-precision numerics and quantization — FP8/E4M3/E5M2, MX and other block-scaled formats, INT8/INT4 — including error analysis and calibration
  • Experience mentoring engineers and setting technical direction across teams
  • Track record of open-source contribution in the kernel, compiler, or ML systems ecosystem
  • Experience with distributed execution and collective communication (NCCL/RCCL-class primitives, tensor and expert parallelism, overlapping communication with compute)
  • Experience building or contributing to high-performance kernel libraries or frameworks — CUTLASS, cuBLAS, cuDNN, CUTE, Triton, Helion, ThunderKittens, oneDNN, Composable Kernel, or comparable internal equivalents
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Experience with pre-silicon software development — architectural simulators, FPGA emulation, performance modeling — and with hardware/software co-design cycles
  • Familiarity with ML framework internals: PyTorch dispatch and eager execution, torch.compile / Inductor, custom operator integration, and inference serving stacks such as vLLM or SGLang

US: $154,003/year to $217,000/year + bonus + equity

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Meta's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Meta's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Meta's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.