Skip to content

Open nowPosted 30 hours agoWe saw it 114 min after it went up

Principal AI Product Engineer

Nscale264 open roles

Pay
$290,000 – $520,000 a year
Where
Houston; New York; San Francisco; Seattle
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowPrincipal AI Product EngineerNscale · Houston; New York; San Francisco; Seattle
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Nscale's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.7% of postings close within 7 days. Measured by our own scanner across the market. Nscale postings stay open a median of 3 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.3%3 days
  3. 7.7%7 days
  4. 14.0%14 days
  5. 33.7%30 days
This job: posted 30 hours ago

Nscale median: 3 days open

The posting

About Nscale

Nscale is taking on the hyperscalers by building a vertically integrated GenAI cloud platform. We own the data centers, software, and applications that power today's AI stack using sustainable technology solutions. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As a Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. Collaboration is key, and we work together swiftly and respectfully, embracing adaptability and resilience in all we do.

About the Role

Nscale is looking for a Principal AI Engineer (Specialised) to lead the inference and post-training pillar of our AI systems engineering organization. You’ll define the multi-year technical roadmap for how models are served, evaluated, and post-trained on Nscale’s GPU cloud, across dedicated and serverless inference and bring-your-own-model deployments. You’ll lead the most consequential architectural programmes in that space and set the engineering standards that 20–50+ engineers build to.

As a Principal engineer, you are one of the deepest technical authorities in the company on AI systems. Your decisions set the cost, latency, and reliability at which Nscale serves tokens and runs post-training workloads, and those numbers have to compete with the world’s leading AI infrastructure providers. The problems span the full stack: kernel efficiency on state-of-the-art GPU systems, fleet-level KV cache and serving architecture, the evals that prove model quality, and RL loops where inference and training share hardware. You frame the solutions the organization executes against, including the API contracts customers see.

How We Work

  • Dog years. We move quickly and compress a lot of learning into a short time.
  • Don’t let perfect be the enemy of good. Ship, measure, iterate.
  • Be relentless. Own the problem end to end and see it through.
  • One team, one mission. Outcomes over process, and no “not my job”.

Responsibilities

  • Define and own the multi-year technical roadmap for Nscale’s inference, evals, and post-training platform, and translate it into architecture that multiple teams can execute against
  • Lead company-scale architectural initiatives in the pillar, such as next-generation serving (disaggregated prefill/decode, KV cache orchestration across GPU, host, and storage tiers, speculative decoding, multi-tenant scheduling), GPU kernel and model efficiency work (custom kernels, FP8/NVFP4/INT8/4 quantization, sparsity, distillation, MoE serving), evals and benchmarking frameworks, and post-training and RL infrastructure
  • Establish engineering standards adopted across all AI teams: API design and compatibility guarantees, benchmarking and evals methodology, training stability norms, and performance testing practices
  • Own the framework by which cost, latency, throughput, and model quality trade-offs are made and measured across the pillar
  • Identify long-horizon systemic risks early (serving engine and framework bets, accelerator support, capability gaps) and resolve them before they block the organization
  • Align AI engineering, research, product, and infrastructure leadership on multi-team technical strategy; frame technical trade-offs in product and commercial terms
  • Mentor and develop Staff and Senior AI Engineers, and grow the next generation of inference technical leaders at Nscale
  • Represent Nscale’s technical approach externally: open-source leadership in the frameworks we depend on, publications, conference talks, and partnerships with GPU vendors and AI labs

Requirements

  • 10–15 years of engineering experience, with a clear track record of pillar-level impact on production AI systems
  • 4+ years of hands-on work with LLMs in inference, GPU performance, evals, or post-training and RL, in production or research
  • Demonstrated ability to define multi-year technical strategy for complex, multi-team AI systems organizations
  • World-class depth in production LLM inference, GPU performance, evals, and/or post-training and RL infrastructure, with strong working knowledge across the rest
  • Demonstrated ownership of the architecture of a large-scale production inference or training platform
  • Proven ability to create architectural frameworks and engineering standards adopted across large engineering organizations
  • Deep understanding of the hardware/software boundary for AI accelerators: CUDA or ROCm, memory bandwidth and interconnect constraints, and distributed compute paradigms
  • Strong history of growing technical leaders (Staff and above) and multiplying technical capability across teams
  • External recognition in the AI systems community through research, open source, or industry contribution

Preferred

  • Prior experience at a top-tier AI lab or major hyperscaler AI infrastructure team
  • Maintainer or core contributor to a foundational inference, kernel, or RL framework (vLLM, SGLang, TensorRT-LLM, LMCache, FlashInfer, Triton, verl, OpenRLHF, TRL, DeepSpeed, Megatron-LM, etc.)
  • Hands-on depth in RL for LLMs (DPO/GRPO-style methods, reward modelling, multi-turn and tool-use RL) and the interaction between inference and training infrastructure
  • Experience defining developer API platforms adopted at scale by external developers
  • Deep experience with control plane / data plane architecture and cell-based deployment patterns in large-scale inference infrastructure
  • Published work in AI systems: MLSys, NeurIPS Systems Track, OSDI, EuroSys, SC, or equivalent
  • Experience with hardware-software co-design: custom accelerator kernels (CUDA, Triton), compiler-level optimization, AI hardware roadmap engagement
  • Experience defining pricing, SLO, and capacity models for a commercial inference product

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$290,000—$520,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Nscale's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Nscale's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Nscale's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.