Skip to content

Open nowPosted today

Director, Software Engineering – Cloud Compute and Infrastructure

NVIDIA2,522 open roles

Where
US, CA, Santa Clara
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowDirector, Software Engineering – Cloud Compute and InfrastructureNVIDIA · US, CA, Santa Clara
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on NVIDIA's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.3% of postings close within 7 days. Measured by our own scanner across the market. NVIDIA postings stay open a median of 28 days.

Share of postings closed within
  1. 1.9%1 day
  2. 4.0%3 days
  3. 8.3%7 days
  4. 15.3%14 days
  5. 34.2%30 days
This job: posted today

NVIDIA median: 28 days open

The posting

NVIDIA is seeking an engineering director to lead the software teams behind our bare-metal datacenters and cloud compute infrastructure. The organization supports hundreds of megawatts of datacenter capacity already online, with more capacity coming. You will build the software that makes this growing physical infrastructure available as reliable, secure, and efficient compute services for NVIDIA engineering.

Our team provides NVIDIA’s continuous integration (CI) infrastructure: the environment where hardware, firmware, drivers, networking, and system software are integrated and tested on their path to production. You will enable engineering teams to bring up preproduction systems, reproduce failures, validate changes, and move platforms toward production readiness. The fleet spans multiple hardware generations and maturity levels: x86 and Arm servers, GPUs, DPUs, high-speed networking, storage, and rack-scale, liquid-cooled systems. Platforms such as Grace Blackwell and Vera Rubin illustrate the breadth of compute and interconnect technology involved. This role combines hands-on systems judgment with leadership of the teams making that diversity manageable at scale.

What you’ll be doing:

  • Lead software engineering teams responsible for cloud compute, bare-metal fleet management, and CI infrastructure, owning architecture, implementation, deployment, and operations.
  • Set the technical direction for compute control planes, resource provisioning, topology-aware placement, reservations, and capacity management across physical servers, virtual machines, and containers.
  • Automate the bare-metal lifecycle: hardware discovery and inventory, server bring-up, imaging, firmware and driver configuration, out-of-band management, health checks, reprovisioning, and recovery.
  • Build CI services that allocate the right hardware, provision repeatable test environments, record hardware and software configurations, collect diagnostics, and restore systems to a known state. Shorten the path from a hardware or software change to actionable validation results.
  • Keep engineering services dependable while supporting evolving preproduction hardware and software. Define service-level objectives, isolate failures, improve observability, and turn incidents and recurring test-infrastructure failures into engineering fixes.
  • Measure and improve hardware onboarding time, provisioning speed, CI queue time, productive fleet utilization, recovery time, and cost efficiency as capacity expands.
  • Set engineering standards for design and code reviews, automated testing, secure development, release quality, and safe changes to production systems.
  • Hire, coach, and retain engineers and engineering managers. Establish clear ownership, develop technical leaders, and build teams that deliver consistently over multiple release cycles.
  • Partner with hardware, firmware, drivers, networking, storage, security, and validation teams to onboard platforms and diagnose failures across system boundaries. Translate engineering users’ needs into clear priorities and technical decisions.
  • Work with datacenter operations and facilities teams to bring additional capacity online. Connect rack power, cooling, physical topology, and hardware-health telemetry to provisioning, placement, serviceability, and operational readiness.

What we need to see:

  • 15+ overall years of experience in software engineering, distributed systems, or cloud infrastructure, including 7+ years leading engineering teams and experience managing engineering managers.
  • Direct engineering ownership of the underlying services of a public or private compute cloud, such as AWS EC2, Google Compute Engine, Azure Compute, OCI Compute, or a comparable infrastructure-as-a-service platform.
  • Strong technical depth in distributed systems, Linux, virtualization, containers, and the networking and storage services that support large compute fleets.
  • Experience building software and automation for bare-metal infrastructure, including server provisioning, hardware inventory, firmware or operating-system lifecycle management, and recovery across heterogeneous systems.
  • Experience running production infrastructure with demanding availability requirements, including failure isolation, incident response, observability, and reliable deployment and recovery mechanisms.
  • The ability to guide architecture, evaluate implementation choices, and debug complex interactions among hardware, firmware, drivers, operating systems, and distributed services with senior engineers.
  • A record of sustained ownership as platforms evolve, with measurable improvements in reliability, scalability, or engineering delivery. Strong communication, collaboration, and people-development skills.
  • A degree in computer science, computer engineering, or a related discipline, or equivalent experience.

Ways to stand out from the crowd:

  • Experience building or operating OpenStack infrastructure, particularly Nova, Neutron, Cinder, or Ironic, or contributing to related open-source projects.
  • Engineering leadership spanning compute control planes and site reliability, including multi-tenant isolation, scheduling, placement, and capacity allocation.
  • Experience with GPU clusters, rack-scale computing, NVLink, InfiniBand or high-speed Ethernet, and AI or high-performance computing workloads.
  • Experience supporting preproduction hardware, new product introduction, or hardware-in-the-loop CI, including repeatable validation environments and systematic regression isolation.
  • Familiarity with liquid-cooled, high-density datacenters and how power, thermal constraints, cooling, and component health affect fleet availability and scheduling and delivery of compute platforms across multiple regions, datacenters, or hybrid-cloud environments, including fleet upgrades and workload migration with minimal service disruption.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 13, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against NVIDIA's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on NVIDIA's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    NVIDIA's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.