Skip to content

Open nowPosted 22 days ago

Senior Site Reliability & Software Engineering Manager

Crux, Inc.29 open roles

Where
Palo Alto
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSenior Site Reliability & Software Engineering ManagerCrux, Inc. · Palo Alto
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Crux, Inc.'s own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.1% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.7%1 day
  2. 3.6%3 days
  3. 8.1%7 days
  4. 15.1%14 days
  5. 34.0%30 days
This job: posted 22 days ago

The posting

Built to set the gold standard for integrated AI infrastructure

Crux AI is a newly formed, U.S.-based integrated AI infrastructure company created to remove the physical and operational constraints on consequential AI ambitions. Crux brings together power, high-density data centers, TPU silicon, networking, orchestration software, and ongoing operations as one integrated system.

Crux is being capitalized to plan every layer together, develop each one to demanding standards, and operate the whole system with efficiency and reliability. That gives hyperscalers, frontier AI labs, sovereign customers, enterprises, and AI-native companies greater freedom to pursue the AI they are here to create.

Crux AI is led by CEO, Ben Treynor Sloss, who spent over two decades in executive technical leadership at Google and founded the Site Reliability Engineering (SRE) discipline. At Crux AI, we treat operations fundamentally as a software engineering problem.

WHAT YOU'LL DO

We are recruiting founding Senior Site Reliability and Software Engineering Managers to build and lead our initial fleet reliability engineering teams in Palo Alto, CA.

In this organization, there is no separate software development team. Your team owns the software, control plane, telemetry, and automated remediation controllers that keep multi-gigawatt TPU clusters provisioned, resilient, and continuously executing customer AI workloads.

This is a true hands-on, builder seat, not a supervisory position. In the early days, you will write the first remediation controllers, set reliability baselines, and take initial pages yourself to stay close to the work before expanding your team. You will lead an elite group of unusually senior software and reliability engineers: engineers with significantly greater technical and software depth than traditional operational SRE orgs. You must thrive on independence and revel in ambiguity, turning unknowns into concrete engineering priorities in a fast-paced, high-growth environment. While you will devise and participate in initial on-call rotations, your core mandate is to combine SRE disciplines with extensive AI/ML automation to drive operational pages down to zero.

In this role, you will:

- Build & Lead Senior Engineering Teams: Recruit, lead, and mentor an initial team of senior software and reliability engineers across Palo Alto and Europe as fleet capacity ramps rapidly.

- Automate Pages to Zero: Own fleet availability end-to-end; establish on-call rotations while relentlessly developing self-healing systems and predictive remediation to eliminate manual pages.

- Embed AI/ML into SRE Disciplines: Apply agentic techniques, machine learning models, and automated diagnostic workflows extensively to telemetry collection, root-cause analysis, and predictive cluster recovery.

- Own Bare-Metal & Fleet Lifecycle Software: Drive software engineering for bare-metal node provisioning, firmware deployment, thermal/stress burn-in validation, host/TPU health monitoring, and decommissioning.

- Control Plane & Fabric Reliability: Own software reliability for cluster orchestration, scheduling, capacity allocation APIs, and high-performance TPU host/interconnect networks.

- Define Observability & SLOs: Establish customer-facing SLIs/SLOs (job goodput, time-to-detect, node availability) and build the telemetry pipelines serving operators, executives, and customers.

SIGNALS OF SUCCESS

After 60 days in this role:

- Baseline reliability framework defined (v1 SLOs, severity structure, change management); first AI/ML-driven automated remediation controller shipped to production; recruiting active for senior engineering hires in Palo Alto.

After 6 months:

- First TPU cluster brought online under your team’s automated acceptance criteria; automated remediation pipeline running with pass rates tracked; observability v1 in daily production use; core senior team onboarded

After 1 year:

- TPU fleet operating against published customer SLOs; >90% of node/fabric faults automatically quarantined and remediated without human paging; team scaled ahead of rapid capacity ramps.

EXPERIENCES, ATTRIBUTES AND MINDSET THAT INDICATE A GOOD MATCH

Experiences

- 10+ years of software or infrastructure engineering experience, with 3+ years managing engineering teams owning direct production SLAs and on-call.

- Deep SRE Discipline: Grounded in foundational SRE principles (SLOs, error budgets, blameless postmortems) paired with a strict "code over heroics" mindset.

- Hands-On Technical Depth (SRE + SWE): Track record shipping production code in Go, Python, or C++, with hands-on systems expertise across Linux OS kernels, bare-metal provisioning, firmware, and/or high-performance networking fabrics. Extensive experience with distributed systems and open source software.

- Extensive AI/ML Adoption: Active utilization of AI agents and automated LLM/ML workflows in modern software engineering and diagnostic operations.

- Senior Talent Magnet: Track record of attracting, evaluating, developing and leading unusually senior software engineers who thrive in fast-paced, high-stakes environments.

Attributes

- Possess a high tolerance for ambiguity. The first clusters will carry customer workloads while the SLOs are still being defined and the team is still being hired. You absorb that, translate unknowns into concrete near-term priorities, and never manufacture false certainty about reliability the data does not support.

- Understands that the customer’s job is the unit of reliability. A node that is “up” while a training run stalls on a flapping link is down. You measure what customers experience — goodput, time-to-recover, lost progress — and hold the whole stack, and Google, to it.

Mindset

- Builder Mindset & Ambiguity: A true "builder, not supervisory" orientation; comfortable operating with high autonomy, navigating ambiguity, and establishing structure amidst rapid growth.

- AI-Agentic First: Fluent with AI agents — or committed to becoming so quickly — and you embed them as first principles in how you and your team work, defaulting to agentic workflows before adding headcount or process.

Nice to have (Preferred, not required):

- Hyperscaler / Neocloud Scale: SRE or fleet leadership at a hyperscaler (Google, AWS, Meta, MSFT) or neocloud (CoreWeave, Lambda, Nebius, Nscale) during rapid fleet ramps.

- Accelerated Compute: Direct TPU experience or large-scale GPU cluster ops (NCCL collective debugging, RDMA/GPU-Direct, Slurm/Kubernetes AI schedulers).

- Custom Fleet Tooling: Hands-on experience building custom remediation controllers, event-driven fleet management software (Go, NetBox/DCIM), or OpenTelemetry/Prometheus pipelines.

- Facility Telemetry & Thermal Signals: Familiarity with high-density, liquid-cooled environments and integrating facility telemetry (power, thermal, flow) into compute health signals.

- Customer SLAs & Reporting: Proven experience constructing customer SLAs/SLOs, credit mechanics, and executive/customer-facing reliability reviews.

Salary Range Information

The annual salary range for this position has been estimated based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Crux

- We offer generous base, bonus and additional incentive based compensation

- Health, dental, and vision coverage for you and your dependents

- Company-paid life insurance and disability

- Full suite of other optional benefits

- 401(k) Plan with 4% company match

- Hybrid schedule offering four days in-office collaboration paired with one remote workday for focused, individual work

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Crux, Inc.'s own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Crux, Inc.'s form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Crux, Inc.'s answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.