Skip to content

Open nowPosted 37 days ago

Founding Senior Infra Engineer, Distributed Infra

goaly7 open roles

Where
Palo Alto, CA, USA
Work mode
Hybrid
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowFounding Senior Infra Engineer, Distributed Infragoaly · Palo Alto, CA, USA
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on goaly's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.7%1 day
  2. 3.5%3 days
  3. 7.8%7 days
  4. 14.6%14 days
  5. 34.1%30 days
This job: posted 37 days ago

The posting

About us

We’re building toward a world where every company can become its own AI lab. Goaly is a stealth AI startup founded by ex-Meta Superintelligence Labs engineers and researchers. Our mission is to dramatically lower the cost, time, and talent barriers to building proprietary AI — and make each generation of models faster and cheaper to build than the last.

Backed by leading AI investors and endorsed by frontier AI researchers and builders, we’re looking for exceptional new grads who want to work on hard, foundational AI systems problems with outsized ownership from day one.

About the role

You will own systems across the lifecycle of our accelerator clusters—from bringing capacity online and upgrading fleets to detecting failures, recovering safely, and retiring capacity. Your work will determine how quickly researchers can start experiments, how efficiently expensive hardware is used, and how reliably long-running workloads complete.

This is a hands-on infrastructure role spanning cloud and datacenter environments, cluster control planes, networking, storage, security, observability, and automation. You will partner with hardware and cloud providers and our Training, RL Systems, Post-Training, Inference, and Security teams to turn heterogeneous compute into a dependable platform. Depending on experience, you may lead multi-quarter initiatives and help set technical direction.

What you'll do

- Design, build, and operate control-plane services and infrastructure-as-code for provisioning, configuration, validation, upgrades, expansion, draining, recovery, and decommissioning; make every change repeatable, auditable, and safe to roll back.

- Bring new accelerator capacity online on schedule by coordinating dependencies across cloud providers, datacenter and hardware partners, networking, storage, security, and internal compute consumers.

- Build high-bandwidth, topology-aware connectivity within and across clusters; diagnose performance and reliability issues spanning hosts, switches, routing, transport, collective communication, and workload placement.

- Make clusters secure by default through identity and access controls, network policy, workload isolation, host and container hardening, secrets management, and trusted software and image supply chains.

- Improve fleet scalability, consistency, and fault tolerance by defining health signals, automating remediation, reducing configuration drift, and designing for partial failure.

- Establish service-level objectives and observability for cluster readiness, provisioning time, usable capacity, job-start latency, infrastructure-caused failures, utilization, and recovery time.

- Lead incident response and blameless postmortems for cluster failures; turn recurring operational pain into automation, safer defaults, and simpler system boundaries.

- Work directly with Training, RL Systems, Post-Training, and Inference engineers to debug cross-layer failures and shape a long-term compute, data, networking, and capacity roadmap.

You may be a good fit if you have

- Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure).

- Strong programming ability in Python, Go, Rust, or another language suited to reliable infrastructure services and automation.

- Hands-on experience with Linux, containers, Kubernetes or another cluster scheduler, infrastructure-as-code, and at least one major cloud platform or substantial bare-metal environment.

- A practical understanding of networking, storage, identity, observability, and reliability, with the ability to trace a failure across multiple layers of a complex system.

- Experience designing systems for safe rollout, fault isolation, idempotency, capacity growth, and recovery from partial or large-scale failures.

- High ownership and clear communication, including comfort coordinating multi-team projects and participating in a healthy on-call rotation.

Strong pluses

- Experience operating large GPU or accelerator fleets for distributed model training, inference, scientific computing, or another communication-intensive workload.

- Depth in Kubernetes internals, custom controllers or operators, device plugins, cluster autoscaling, scheduler extensions, Slurm, or comparable orchestration systems.

- Experience with high-performance networking such as RDMA, InfiniBand, RoCE, BGP, cloud interconnects, multi-NIC hosts, CNI or eBPF networking, or topology-aware placement.

- Experience with Terraform, workflow orchestration, and automated qualification of hosts, drivers, firmware, networks, and new hardware.

- Knowledge of GPU systems, NCCL, NVLink or NVSwitch, and the failure modes of large distributed jobs.

How we work

- Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.

- High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.

- Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.

- Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.

- Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.

- Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.

Location, visa sponsorship & benefits

- Location-based hybrid policy. This is a location-based hybrid role. We currently expect all staff to work from one of our offices at least three days per week. Exact office options will be confirmed during the recruiting process.

- Visa sponsorship. We do sponsor visas. However, we cannot successfully sponsor a visa for every role and every candidate. If we make you an offer, we will make every reasonable effort to secure the necessary visa, and we retain immigration counsel to support the process.

- Meals and office benefits. We provide complimentary lunch and dinner in our offices, along with snacks and beverages.

A note on qualifications. We care more about exceptional evidence than a perfect keyword match. If the work excites you and you can show unusual strength, learning speed, or ownership, we encourage you to apply even if your background does not match every preferred qualification.

Equal opportunity

We are an equal opportunity employer and consider qualified applicants without regard to any characteristic protected by applicable law. Reasonable accommodations are available throughout the hiring process.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against goaly's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on goaly's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    goaly's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.