Skip to content

Open nowPosted 36 days ago

Founding Senior AI Infrastructure Engineer

goaly7 open roles

Where
Palo Alto, CA, USA
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowFounding Senior AI Infrastructure Engineergoaly · Palo Alto, CA, USA
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on goaly's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.8% of postings close within 7 days. Measured by our own scanner across the market.

Share of postings closed within
  1. 1.7%1 day
  2. 3.5%3 days
  3. 7.8%7 days
  4. 14.6%14 days
  5. 34.1%30 days
This job: posted 36 days ago

The posting

ABOUT US

We’re building toward a world where every company can become its own AI lab. Goaly is a stealth AI startup founded by ex-Meta MSL engineers and researchers. Our mission is to dramatically lower the cost, time, and talent barriers to building proprietary AI — and make each generation of models faster and cheaper to build than the last.

Backed by leading AI investors and endorsed by frontier AI researchers and builders, we’re looking for exceptional new grads who want to work on hard, foundational AI systems problems with outsized ownership from day one.

ABOUT THE ROLE

Running modern AI workloads at scale creates systems problems that rarely fit within a single layer of the stack. A slowdown that appears in a training job may originate in a GPU kernel, collective communication, data movement, container runtime, storage path, scheduler, or interaction between model architecture and hardware topology.

You will identify these bottlenecks and build systems that improve the throughput, efficiency, and robustness of our largest distributed workloads. Your scope will span post-training, agentic reinforcement learning, model training, rollout inference, and the GPU cluster platform beneath them. You will work closely with researchers and systems engineers, develop a quantitative understanding of performance, and turn one-off investigations into durable infrastructure improvements.

This role is a strong fit for an exceptional systems or performance engineer who wants to work on frontier AI infrastructure. Deep ML experience is valuable but not required; we care most about a track record of solving difficult systems problems at scale and the ability to become fluent in new parts of the ML stack quickly.

WHAT YOU'LL DO

- Profile end-to-end AI workloads and identify limiting resources across model code, GPU kernels, memory, collective communication, networking, storage, orchestration, and environment execution.

- Build low-latency, high-throughput sampling and inference systems for large language models, including batching, scheduling, caching, load balancing, and efficient weight updates.

- Optimize GPU execution through kernel and graph profiling, memory-layout improvements, reduced-precision computation, communication overlap, compilation, and targeted CUDA or Triton work.

- Improve distributed training and reinforcement-learning performance across heterogeneous GPU and CPU workloads, variable-length rollouts, complex network topologies, and changing model architectures.

- Design quantitative performance and capacity models that predict bottlenecks, explain scaling behavior, guide hardware and topology choices, and prioritize engineering work.

- Build scheduling and load-balancing mechanisms that improve accelerator utilization while respecting memory, locality, topology, latency, and fault-domain constraints.

- Design fault-tolerant distributed systems that detect failures early, isolate their impact, recover efficiently, and preserve correctness during long-running jobs.

- Investigate difficult production issues such as kernel-level stalls, network-latency spikes, collective timeouts, memory fragmentation, stragglers, and performance regressions in containerized environments.

- Develop benchmarks, profiling tools, performance dashboards, and regression tests that make system behavior visible and allow improvements to be measured under realistic workloads.

- Partner with researchers to understand new models and algorithms, remove infrastructure constraints from the experimental loop, and translate successful optimizations into reusable platform capabilities.

WHAT SUCCESS LOOKS LIKE

- Within your first three months, you have built a quantitative understanding of at least one critical workload, identified its dominant bottlenecks, and shipped a measurable performance or reliability improvement.

- Within six to twelve months, you have delivered sustained gains in throughput, accelerator utilization, latency, scaling efficiency, or cost across real training, rollout, or inference workloads.

- Performance investigations become faster and more rigorous because the team has better benchmarks, models, profiles, and observability—not just undocumented fixes.

- Large distributed jobs run predictably across complex hardware and network topologies, and failures or regressions can be detected, explained, and recovered from with minimal researcher intervention.

YOU MAY BE A GOOD FIT IF YOU HAVE

- Significant software-engineering, distributed-systems, high-performance-computing, or ML-infrastructure experience, particularly with performance-critical systems operating at large scale.

- Exceptional programming and debugging ability in Python and at least one systems language such as C++, Rust, or Go.

- Strong systems fundamentals, including operating systems, concurrency, memory, networking, storage, scheduling, containerization, and failure handling.

- A track record of using measurement to solve ambiguous performance problems: forming hypotheses, designing representative benchmarks, reading profiles and traces, identifying root causes, and validating improvements under production conditions.

- The ability to reason across abstraction boundaries, from model architecture and framework execution to accelerator behavior, distributed runtimes, cluster topology, and infrastructure services.

- A results-oriented mindset, flexibility about where in the stack to work, and a willingness to take ownership beyond a narrowly defined job description.

- Strong communication and collaboration skills, including the ability to work directly with researchers, explain complex systems behavior clearly, and turn repeated investigations into maintainable tools and abstractions.

- Interest in developing deep expertise in machine learning systems, even if your prior work has been primarily in distributed systems, HPC, operating systems, compilers, databases, or networking.

STRONG PLUSES

- Experience building or operating high-performance, large-scale ML training, post-training, or inference systems.

- Experience with GPU or accelerator programming, including CUDA, Triton, custom kernels, low-precision computation, memory optimization, or compiler stacks.

- Familiarity with ML framework internals and distributed stacks such as PyTorch, JAX, torch.distributed, FSDP, Megatron, DeepSpeed, Ray, vLLM, SGLang, or TensorRT-LLM.

- Knowledge of NCCL or RCCL, RDMA, InfiniBand, RoCE, NVLink, collective algorithms, network topology, or topology-aware workload placement.

- Experience with Linux or operating-system internals, container runtimes, Kubernetes, cluster schedulers, storage systems, or production observability.

- Familiarity with transformer architectures, language-model training, reinforcement learning, online sampling, or the performance characteristics of mixture-of-experts models.

- Meaningful contributions to open-source systems, ML frameworks, compilers, kernels, networking software, or performance tooling.

HOW WE WORK

- Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.

- High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.

- Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.

- Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.

- Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.

- Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.

A note on qualifications. We care more about exceptional evidence than a perfect keyword match. If the work excites you and you can show unusual strength, learning speed, or ownership, we encourage you to apply even if your background does not match every preferred qualification.

EQUAL OPPORTUNITY

We are an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. We provide reasonable accommodations for candidates who need them during the hiring process.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against goaly's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on goaly's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    goaly's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.