Skip to content

Open nowPosted today

Software Development Engineer, ML Acceleration, Trainium AI Systems, Annapurna Labs

Amazon / AWS22,907 open roles

Where
Austin, Texas, United States
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowSoftware Development Engineer, ML Acceleration, Trainium AI Systems, Annapurna LabsAmazon / AWS · Austin, Texas, United States
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Amazon / AWS's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.0% of postings close within 7 days. Measured by our own scanner across the market. Amazon / AWS postings stay open a median of 7 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 8.0%7 days
  4. 15.0%14 days
  5. 34.2%30 days
This job: posted today

Amazon / AWS median: 7 days open

The posting

Annapurna Labs designs the silicon behind AWS machine learning acceleration. Trainium and Inferentia servers train and serve the largest models our customers run, and before any of that hardware carries customer traffic it has to prove it works. MLA Vetting owns that gate. We build the diagnostic tests that every accelerator server runs before it becomes sellable, and we decide which failures block a host, which route to a technician for repair, and which are noise.

That gate only works if we can see through it. We are hiring a Software Development Engineer II to build the data and analytics systems that tell us what our tests are actually doing across the fleet. You will build the pipelines that ingest diagnostic results from tens of thousands of servers, the dashboards that make test behavior legible, and the metrics and alarms that tell us a test has started failing on a hardware generation before it costs us a week of capacity. When the team decides whether a diagnostic has enough evidence to move from observation into blocking production, your data answers that question.

This is a software engineering role, not a reporting role. You will write and operate production code, own the alarms and metrics that page us, and build detection for trends nobody is watching yet. You will work directly with hardware, firmware, provisioning, and data center operations teams who act on what you build.

You are fluent in Python and SQL and have built production data pipelines at scale. Experience with Amazon Redshift, Apache Spark, workflow orchestration, and dashboarding in Grafana or QuickSight maps directly to what you will work on here. Experience with telemetry or diagnostic data from a large server fleet, or with defining metrics and anomaly detection on operational time-series data, will have you productive faster. Familiarity with server hardware or data center operations is a plus, not a requirement — we will teach you the hardware.

Key job responsibilities Design, build, and operate production data pipelines that ingest diagnostic, telemetry, and repair-ticket data from the Trainium and Inferentia fleet into a warehouse other teams query with confidence.

Build and own the metrics, alarms, and anomaly detection that surface test regressions, failure-rate shifts, and new failure signatures across hardware generations without a human going looking for them.

Build dashboards and visualizations that make fleet and test health legible to engineers, hardware partners, and leadership, covering failure rates, failure-signature breakdowns, first pass yield, and repair latency.

Analyze large-scale fleet data to find root cause behind failure trends, and separate genuine hardware faults from software defects and test noise — a distinction that decides whether a failure reaches a technician or an engineer.

Define the evidence standard that gates operational decisions, including whether a diagnostic has soaked long enough and cleanly enough in the fleet to move from observation mode into blocking production.

Improve data quality and pipeline reliability so downstream consumers trust the numbers without re-deriving them.

Write clear analyses and design documents for technical and non-technical readers, including leadership.

A day in the life You start by checking an alarm that fired overnight: a diagnostic's failure rate climbed on one server generation. You query the failure signatures, find the increase concentrates in one component and one data center, and hand the breakdown to the hardware team by mid-morning. The rest of the day goes to a pipeline you are building to join repair outcomes against diagnostic results, so the team can measure how often a repair actually fixes the fault it was dispatched for. You close the day reviewing a teammate's code review and answering an operations partner who needs a new cut of yield data.

About the team MLA Vetting sits between hardware manufacturing and customer-ready capacity. We own the diagnostic test suites and the triage framework that turn a raw hardware failure into a specific, actionable repair, and we own the fleet-scale evidence that drives hardware and firmware improvements upstream. Our work is measured in how fast a server reaches sellable and how often it gets there on the first try.

- 3+ years of non-internship professional software development experience - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience - Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field - Experience programming with at least one software programming language

- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience - Bachelor's degree in computer science or equivalent

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, TX, Austin - 143,700.00 - 194,400.00 USD annually

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Amazon / AWS's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Amazon / AWS's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Amazon / AWS's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.