Skip to content

Open nowPosted 6 hours agoWe saw it 91 min after it went up

Hardware Diagnostics Engineer - Infrastructure

TensorWave38 open roles

Where
Las Vegas, Nevada
Work mode
On site
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowHardware Diagnostics Engineer - InfrastructureTensorWave · Las Vegas, Nevada
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on TensorWave's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

7.7% of postings close within 7 days. Measured by our own scanner across the market. TensorWave postings stay open a median of 29 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.3%3 days
  3. 7.7%7 days
  4. 14.0%14 days
  5. 33.7%30 days
This job: posted 6 hours ago

TensorWave median: 29 days open

The posting

About TensorWave

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

About the Role

We are looking for a Hardware Diagnostics Engineer to run burn-in, triage what fails, work servers out-of-band, and own RMAs end to end. If you like hardware that misbehaves in ways that take real work to explain, this is a good seat.

Before any GPU server carries a customer workload, it has to prove it works — under load, at temperature, for hours. When it doesn't, somebody has to figure out why, get replacement hardware in, and send the failed part back to the vendor.

What You’ll Do

- Run server and GPU burn-in and stress testing, interpret the results, and decide whether hardware is production-ready

- Triage failures across GPUs, memory, drives, NICs, PSUs, and cabling: reproduce the failure, isolate the faulty component, and document what proved it

- Work servers out-of-band through IPMI and Redfish for power control, boot configuration, BIOS settings, and sensor and event log collection

- Apply firmware updates across the fleet following the team's qualified baselines and rollout process

- Drive RMAs with vendors from ticket through replacement, installation, and return of the failed part

- Keep asset, serial, and replacement history accurate in NetBox so we know what's actually in every rack

- Track failure patterns across the fleet and raise them when the same part or firmware version keeps turning up

- Improve the runbooks you work from, and script the steps you find yourself repeating

- Partner with datacenter operations on hands-on work during turn-ups and expansions

- Take part in an on-call and escalation rotation for hardware issues

Who You Are

Required Qualifications

- 3–6 years in datacenter operations, systems administration, hardware support, or infrastructure engineering

- Hands-on experience with enterprise server hardware: component replacement, POST and boot failures, and reading hardware behavior at the rack

- Practical experience with BMCs and out-of-band management: IPMI, Redfish, iDRAC, iLO, or equivalent

- Strong Linux troubleshooting: boot process, driver and device issues, and diagnostic tools such as {{dmesg}}, {{lspci}}, {{ipmitool}}, and SMART

- Comfort reading sensor data, event logs, and thermal and power telemetry well enough to tell a real failure from noise

- Working scripting ability in Bash or Python — enough to automate a repetitive task and read someone else's tooling

- Experience running hardware RMAs with vendors, or a clear track record of driving issues to closure with outside parties

- A methodical troubleshooting habit: you isolate variables, you don't change three things at once, and you can say what evidence led to your conclusion

- Clear written communication for tickets, runbooks, and vendor cases

Preferred Qualifications

- GPU server experience, especially AMD GPUs and ROCm

- Burn-in, stress testing, or node validation tooling in a GPU or HPC environment

- Familiarity with firmware update processes and why fleet-wide changes get staged

- NetBox or other DCIM and IPAM tooling

- Ansible, or Python against REST APIs

- Prior work in a high-volume hardware environment: hyperscaler, colo, integrator, or manufacturing test

First Six Months

- By 90 days you'll run burn-in cycles and triage failures independently from our runbooks, and you'll have driven at least one RMA to closure. By six months you're the person who spots the pattern before anyone else — this batch, this firmware, this part — and you've automated at least one step you used to do by hand.

What We Offer

- Stock Options

- 100% paid Medical, Dental, and Vision insurance for Employees

- Company Health Savings Account Contributions

- 100% paid Short Term and Long Term Disability Insurance for Employees

- Life and Voluntary Supplemental Insurance Options

- Other Insurance Options, such as Pet & Legal Insurance

- Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support

- Flexible Spending Account

- 401(k)

- Employee Assistance Program

- Flexible PTO

- Paid Holidays

- Parental Leave

- Other In-Office Perks

Equal Employment Opportunity

TensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.

Reasonable Accommodations

TensorWave provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process, please contact [email protected].

Employment Eligibility

All offers of employment are contingent upon verification of identity and authorization to work in United States, as required by law.

Background Checks

Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.

Data Privacy Notice

By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against TensorWave's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on TensorWave's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    TensorWave's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.