Skip to content

Open nowFirst seen 4 hours ago

DPU RAS and Debug Architect, Infrastructure Silicon

Meta1,030 open roles

Where
Sunnyvale, CA; Austin, TX; Menlo Park, CA
Get the CV for this job

From $25 per CV, paid once. No subscription.

Your applicationOpen nowDPU RAS and Debug Architect, Infrastructure SiliconMeta · Sunnyvale, CA; Austin, TX; Menlo Park, CA
  1. YouYes, apply to this one.

  2. CV RocketCV written for this posting.

  3. 25 readersRecruiter, hiring manager, skeptic. Round after round.

  4. CV RocketApplied on Meta's own form.

The reply lands in your private mailbox

3×more interviews than doing it yourself with ChatGPT.

The clock on this job

Early applications get read.

8.0% of postings close within 7 days. Measured by our own scanner across the market. Meta postings stay open a median of 35 days.

Share of postings closed within
  1. 1.6%1 day
  2. 3.6%3 days
  3. 8.0%7 days
  4. 15.0%14 days
  5. 34.2%30 days
This job: first seen 4 hours ago

Meta median: 35 days open

The posting

Meta's Infrastructure Silicon organization designs custom silicon that powers our data center infrastructure — SmartNICs/IPUs/DPUs, AI accelerators, and networking ASICs.

We are looking for a Data Processing Unit (DPU) RAS and Debug Architect to define the reliability, availability, and serviceability architecture for our DPU ASICs. You'll set FIT-rate targets from the intended usages and deployment models, define how errors are detected, corrected, and reported, and define the trace, debug, and performance-monitoring architecture that supports both post-silicon debug and software and firmware debugging. You'll decide what belongs in hardware versus firmware versus software, define how the hardware presents itself to firmware and software, and work hand in hand with the firmware, driver, and RTL teams - carrying your designs from early path-finding through implementation and silicon bring-up.

Responsibilities

  • Own the RAS and debug architecture for our DPUs - error detection, correction, and reporting, FIT budgeting, and the trace, debug, and telemetry infrastructure - from early path-finding through silicon bring-up
  • Set FIT-rate targets from the intended usages and deployment models, and budget them across the design - memories, logic, on-chip interfaces, and links
  • Define the error detection, correction, reporting, and containment architecture -- parity, ECC/SECDED, poisoning and poison propagation, error logging, and how errors are surfaced to firmware, host software, and platform management
  • Define the trace, debug, and performance-monitoring architecture for post-silicon debug and for software and firmware debugging - on-chip trace, event and counter telemetry, crash and state capture, and the JTAG/debug-access model
  • Decide what belongs in hardware versus firmware versus software, and define how the hardware presents itself to them - register and programming models, error and interrupt models, and trace/telemetry interfaces
  • Work closely with design, DV, and PD teams on feature definition and PPA tradeoff; refine architecture to meet the design and PD constraints. Define architecture to maximize DV complexity including defining specific features to help ease DV
  • Support Design, DV, PD and DFT teams to resolve interface and integration issues as they come up. Support post-silicon bring-up

Minimum Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • 8+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs
  • Experience with RAS concepts: FIT-rate estimation and budgeting, failure modes (including silent data corruption), error detection, correction, and containment, and reliability targets for data center deployments. Familiarity with data center reliability, serviceability, and manageability requirements
  • Experience with error-protection mechanisms - parity, ECC/SECDED, CRC, data poisoning and poison propagation, and lockstep/redundancy techniques - and their area, latency, and power trade-offs
  • Experience with memory and interface RAS -- LPDDR/DDR RAS (ECC, on-die ECC, error scrubbing, post-package repair) and PCIe RAS (Advanced Error Reporting, ECRC, link error detection and recovery)
  • Experience with error reporting and handling architectures (ARM and/or x86) - error logging and registers, interrupts, machine-check and AER-style reporting, and escalation to firmware, host software, and platform/BMC management
  • Experience with on-chip debug and trace architectures -- JTAG and debug access, on-chip trace, breakpoints and watchpoints, and crash and state capture (e.g., ARM CoreSight or comparable)
  • Experience with performance-monitoring architectures - hardware performance counters and event telemetry - and their use in post-silicon and software/firmware debug
  • Experience with DFT concepts (scan, MBIST/LBIST, boundary scan) and how they interact with RAS and debug
  • Experience with processor ISA debug mechanisms (ARM and/or x86) and instruction/execution trace mechanisms such as ETM
  • Experience with relevant industry standards and specifications, such as OCP (server, RAS, and telemetry/manageability specifications), JEDEC (LPDDR/DDR), and PCIe
  • Experience driving analysis independently and influencing architectural direction through data

Preferred Qualifications

  • PhD in Computer Science, Computer Engineering or Electrical Engineering
  • 15+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs
  • Experience with hardware description languages (e.g., SystemVerilog, VHDL) and simulation environments used in ASIC development flows
  • Experience designing on-chip trace and debug subsystems (e.g., ARM CoreSight) and the associated post-silicon debug tooling
  • Experience defining FIT budgets and reliability targets for hyperscale data center silicon, including soft-error rate (SER) analysis and mitigation
  • Familiarity with functional-safety standards (e.g., ISO 26262) and reliability qualification methods
  • Experience developing Python-based (or other scripting) automation pipelines for debug, telemetry collection, and data analysis
  • Experience designing error-reporting and machine-check/AER architectures and the firmware error-handling flows built on them

US: $178,000/year to $250,000/year + bonus + equity

From $25, paid onceGet the CV for this job

What happens when you press

One press. We do the rest.

  1. A CV for this posting

    Written against Meta's own wording, from every piece of relevant proof in your profile.

  2. 25 readers review it

    Recruiter, hiring manager, skeptic and more read every draft, round after round. You get the best round.

    The review screen in CV Rocket: how each CV was read, round by round.
  3. We apply on Meta's form

    Our application engine gets through the hardest forms there are. Where a question needs you, AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

    An application in CV Rocket: every answer filled in on the employer's form.
  4. Every reply, sorted

    Meta's answer lands in your private mailbox, and we classify it on arrival: interview, question, rejection.

    The CV Rocket inbox: each employer reply classified as an interview, an action or a rejection.
  5. Reply with AI

    AI helps you write the email, checks it and sends it. We show you whether the recruiter read it.

  6. The interview in your calendar

    Full integration with your calendar. The invitation goes straight in.

    An interview invitation in the CV Rocket inbox, added to the candidate's calendar.
Get the CV for this job

From $25 per CV, paid once. No subscription.

Why it works

3×

more interviews than doing it yourself with ChatGPT.

ChatGPT writes a CV and never learns what happened to it. We see every reply. For each CV we know:

  • How it was written, and how the review scored it
  • When we applied, and how long after the posting went up
  • Which posting, which company, which city
  • Who got the interview, and who heard nothing

That is how we know which CVs get called.

Get the CV for this job

From $25 per CV, paid once. No subscription.

The numbers game

More applications. More interviews.

Every application goes out with its own CV, written for that posting and paid once. Send enough of them and the law of large numbers finds you the job.

By hand5–10
With CV Rocket100
applications a day

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

Before you press

Straight answers

Get the CV for this job

From $25 per CV, paid once. No subscription.

What if my background isn't good enough?

We make the most of the background you have. The CV uses every piece of relevant proof your profile holds, and one of the 25 readers reads your whole profile and flags what the CV left out.

Do you really apply for me?

Yes, on the employer's own form, the hardest ones included. Where a question needs you, you answer it right there and AI suggests the best answer. Don't want us applying from our IP addresses? Use our Chrome extension: we apply straight from your own browser.

Is it a subscription?

No. You pay once per CV, from $25. Every application goes out with its own CV, written for that posting.

One job. One CV.
Paid once.

Pick the posting you want. We write for it, apply for you and catch the reply.

Get the CV for this job

From $25 per CV, paid once. No subscription.