Skip to content

Senior Datacenter Platform/Debug Engineer

AMD

Austin, Texas, United States

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. THE ROLE: Join AMD's Datacenter Platform Engineering Group (DPEG) and help support the deployment, availability, and operational success of next-generation AI and HPC infrastructure. As a Platform Systems Engineer, you will work on cutting-edge GPU and server platforms, partnering with hardware, firmware, software, validation, and datacenter engineering teams to troubleshoot complex system issues and ensure reliable operation of large-scale compute environments. This role offers the opportunity to work directly with advanced datacenter technologies, participate in system bring-up and deployment activities, and become a key contributor in resolving critical platform-level issues. Ideal candidates enjoy solving challenging technical problems, collaborating across multiple engineering disciplines, and making a direct impact on the success of AMD's datacenter infrastructure. THE PERSON: The ideal candidate is a hands-on systems engineer who enjoys deep technical troubleshooting and thrives in fast-paced datacenter environments. They possess strong analytical skills, can quickly isolate and resolve complex issues, and are comfortable working across hardware, firmware, and software layers of a system. Successful candidates will demonstrate: Strong troubleshooting and root cause analysis skills Excellent communication and collaboration abilities A proactive and self-driven approach to problem solving Ability to mentor and guide junior engineers Strong documentation and organizational skills Comfort operating in highly technical and mission-critical environments A passion for learning new technologies and solving complex engineering challenges KEY RESPONSIBILITIES: Support datacenter deployments and help maintain the availability and uptime of large-scale compute systems. Perform system-level debugging and triage across hardware, firmware, software, and operating system layers. Investigate and resolve complex platform issues impacting GPU and server infrastructure. Support system bring-up, initialization, validation, and operational readiness activities. Utilize industry-standard debug tools and diagnostic methods to identify root causes. Provide technical leadership and guidance to junior engineers during troubleshooting activities. Document debug methodologies, troubleshooting procedures, and best practices. Collaborate with cross-functional engineering teams to drive issue resolution and continuous improvement. PREFERRED EXPERIENCE: System-level hardware, firmware, and software debugging experience Datacenter, server, HPC, or AI infrastructure environments Root cause analysis and triage of complex platform issues GPU, PCIe, memory, retimer, networking, and system architecture knowledge RAS (Reliability, Availability, Serviceability) concepts and methodologies Linux operating system experience Python, Bash, or similar scripting experience Hands-on experience with industry-standard debug tools and diagnostics Server bring-up, system initialization, and validation activities Technical leadership, mentoring, and cross-functional collaboration ACADEMIC CREDENTIALS: Bachelor's or Master’s degree preferred in Computer Engineering, Electrical Engineering, Computer Science, or a related technical discipline LOCATION: Rockdale, Texas | 100% Onsite THIS ROLE IS NOT ELIGIBLE FOR VISA SUPPORT LI-CS1 Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.

THE ROLE: Join AMD's Datacenter Platform Engineering Group (DPEG) and help support the deployment, availability, and operational success of next-generation AI and HPC infrastructure. As a Platform Systems Engineer, you will work on cutting-edge GPU and server platforms, partnering with hardware, firmware, software, validation, and datacenter engineering teams to troubleshoot complex system issues and ensure reliable operation of large-scale compute environments. This role offers the opportunity to work directly with advanced datacenter technologies, participate in system bring-up and deployment activities, and become a key contributor in resolving critical platform-level issues. Ideal candidates enjoy solving challenging technical problems, collaborating across multiple engineering disciplines, and making a direct impact on the success of AMD's datacenter infrastructure. THE PERSON: The ideal candidate is a hands-on systems engineer who enjoys deep technical troubleshooting and thrives in fast-paced datacenter environments. They possess strong analytical skills, can quickly isolate and resolve complex issues, and are comfortable working across hardware, firmware, and software layers of a system. Successful candidates will demonstrate: Strong troubleshooting and root cause analysis skills Excellent communication and collaboration abilities A proactive and self-driven approach to problem solving Ability to mentor and guide junior engineers Strong documentation and organizational skills Comfort operating in highly technical and mission-critical environments A passion for learning new technologies and solving complex engineering challenges KEY RESPONSIBILITIES: Support datacenter deployments and help maintain the availability and uptime of large-scale compute systems. Perform system-level debugging and triage across hardware, firmware, software, and operating system layers. Investigate and resolve complex platform issues impacting GPU and server infrastructure. Support system bring-up, initialization, validation, and operational readiness activities. Utilize industry-standard debug tools and diagnostic methods to identify root causes. Provide technical leadership and guidance to junior engineers during troubleshooting activities. Document debug methodologies, troubleshooting procedures, and best practices. Collaborate with cross-functional engineering teams to drive issue resolution and continuous improvement. PREFERRED EXPERIENCE: System-level hardware, firmware, and software debugging experience Datacenter, server, HPC, or AI infrastructure environments Root cause analysis and triage of complex platform issues GPU, PCIe, memory, retimer, networking, and system architecture knowledge RAS (Reliability, Availability, Serviceability) concepts and methodologies Linux operating system experience Python, Bash, or similar scripting experience Hands-on experience with industry-standard debug tools and diagnostics Server bring-up, system initialization, and validation activities Technical leadership, mentoring, and cross-functional collaboration ACADEMIC CREDENTIALS: Bachelor's or Master’s degree preferred in Computer Engineering, Electrical Engineering, Computer Science, or a related technical discipline LOCATION: Rockdale, Texas | 100% Onsite THIS ROLE IS NOT ELIGIBLE FOR VISA SUPPORT LI-CS1

Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.

Seen 3 days ago · within 32 minutes of the employer posting it · AMD postings close after a median of 27 days.

Original posting on AMD's site ↗

Posting text belongs to the employer. Removal requests: contact us.

Nearby

Live postings like this one

Same employer first, then the same role elsewhere.

One job at a time

One posting. One CV. $25.

Pick the job you actually want and we write for it.

Get my CV for this job