ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. THE ROLE: AMD is looking for a highly motivated and experienced Server Platform Debug Engineer with strong expertise in Reliability, Availability, and Serviceability (RAS) technologies to lead complex platform investigations and drive issue resolution across AMD's next-generation server platforms. This role requires significantly more than firmware development expertise. The successful candidate must possess strong platform-level debugging skills and the ability to investigate issues spanning BIOS, BMC, silicon, operating systems, drivers, hardware, validation environments, and customer platforms. The engineer is expected to understand the complete system architecture, analyze interactions across multiple domains, and use evidence from logs, traces, registers, schematics, crash data, and validation results to identify root causes and deliver robust solutions. The ideal candidate will serve as the technical owner and coordinator for critical customer and internal platform issues. In addition to hands-on debugging, this individual will lead cross-functional technical discussions, define debug strategies, establish clear problem statements and ownership, identify missing data or incomplete investigations, close validation gaps, assign and track action items, and ensure all required workstreams are properly executed. Success in this role requires the ability to recognize what has been completed, what remains unresolved, and what additional analysis or validation is needed to accelerate issue closure. THE PERSON: The candidate must be comfortable leading technical meetings and coordinating efforts across BIOS, BMC, silicon, operating system, driver, validation, hardware, tool, and customer teams. The engineer will drive issues end to end, from initial problem definition and data collection through issue isolation, root-cause identification, corrective action, solution validation, and final closure, while escalating technical risks and blockers when necessary. Strong RAS expertise is a required core competency for this position. The candidate should have hands-on experience with platform error handling, containment and recovery flows, firmware and operating-system error reporting, in-band and out-of-band telemetry, fault injection, and system-level reliability validation. The ability to connect RAS architecture and error behavior with real customer platform issues is essential. The successful candidate will combine broad platform knowledge, deep technical expertise, strong ownership, structured problem-solving, and the ability to influence without direct authority. This individual should be passionate about solving challenging system-level problems, improving platform quality, strengthening debug and issue-management processes, and enabling successful customer deployments. KEY RESPONSIBILITIES Own complex customer and internal server platform issues from initial report through verified closure, serving as the technical lead for the overall investigation. Lead cross-functional debug meetings, establish clear problem statements, align teams on investigation priorities, and maintain an actionable debug plan. Identify missing data, incomplete analysis, validation gaps, unclear ownership, and unexecuted actions that could delay root-cause identification or issue closure. Assign and track action items, follow up on deliverables, escalate blockers and technical risks, and ensure required workstreams are completed on schedule. Apply BIOS/UEFI and firmware knowledge as part of broader platform investigations, with particular focus on RAS features, error-handling flows, containment, reporting, and recovery behavior. Perform platform-level debugging across processor, memory, PCIe, CXL, interconnect, BMC, operating system, driver, and platform-management domains rather than limiting analysis to firmware. Analyze complex hardware, firmware, and software failures using BIOS logs, POST codes, register data, CPER and event records, traces, crash data, schematics, and platform debug tools. Drive structured root-cause analysis by defining reproduction conditions, forming technical hypotheses, designing experiments, isolating failing components, and validating conclusions with evidence. Design and execute error-injection and fault-validation strategies for correctable, uncorrectable, fatal, and recoverable error scenarios. Validate firmware, operating-system, and management reporting paths, including in-band and out-of-band error reporting, event logging, telemetry, containment, and recovery behavior. Work with IBV code bases and customer BIOS implementations to integrate fixes, verify workarounds, perform regression testing, and ensure solution readiness. Collaborate with silicon, AGESA/firmware, BMC, OS, driver, validation, tools, hardware, and customer teams to resolve cross-component issues. Support customer platform bring-up, feature enablement, technical workshops, and critical escalations, including on-site support when required. Provide concise technical status updates that clearly communicate issue impact, current findings, open questions, ownership, next actions, risks, and closure criteria. Create technical specifications, debug guides, validation procedures, training materials, and reusable knowledge to improve team and customer capabilities. Contribute to automation that improves log collection, error analysis, test execution, result tracking, action-item management, and issue triage. BASIC QUALIFICATIONS: Professional experience in server platform debug, BIOS, UEFI, embedded firmware, platform firmware, or low-level system software development and validation. Strong platform-level debugging capability, with experience analyzing interactions among hardware, firmware, operating systems, drivers, management controllers, and validation environments. Strong C programming skills and experience with firmware development, code review, source control, and defect tracking. Solid understanding of x86 server architecture, multiprocessor systems, memory architecture, PCIe, ACPI, power management, and system management. Proven ability to lead complex technical investigations across multiple teams, define ownership and next actions, and drive issues to closure. Strong analytical and structured problem-solving skills, including issue reproduction, data collection, isolation, root-cause analysis, corrective action, and validation. Excellent meeting facilitation, stakeholder coordination, action-item tracking, and written and spoken English communication skills. PREFERRED EXPERIENCE: Deep knowledge of server RAS concepts, architecture, firmware implementation, platform error handling, containment, reporting, and recovery. Experience with Machine Check Architecture, error severity classification, recovery policy, error containment, and reset/recovery flows. Experience with memory RAS, including ECC, parity, patrol scrub, retry and recovery behavior, sparing, and error reporting. Experience with PCIe Advanced Error Reporting, DPC, hot-plug, link failure handling, and PCIe/CXL error scenarios. Experience with ACPI-based firmware and OS error interfaces, including EINJ, ERST, HEST, BERT, and CPER-based reporting. Experience with in-band and out-of-band RAS reporting and BIOS-to-BMC/platform-management interactions. Experience using hardware and firmware debug tools, protocol analyzers, JTAG, logic analyzers, register-access tools, error-injection frameworks, and electrical or board-level debug data. Experience correlating evidence across firmware logs, OS/kernel logs, crash dumps, BMC logs, hardware traces, schematics, and validation results. Experience with server BIOS vendors or code bases such as AMI, Insyde, or Phoenix. Experience with Linux or Windows server environments, virtualization, kernel/driver logs, crash analysis, and system-level validation. Experience developing scripts or automation in Python, shell, or similar languages for test, data collection, log analysis, and workflow tracking. Experience supporting cloud, hyperscale, OEM, ODM, or enterprise server customers is a plus. PERSONAL ATTRIBUTES Strong technical ownership and a genuine interest in driving the overall issue-resolution process, not only completing an assigned firmware task. Curious and system-oriented, with the willingness to investigate beyond firmware and understand the full platform behavior. Comfortable leading technical meetings, challenging assumptions, identifying missing work, and requesting the evidence needed for informed decisions. Highly organized and disciplined in defining owners, tracking action items, following up on commitments, and maintaining issue momentum. Able to influence without direct authority and align global cross-functional teams around a common debug strategy. Customer-focused and accountable, with the ability to manage high-priority escalations and drive issues through verified closure. Proactive in knowledge sharing, documentation, mentoring, and continuous improvement of debug and validation processes. Detail-oriented, data-driven, and committed to platform quality, reliability, and engineering excellence. ACADEMIC CREDENTIALS: Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Electronic Engineering, or a related field, or equivalent practical experience. LOCATION: Nangang - Taiwan #LI-VJ1 Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.
THE ROLE: AMD is looking for a highly motivated and experienced Server Platform Debug Engineer with strong expertise in Reliability, Availability, and Serviceability (RAS) technologies to lead complex platform investigations and drive issue resolution across AMD's next-generation server platforms. This role requires significantly more than firmware development expertise. The successful candidate must possess strong platform-level debugging skills and the ability to investigate issues spanning BIOS, BMC, silicon, operating systems, drivers, hardware, validation environments, and customer platforms. The engineer is expected to understand the complete system architecture, analyze interactions across multiple domains, and use evidence from logs, traces, registers, schematics, crash data, and validation results to identify root causes and deliver robust solutions. The ideal candidate will serve as the technical owner and coordinator for critical customer and internal platform issues. In addition to hands-on debugging, this individual will lead cross-functional technical discussions, define debug strategies, establish clear problem statements and ownership, identify missing data or incomplete investigations, close validation gaps, assign and track action items, and ensure all required workstreams are properly executed. Success in this role requires the ability to recognize what has been completed, what remains unresolved, and what additional analysis or validation is needed to accelerate issue closure. THE PERSON: The candidate must be comfortable leading technical meetings and coordinating efforts across BIOS, BMC, silicon, operating system, driver, validation, hardware, tool, and customer teams. The engineer will drive issues end to end, from initial problem definition and data collection through issue isolation, root-cause identification, corrective action, solution validation, and final closure, while escalating technical risks and blockers when necessary. Strong RAS expertise is a required core competency for this position. The candidate should have hands-on experience with platform error handling, containment and recovery flows, firmware and operating-system error reporting, in-band and out-of-band telemetry, fault injection, and system-level reliability validation. The ability to connect RAS architecture and error behavior with real customer platform issues is essential. The successful candidate will combine broad platform knowledge, deep technical expertise, strong ownership, structured problem-solving, and the ability to influence without direct authority. This individual should be passionate about solving challenging system-level problems, improving platform quality, strengthening debug and issue-management processes, and enabling successful customer deployments. KEY RESPONSIBILITIES Own complex customer and internal server platform issues from initial report through verified closure, serving as the technical lead for the overall investigation. Lead cross-functional debug meetings, establish clear problem statements, align teams on investigation priorities, and maintain an actionable debug plan. Identify missing data, incomplete analysis, validation gaps, unclear ownership, and unexecuted actions that could delay root-cause identification or issue closure. Assign and track action items, follow up on deliverables, escalate blockers and technical risks, and ensure required workstreams are completed on schedule. Apply BIOS/UEFI and firmware knowledge as part of broader platform investigations, with particular focus on RAS features, error-handling flows, containment, reporting, and recovery behavior. Perform platform-level debugging across processor, memory, PCIe, CXL, interconnect, BMC, operating system, driver, and platform-management domains rather than limiting analysis to firmware. Analyze complex hardware, firmware, and software failures using BIOS logs, POST codes, register data, CPER and event records, traces, crash data, schematics, and platform debug tools. Drive structured root-cause analysis by defining reproduction conditions, forming technical hypotheses, designing experiments, isolating failing components, and validating conclusions with evidence. Design and execute error-injection and fault-validation strategies for correctable, uncorrectable, fatal, and recoverable error scenarios. Validate firmware, operating-system, and management reporting paths, including in-band and out-of-band error reporting, event logging, telemetry, containment, and recovery behavior. Work with IBV code bases and customer BIOS implementations to integrate fixes, verify workarounds, perform regression testing, and ensure solution readiness. Collaborate with silicon, AGESA/firmware, BMC, OS, driver, validation, tools, hardware, and customer teams to resolve cross-component issues. Support customer platform bring-up, feature enablement, technical workshops, and critical escalations, including on-site support when required. Provide concise technical status updates that clearly communicate issue impact, current findings, open questions, ownership, next actions, risks, and closure criteria. Create technical specifications, debug guides, validation procedures, training materials, and reusable knowledge to improve team and customer capabilities. Contribute to automation that improves log collection, error analysis, test execution, result tracking, action-item management, and issue triage. BASIC QUALIFICATIONS: Professional experience in server platform debug, BIOS, UEFI, embedded firmware, platform firmware, or low-level system software development and validation. Strong platform-level debugging capability, with experience analyzing interactions among hardware, firmware, operating systems, drivers, management controllers, and validation environments. Strong C programming skills and experience with firmware development, code review, source control, and defect tracking. Solid understanding of x86 server architecture, multiprocessor systems, memory architecture, PCIe, ACPI, power management, and system management. Proven ability to lead complex technical investigations across multiple teams, define ownership and next actions, and drive issues to closure. Strong analytical and structured problem-solving skills, including issue reproduction, data collection, isolation, root-cause analysis, corrective action, and validation. Excellent meeting facilitation, stakeholder coordination, action-item tracking, and written and spoken English communication skills. PREFERRED EXPERIENCE: Deep knowledge of server RAS concepts, architecture, firmware implementation, platform error handling, containment, reporting, and recovery. Experience with Machine Check Architecture, error severity classification, recovery policy, error containment, and reset/recovery flows. Experience with memory RAS, including ECC, parity, patrol scrub, retry and recovery behavior, sparing, and error reporting. Experience with PCIe Advanced Error Reporting, DPC, hot-plug, link failure handling, and PCIe/CXL error scenarios. Experience with ACPI-based firmware and OS error interfaces, including EINJ, ERST, HEST, BERT, and CPER-based reporting. Experience with in-band and out-of-band RAS reporting and BIOS-to-BMC/platform-management interactions. Experience using hardware and firmware debug tools, protocol analyzers, JTAG, logic analyzers, register-access tools, error-injection frameworks, and electrical or board-level debug data. Experience correlating evidence across firmware logs, OS/kernel logs, crash dumps, BMC logs, hardware traces, schematics, and validation results. Experience with server BIOS vendors or code bases such as AMI, Insyde, or Phoenix. Experience with Linux or Windows server environments, virtualization, kernel/driver logs, crash analysis, and system-level validation. Experience developing scripts or automation in Python, shell, or similar languages for test, data collection, log analysis, and workflow tracking. Experience supporting cloud, hyperscale, OEM, ODM, or enterprise server customers is a plus. PERSONAL ATTRIBUTES Strong technical ownership and a genuine interest in driving the overall issue-resolution process, not only completing an assigned firmware task. Curious and system-oriented, with the willingness to investigate beyond firmware and understand the full platform behavior. Comfortable leading technical meetings, challenging assumptions, identifying missing work, and requesting the evidence needed for informed decisions. Highly organized and disciplined in defining owners, tracking action items, following up on commitments, and maintaining issue momentum. Able to influence without direct authority and align global cross-functional teams around a common debug strategy. Customer-focused and accountable, with the ability to manage high-priority escalations and drive issues through verified closure. Proactive in knowledge sharing, documentation, mentoring, and continuous improvement of debug and validation processes. Detail-oriented, data-driven, and committed to platform quality, reliability, and engineering excellence. ACADEMIC CREDENTIALS: Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, Electronic Engineering, or a related field, or equivalent practical experience. LOCATION: Nangang - Taiwan #LI-VJ1
Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.
Seen 5 hours ago · within 59 minutes of the employer posting it.
Original posting on AMD's site ↗
Posting text belongs to the employer. Removal requests: contact us.
Nearby
Live postings like this one
Same employer first, then the same role elsewhere.
- 3h ago
- 5h ago
- 5h ago
- 6h ago
- 7h ago
- 13h ago
- 14h ago
- 15h ago
One job at a time
One posting. One CV. $25.
Pick the job you actually want and we write for it.