Description
ADVANCE YOUR CAREER. ADVANCE THE WORLD.
At AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.
Whether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we'll advance your career.
THE ROLE:
Join AMD's growing Datacenter Infrastructure team in Rockdale, Texas, supporting the deployment, operation, and debug of next-generation high-performance computing and AI platforms. This role offers the opportunity to work directly with cutting-edge server technologies, collaborating with hardware, software, firmware, validation, and datacenter engineering teams to bring advanced computing systems online and keep them operating at peak performance.
As a Senior Member Lab Systems Engineer, you will be a technical leader within the lab environment, responsible for deploying and debugging complex server and rack-scale platforms while driving root cause analysis across hardware, firmware, software, and operating system layers. The ideal candidate enjoys solving difficult technical problems, thrives in a fast-paced environment, and is excited by the opportunity to work on the latest AMD technologies before they reach customers.
This position provides exposure to AMD's newest HPC and AI infrastructure platforms and offers opportunities to collaborate with engineering teams across multiple AMD sites. Travel to other AMD datacenter locations may be required to support deployments, troubleshooting, and operational initiatives.
THE PERSON:
The successful candidate is a highly motivated and self-directed engineer who enjoys tackling complex technical challenges and working across multiple disciplines. They are equally comfortable troubleshooting independently and collaborating with cross-functional teams to resolve critical issues.
- Strong troubleshooting and analytical thinking skills
- Proactive approach to identifying and resolving issues
- Excellent collaboration and communication skills
- Structured and methodical approach to debugging
- Ability to mentor and assist other team members
- Strong documentation and organizational skills
- Flexibility to support business-critical activities when needed
- Passion for learning emerging technologies and solving difficult technical problems
KEY RESPONSIBILITIES:
- Deploy, configure, and maintain server, rack, and datacenter infrastructure platforms.
- Execute hardware installation activities supporting lab and datacenter environments.
- Troubleshoot and debug hardware, firmware, software, operating system, and networking issues.
- Perform root cause analysis of complex system failures and drive resolution efforts.
- Utilize system diagnostics, logs, telemetry, and low-level debug tools to identify faults.
- Support execution of engineering workloads, validation activities, and test environments.
- Develop and maintain scripts and tools to improve operational efficiency and automation.
- Document investigations, test results, corrective actions, and technical procedures in a clear and structured manner.
- Collaborate with hardware, firmware, software, validation, and infrastructure teams to resolve technical challenges.
- Travel to AMD datacenter sites as needed to support deployments and operational initiatives.
PREFERRED EXPERIENCE:
- Experience supporting server, datacenter, HPC, or AI infrastructure environments
- Knowledge of computer hardware architecture including CPUs, GPUs, memory subsystems, networking, and I/O devices
- Experience debugging hardware, firmware, operating systems, and software applications
- Linux system administration experience
- Experience with Python, Bash, C, or C++ scripting/programming
- Familiarity with server setup, deployment, and administration
- Experience working in rack-scale computing environments
- Networking fundamentals including system connectivity and troubleshooting
- Strong technical documentation and reporting skills
- Experience performing structured triage, troubleshooting, and root cause analysis
- Familiarity with technologies such as PCIe, Redfish, IPMI, I2C, SPI, NVLink, xGMI, Infinity Fabric, or related protocols
- Experience supporting large-scale compute or validation environments
ACADEMIC CREDENTIALS:
- Bachelor's or master's degree preferred in Computer Engineering, Electrical Engineering, Computer Science, or a related technical discipline
LOCATION: Rockdale, Texas (Onsite)
THIS ROLE IS NOT ELIGIBLE FOR VISA SUPPORT
LI-CS1
Benefits offered are described: AMD benefits at a glance.
AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.
AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's “Responsible AI Policy” is available here.
This posting is for an existing vacancy.
Apply on company website