Role overview
This position focuses on bringing up, validating, and operating large-scale bare-metal compute environments with an emphasis on GPU-enabled systems. The engineer will sit at the intersection of hardware, systems software, and automation, owning diagnostics, qualification, and the tooling that keeps clusters ready for demanding AI/ML and HPC workloads.
Responsibilities
- Own and evolve image management, deployment, and validation pipelines across bare-metal infrastructure, including firmware, driver, and OS qualification for GPU-enabled systems
- Operate and maintain test clusters used for bring-up and diagnostics, supporting hardware qualification efforts for next-generation platforms
- Diagnose and resolve complex issues spanning GPUs, drivers, OS, and underlying hardware; analyze performance using tools such as NVIDIA DCGM
- Build Python-based automation for provisioning, validation, and system bring-up, improving reliability, repeatability, and scalability
- Manage Linux-based production and validation environments, including virtualization and PXE/image-based bare-metal provisioning workflows
- Partner with infrastructure, hardware, and data center teams, plus platform and ML stakeholders, to ensure systems meet workload requirements and contribute to provisioning and lifecycle best practices
Requirements
- 5+ years in infrastructure engineering, systems engineering, or a closely related role
- Strong Linux systems experience in production environments
- Hands-on experience with GPU-enabled systems and diagnostic tooling such as NVIDIA DCGM
- Familiarity with bare-metal provisioning and system bring-up workflows
- Proficiency in Python or comparable scripting/programming languages for automation
- Ability to debug complex issues across hardware, OS, GPUs, and system software
Nice to have
- Experience with high-performance interconnects such as InfiniBand or NVLink
- Familiarity with PXE boot environments, LiveCD systems, or image-based provisioning workflows
- Working knowledge of hardware management interfaces (iDRAC, IPMI, Redfish)
- Data center operations experience with physical hardware
- Background supporting AI/ML or HPC workloads at scale
- Experience with GPU validation frameworks or large-scale hardware qualification processes
Benefits and work setup
- Anticipated annual base salary range of $180,000–$220,000 USD, plus discretionary bonus and equity
- Comprehensive medical, dental, and vision coverage for employees and eligible dependents
- Retirement savings support (401(k) matching in the U.S., pension contributions in the U.K.)
- Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break
- Paid parental and family leave, plus a four-week paid sabbatical after four years of service
- Annual professional development allowance, plus wellness and work-from-home stipends
- Flexible schedules with a hybrid model for office-based teams and complimentary in-office meals at hubs
- Work may be fully remote within the U.S. or hybrid out of office hubs in NYC, SF, Seattle, or London, with occasional team and company offsites; visa sponsorship is not available for this role
#J-18808-Ljbffr