Infrastructure Engineer (GPU & Compute)

Infrastructure Engineer (GPU & Compute)

Vollzeit Kein Homeoffice möglich
L

Role overview

This position focuses on bringing up, validating, and operating large-scale bare-metal compute environments with an emphasis on GPU-enabled systems. The engineer will sit at the intersection of hardware, systems software, and automation, owning diagnostics, qualification, and the tooling that keeps clusters ready for demanding AI/ML and HPC workloads.

Responsibilities

  • Own and evolve image management, deployment, and validation pipelines across bare-metal infrastructure, including firmware, driver, and OS qualification for GPU-enabled systems
  • Operate and maintain test clusters used for bring-up and diagnostics, supporting hardware qualification efforts for next-generation platforms
  • Diagnose and resolve complex issues spanning GPUs, drivers, OS, and underlying hardware; analyze performance using tools such as NVIDIA DCGM
  • Build Python-based automation for provisioning, validation, and system bring-up, improving reliability, repeatability, and scalability
  • Manage Linux-based production and validation environments, including virtualization and PXE/image-based bare-metal provisioning workflows
  • Partner with infrastructure, hardware, and data center teams, plus platform and ML stakeholders, to ensure systems meet workload requirements and contribute to provisioning and lifecycle best practices

Requirements

  • 5+ years in infrastructure engineering, systems engineering, or a closely related role
  • Strong Linux systems experience in production environments
  • Hands-on experience with GPU-enabled systems and diagnostic tooling such as NVIDIA DCGM
  • Familiarity with bare-metal provisioning and system bring-up workflows
  • Proficiency in Python or comparable scripting/programming languages for automation
  • Ability to debug complex issues across hardware, OS, GPUs, and system software

Nice to have

  • Experience with high-performance interconnects such as InfiniBand or NVLink
  • Familiarity with PXE boot environments, LiveCD systems, or image-based provisioning workflows
  • Working knowledge of hardware management interfaces (iDRAC, IPMI, Redfish)
  • Data center operations experience with physical hardware
  • Background supporting AI/ML or HPC workloads at scale
  • Experience with GPU validation frameworks or large-scale hardware qualification processes

Benefits and work setup

  • Anticipated annual base salary range of $180,000–$220,000 USD, plus discretionary bonus and equity
  • Comprehensive medical, dental, and vision coverage for employees and eligible dependents
  • Retirement savings support (401(k) matching in the U.S., pension contributions in the U.K.)
  • Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break
  • Paid parental and family leave, plus a four-week paid sabbatical after four years of service
  • Annual professional development allowance, plus wellness and work-from-home stipends
  • Flexible schedules with a hybrid model for office-based teams and complimentary in-office meals at hubs
  • Work may be fully remote within the U.S. or hybrid out of office hubs in NYC, SF, Seattle, or London, with occasional team and company offsites; visa sponsorship is not available for this role

#J-18808-Ljbffr

L

Kontaktdaten:

lightningai Recruiting-Team