- Own the end-to-end platform architecture through architecture proposals, high-level and low-level designs, reviews, and maintenance of the baseline
- Lead and line-manage a distributed team of about twelve engineers across backend, frontend, DevOps, QA, and documentation
- Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input
- Design, build, and operate a managed Slurm service for research users
- Own controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation
- Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal
- Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement
- Own managed inference architecture, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity
- Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers
- Lead incident response, post-incident reviews, and development of a sustainable on-call model
- Serve as the primary technical interface to infrastructure partners and vendors
- Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing
- Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity
- Complete the platform team and set the technical bar for new engineers
Requirements
- Eight or more years of hands-on engineering experience
- At least three years leading teams that build and operate infrastructure platforms other teams depend on
- Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience
- Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system
- Experience operating an HPC or GPU training cluster for a research population is ideally preferred
- Experience operating NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance
- Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems
- Deep Linux systems knowledge, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning
- Production Kubernetes operations experience, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design
- Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets
- Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review
- Working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions
- Experience shipping a platform with real users, such as a multi-tenant IaaS/PaaS or research computing service
- People management across time zones, cross-track review, written architecture decisions, and partner/executive communication
- Excellent written and spoken English
- Based between UTC and UTC+5:30
- Desirable: Slurm operators on Kubernetes or Kubernetes-native schedulers
- Desirable: modern serving stacks such as vLLM, SGLang, and TensorRT-LLM
- Desirable: VM/container isolation, confidential computing, Cluster API, kubeadm, Cilium, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and hardware-provider partnership experience
#J-18808-Ljbffr
Technical Lead – GPU Infrastructure Arbeitgeber: Jobtailor
Als Front Office Supervisor in unserem dynamischen Team bieten wir Ihnen die Möglichkeit, in einem unterstützenden und freundlichen Arbeitsumfeld zu wachsen. Wir legen großen Wert auf die berufliche Entwicklung unserer Mitarbeiter und bieten regelmäßige Schulungen sowie die Chance, Verantwortung zu übernehmen. Unsere Lage ermöglicht es Ihnen, Teil einer lebendigen Gemeinschaft zu sein, während Sie gleichzeitig die Standards unseres Franchise-Partners einhalten und unseren Gästen einen unvergesslichen Aufenthalt bieten.