Overview
In this role you will guide NVIDIA Cloud Partners through Day 2 operations to ensure reliable, large-scale NVIDIA accelerated infrastructure in production. You’ll collaborate with partner teams to implement health, observability, remediation, and lifecycle practices that align with NVIDIA workloads and partner environments. The role sits at the intersection of distributed systems, cloud infrastructure, and production operations, offering hands-on work to scale and stabilize complex AI compute deployments. You will shape practical operating models, runbooks, and automation to sustain performance and readiness at scale.
Verantwortungsbereiche
- Lead Day 2 operational readiness initiatives for NVIDIA Cloud Partners
- Set up systems, procedures, automation, and methods to manage NVIDIA accelerated infrastructure post-deployment
- Develop continuous infrastructure validation across GPU, CPU, storage, and network for large AI clusters
- Establish observability through telemetry, monitoring, dashboards, and alerts across compute, networking, storage, Kubernetes, and AI workloads
- Create automated detection and remediation workflows to minimize disruption to workloads
- Refine fleet lifecycle management including driver/firmware lifecycle, node maintenance, OS patching, and configuration drift detection
- Translate NVIDIA reference architectures into production operating practices, validation criteria, runbooks, and standards
- Define health signals, SLOs, metrics, acceptance criteria, and ongoing validation mechanisms
- Build reusable operational frameworks, playbooks, runbooks, and reference implementations across multiple NCP environments
Zentrale Anforderungen
- 8+ years in infrastructure engineering, SRE, DevOps, or similar roles supporting large-scale production environments
- Strong experience with Linux-based distributed systems and production cloud infrastructure
- Deep understanding of Kubernetes, cluster scheduling, and large multi-node environments
- Proficiency in production observability (metrics, logging, alerting, dashboards) guided by SLAs
- Experience automating infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management
- Strong networking fundamentals with troubleshooting across compute, network, and storage layers
- Programming/automation experience with Python, Go, or shell scripting
- BS/MS/PhD in computer science, computer/electrical engineering, or related field, or equivalent experience
- Collaborative mindset for cross-team work with partners and internal teams
- Strong communication and problem-solving abilities
- Analytical thinking and attention to detail
- Kubernetes and cluster lifecycle management
- Linux-based distributed systems
- Observability tooling (Prometheus, Grafana, OpenTelemetry, Alertmanager)
Senior Engineer, NCX in Düsseldorf Arbeitgeber: NVIDIA
NVIDIA ist ein herausragender Arbeitgeber, der nicht nur wettbewerbsfähige Gehälter und umfassende Sozialleistungen bietet, sondern auch eine dynamische Arbeitskultur fördert, die Innovation und Zusammenarbeit in der Robotikbranche anregt. In dieser Schlüsselposition in der EMEA-Region haben Sie die Möglichkeit, ein leistungsstarkes Team zu leiten, strategische Beziehungen aufzubauen und bedeutende Wachstumschancen zu entwickeln, während Sie gleichzeitig von einem Umfeld profitieren, das persönliche und berufliche Weiterentwicklung unterstützt. Werden Sie Teil eines Unternehmens, das die Zukunft intelligenter Maschinen gestaltet und dabei auf eine positive Work-Life-Balance Wert legt.