Overview
In this role you guide EMEA customers in deploying and optimizing large-scale AI inference on multi-node GPU clusters. You act as a technical lead to ensure the NVIDIA inference stack delivers peak performance, efficiency, and reliability in demanding production environments. You will shape solutions for dense and MoE models, collaborate with product teams, and drive community engagement through workshops and reference architectures. This position offers a meaningful impact on scalable AI inference across leading enterprises and AI-native ecosystems.
Leistungen / Benefits
- comprehensive benefits package
- competitive salaries
Verantwortungsbereiche
- Guide EMEA customers in deploying and optimizing large-scale inference workloads on multi-node GPU clusters
- Architect efficient inference pipelines for dense and sparse/latent MoE models distributing workloads across thousands of GPUs
- Improve inference efficiency through quantization, speculative decoding, disaggregated prefill/decode, KV cache management, and WideEP for MoE deployments
- Collaborate with NVIDIA product teams (Dynamo, TensorRT-LLM, NIXL) to accelerate customer success
- Lead and animate the AI inference developer community across EMEA via technical workshops, hackathons, and reference architectures
Zentrale Anforderungen
- MS or PhD in Computer Science, Engineering, High-Performance Computing, or equivalent professional experience
- 5+ years of experience in Neural Networks inference optimization
- Strong understanding of transformers inference optimization (quantization, disaggregated inference, speculative decoding, continuous batching, KV cache optimization)
- Hands-on experience with MoE inference at scale (expert parallelism, WideEP, all-to-all communication, routing overhead, load balancing)
- Ability to engage with ML engineers, researchers, and systems architects at a deep technical level
- strong collaboration across cross-functional teams
- deep technical communication with both researchers and engineers
- customer-focused technical leadership
- NVIDIA Dynamo
- NIXL
- TensorRT-LLM
Senior Solutions Architect – Large Scale AI Inference in Halle Arbeitgeber: NVIDIA
NVIDIA ist ein herausragender Arbeitgeber, der nicht nur wettbewerbsfähige Gehälter und umfassende Sozialleistungen bietet, sondern auch eine dynamische Arbeitskultur fördert, die Innovation und Zusammenarbeit in der Robotikbranche anregt. In dieser Schlüsselposition in der EMEA-Region haben Sie die Möglichkeit, ein leistungsstarkes Team zu leiten, strategische Beziehungen aufzubauen und bedeutende Wachstumschancen zu entwickeln, während Sie gleichzeitig von einem Umfeld profitieren, das persönliche und berufliche Weiterentwicklung unterstützt. Werden Sie Teil eines Unternehmens, das die Zukunft intelligenter Maschinen gestaltet und dabei auf eine positive Work-Life-Balance Wert legt.