The Role
You will own the technical side of our largest and most complex inference deals, from first discovery call to go-live. You're the trusted technical counterpart for our customers' CTOs and ML leads, you design the setups that run their models, and you work hand in hand with our commercial team to get deals closed.
You'll be one of two senior technical owners of our inference deals. Together with our Inference Sales Engineering lead, you'll also turn what we learn in the field into a repeatable playbook and product.
What You'll Do
-
Own the technical side of our largest and most complex inference deals end to end, from discovery to go-live
-
Be the senior technical counterpart for customers' CTOs and ML leads on architecture, sizing, SLAs, latency/throughput trade-offs and cost
-
Design dedicated inference setups: model choice, GPU type and count, parallelism, quantization, inference engine and configuration
-
Lead benchmarks and proofs of concept, and turn the results into clear recommendations
-
Scope customer customizations with our engineering team and decide what becomes product
-
Build the playbook and tooling for matching workloads to GPUs, and feed our roadmap with what you learn in the field
-
Work in tandem with our commercial team on pricing, proposals and the technical parts of contracts
What We're Looking For
-
A degree in computer science, data science or a closely related field
-
4+ years in a customer-facing technical role, e.g. solutions or sales engineer, forward deployed engineer, ML engineer working closely with customers, or technical consultant
-
Hands-on experience with LLM inference and model serving (e.g. vLLM, SGLang, TensorRT-LLM, Triton) and GPU sizing
-
A track record of owning the technical side of complex B2B deals, ideally with enterprise customers
-
You're credible with customer CTOs and engineers alike, and you explain trade-offs clearly
-
Entrepreneurial mindset: give you an outcome, and you find a way without getting blocked
-
Fluent English
Bonus Points
-
Experience at an inference provider, GPU cloud or AI infrastructure company
-
Benchmarking and performance optimization: throughput, latency, cost per token
-
Experience with quantization, parallelism strategies, KV-cache and batching
-
Startup experience
-
German
Why us?
-
Outstanding team: Work with some of the best engineers in the world, coming from hedge funds, big tech, AI startups and top universities
-
Once in a lifetime opportunity: Early-stage company in the fastest-growing market in the world
-
Ownership: Shape how European AI companies access GPU compute
-
European mission: Build sovereign, GDPR-compliant AI infrastructure for the next generation of deep-tech
Forward Deployed Engineer AI Inference in Berlin Arbeitgeber: Lyceum
Als Compute Markets Associate bei uns erwartet Sie eine herausragende Arbeitsumgebung, in der Sie mit einem talentierten Team von Ingenieuren zusammenarbeiten, die aus Hedgefonds, großen Technologieunternehmen und führenden Universitäten stammen. Wir bieten Ihnen die Möglichkeit, in einem schnell wachsenden Markt zu arbeiten und aktiv an der Gestaltung der europäischen AI-Infrastruktur mitzuwirken. Darüber hinaus fördern wir Ihre persönliche und berufliche Entwicklung durch innovative Projekte und ein unterstützendes Arbeitsklima, das auf Zusammenarbeit und Kreativität setzt.