ML Platform Engineer

ML Platform Engineer

Vollzeit 63000 - 77000 € / Jahr (geschätzt) Kein Homeoffice möglich
K

Auf einen Blick

  • Aufgaben: Entwickle und skaliere die Infrastruktur für KI-gestützte klinische Assistenzsysteme.
  • Unternehmen: Kaiko, ein innovatives Unternehmen im Gesundheitswesen mit internationalem Team.
  • Vorteile: Attraktives Gehalt, 25 Urlaubstage, Weiterbildungsgeld und flexible Arbeitszeiten.
  • Weitere Informationen: Dynamisches Umfeld mit großartigen Teamevents und Entwicklungsmöglichkeiten.
  • Warum dieser Job: Gestalte die Zukunft der Gesundheitsversorgung mit modernster Technologie und direktem Einfluss.
  • Qualifikationen: Erfahrung in ML-Plattformen und Teamarbeit, Offenheit für neue Programmiersprachen.

Das prognostizierte Gehalt liegt zwischen 63000 - 77000 € pro Jahr.

Kaiko is building a next-generation agentic clinical AI assistant that helps clinicians reason across patient data, guidelines, and diagnostics.

Healthcare decisions are rarely made by a single person or from a single data source. kaiko's assistant maintains longitudinal patient context across encounters, clinicians, and institutions, enabling collaboration, second opinions, and complex diagnostic workflows.

The system is designed to operate safely in real clinical environments, with human oversight, auditability, and regulatory alignment at its core.

Our assistant core supports broadly applicable clinical tasks such as patient data navigation, guideline interaction, multimodal interaction (chat and voice), and care coordination.

On top of this foundation, we are developing specialized diagnostic agents in areas such as oncology, radiology, and pathology.

We build in close collaboration with leading hospitals and research centers, including the Netherlands Cancer Institute (NKI). kaiko is a well-funded company with a growing international team, operating from Zurich and Amsterdam.

About the role As an ML Platform Engineer, you'll help architect, scale, and evolve the infrastructure that powers kaiko's foundation-model training and serving across the full ML lifecycle: from compute orchestration that runs large-scale training jobs, through experiment tracking and model registry, to GPU-backed serving in production.

The Data openness to picking up Go or other programming languages where the platform demands it.

Collaborative by default.

You work across team boundaries (research, product, data engineering) and design interfaces that serve multiple stakeholders without requiring everyone to understand the underlying plumbing.

Nice to have: Direct experience supporting large-scale foundation-model training or inference, with a working sense of how a run behaves from the platform's perspective: MFU, GPU utilisation, communication overhead, and dataloader stalls on the training side; serving frameworks (v LLM, Triton, Torch Serve), KV cache management, and batching strategies on the inference side.

Familiarity with lower-level GPU communication and I/O concepts: RDMA, GPUDirect Storage (GDS), GPUDirect RDMA, NCCL, and how these affect training throughput and cluster design.

Experience with Kubernetes-native scheduling stacks for accelerator-heavy workloads, such as Volcano, KAI Scheduler, Yuni Korn, or similar, so scheduler tradeoffs are understood rather than defaulted to.

High performance (parallel) filesystems that feed GPU's with high throughput and low latency (Hammerspace, CEPH, WEKA, VAST).

On-prem and hybrid GPU environments, especially in regulated settings (healthcare, finance, public sector).

Exposure to Mo E architectures or large-scale distributed training, useful for anticipating what future workloads will demand of the platform.

We are excited to gather a broad range of perspectives in our team, as we believe it will help us build better products to support a broader set of people.

If you’re excited about us but don’t fit every single qualification, we still encourage you to apply: we’ve had incredible team members join us who didn’t check every box!

At kaiko, we believe the best ideas come from collaboration, ownership and ambition.

We’ve built a team of international experts where your work has direct impact.

Here’s what we value: Ownership: You’ll have the autonomy to set your own goals, make critical decisions, and see the direct impact of your work.

Collaboration: You’ll have to approach disagreement with curiosity, build on common ground and create solutions together.

Ambition: You’ll be surrounded by people who set high standards for themselves and others, who see obstacles as opportunities, and who are relentless in their work to create better outcomes for patients.

In addition, we offer: An attractive and competitive salary, a good pension plan and 25 vacation days per year.

Great offsites and team events to strengthen the team and celebrate successes together.

A EUR 1000 learning and development budget to help you grow.

Autonomy to do your work the way that works best for you, whether you have a kid or prefer early mornings.

An annual commuting subsidy.

Our interview process is designed to assess mutual fit across skills, motivation, and values.

It typically includes the following steps: Screening call: A short conversation to align on your motivation, career goals, and initial fit for the role.

Coding assessment

ML Platform Engineer Arbeitgeber: kaiko.ai

Kaiko.ai in Zürich ist ein hervorragender Arbeitgeber, der eine dynamische und innovative Arbeitsumgebung bietet, in der Mitarbeiter die Möglichkeit haben, an bedeutenden Projekten im Bereich der Gesundheitsdaten und KI zu arbeiten. Mit einem wettbewerbsfähigen Gehalt, einem Lernbudget und 25 Urlaubstagen jährlich fördert das Unternehmen die persönliche und berufliche Weiterentwicklung seiner Mitarbeiter und schafft eine Kultur des kontinuierlichen Lernens und der Zusammenarbeit.

K

Kontaktdaten:

kaiko.ai Recruiting-Team

Wir glauben, dass du diese Fähigkeiten brauchst, um ML Platform Engineer mit Bravour zu bestehen

Architektur von ML-Infrastrukturen
Skalierung von ML-Plattformen
Compute-Orchestrierung
Experiment-Tracking
Modell-Registry
GPU-gestütztes Serving
Go oder andere Programmiersprachen