Auf einen Blick
- Aufgaben: Entwickle ML-Infrastrukturen für humanoide Roboter und optimiere Trainingsdatenpipelines.
- Unternehmen: Flexion, ein innovatives Unternehmen im Bereich humanoider Robotik.
- Vorteile: Wettbewerbsfähiges Gehalt, spannende Projekte und ein dynamisches Team.
- Weitere Informationen: Tolle Karrierechancen in einem schnell wachsenden Umfeld.
- Warum dieser Job: Sei Teil der Zukunft der Robotik und arbeite an bahnbrechenden Technologien.
- Qualifikationen: Erfahrung in ML-Infrastruktur und starke Python-Kenntnisse erforderlich.
Das prognostizierte Gehalt liegt zwischen 75000 - 105000 € pro Jahr.
About Flexion At Flexion, we're building the intelligence layer powering the next generation of humanoid robots.
Our mission is to accelerate the transition from fragile prototypes to real-world humanoid deployment.
We are founded by leading scientists in robot reinforcement learning (ex-Nvidia, ex-ETH Zürich), and backed by leading international VC firms.
In just months, we’ve gone from our first line of code to deploying real humanoid capabilities.
The Role This is a senior ML engineering role, building out our core compute and data platforms.
We’re building the brain for humanoid robots, which involves training large-scale foundational models with vast amounts of data.
You'll architect the pipelines that move data from simulators and robots into model training, optimize training workloads and create the platforms that help our AI engineers train, evaluate, and iterate fast.
You'll join Flexion's experienced Infrastructure team (ex-Google, Meta, Amazon) and take significant ownership of the systems behind our data collection, training and experimentation workflows: from strategic infrastructure decisions, cluster orchestration and distributed training optimization to data platforms, CI, and experiment tooling.
This is a senior, on-site role at our Zürich office.
Key Responsibilities Architect training data platforms and pipelines: build the storage, processing, and serving layers that handle the full data lifecycle: from simulator output and robot telemetry to training datasets.
This includes building infrastructure with object storage (S3), parallel filesystems (Lustre), and common data formats (Parquet, Web Dataset, Le Robot).
Use distributed processing frameworks (Ray, Spark) to transform and validate data at scale.
Optimize distributed training: work with our AI engineers to scale workloads across multi-node GPU clusters, profiling and improving throughput, device utilization, and communication efficiency.
This includes optimizing our distributed Isaac Lab-based sim-to-real training.
Evaluate and adopt new platforms and technologies: compare cloud providers, GPUaa S platforms, and emerging tooling, owning the decisions on what we adopt as we grow our compute footprint.
Professional experience building infrastructure, tooling, or platforms for large-scale ML training workloads.
Strong experience with ML data infrastructure: distributed processing, object storage, metadata/catalog systems, dataset versioning, streaming, shuffling, caching, and high-throughput dataloading.
Hands-on experience training, scaling, or supporting distributed ML workloads, with understanding of DDP, FSDP, NCCL, checkpointing, fault tolerance, and training performance bottlenecks.
Experience with cloud infrastructure such as AWS, GCP, or similar, including compute, networking, storage, and cost/performance tradeoffs.
Experience with job scheduling or orchestration systems such as Slurm, Kubernetes, Ray, or similar.
Proficiency in Python and working knowledge of Py Torch.
Ownership mindset: comfortable making architectural decisions, setting direction, and delivering independently in a fast-moving environment.
Nice to have Experience with job scheduling and orchestration tools: Slurm, Kubernetes, or both.
Familiarity with common data formats (Parquet, Web Dataset, Le Robot).
Familiarity with robotics simulation environments (Isaac Lab, Isaac Gym, Mu Jo Co).
Experience with infrastructure-as-code and configuration management (Terraform, Ansible).
Familiarity with experiment tracking platforms (Weights
Biases, MLflow).
Competitive compensation package A front-row seat at one of Europe’s most ambitious robotics companies An energetic, collaborative team with a bias for action #J-18808-Ljbffr
Machine Learning Engineer Arbeitgeber: Flexion Robotics
pFlexion Robotics in Zürich bietet eine dynamische Arbeitsumgebung, in der Innovation und Teamarbeit großgeschrieben werden. Als IT & Security Engineer haben Sie die Möglichkeit, Ihre Fähigkeiten in einem der führenden Unternehmen der Robotikbranche einzubringen und weiterzuentwickeln. Flexible Arbeitszeiten und ein engagiertes Team fördern nicht nur Ihre berufliche Entwicklung, sondern auch eine ausgewogene Work-Life-Balance./p