- Our engineering team builds next-generation, cloud-native bare-metal orchestration platforms for AI neo-cloud environments. We operate on a T-shaped engineering model: while each engineer brings deep expertise in a primary domain, every team member maintains practical fluency across neighboring systems
- This shared foundation enables rigorous design reviews, effective cross-domain code reviews, and dependable on-call coverage. Our development is fundamentally based on open-source projects, and contributing to them is a major part of our work
- We develop primary control plane software in Rust and Go, leverage Kubernetes operators, and foster an OpenTelemetry-first operational culture
- We are seeking a Senior Software Engineer who combines strong software architecture principles with understanding of Datacenter Networking and DPU architectures. Fitting into our T-shaped engineering model, you will drive the development of control planes powering high-density, high-throughput infrastructure for AI training and inference workloads
- Control Plane Development: Design, build, and maintain production-grade control plane microservices, custom Kubernetes operators, and robust reconciliation engines in Rust and Go
- Architecture & API Design: Author clear Functional (FR) and Non-Functional Requirements (NFR), architectural specs, system sequence diagrams, and clean gRPC/Protobuf and REST API schemas
- Network Integration: Develop custom integration modules and integrations for modern network operating systems (SONiC, NVUE, Cumulus) and DPU hardware platforms
- Observability & Diagnostics: Instrument services end-to-end using OpenTelemetry (OTel) traces and metrics, performing systematic root-cause analysis across polyglot distributed systems
- Quality & Engineering Excellence: Participate in rigorous, review-gated pull request workflows across Rust, Go, and SQL codebases while ensuring high test coverage via mock-driven testing
Familiarity with automated bare-metal provisioning and lifecycle management is a strong assetKubernetes Control Planes: Experience writing custom controllers/operators using kube-rs with CRD-driven reconciliation loop patternsPrimary mastery of Rust (Tokio async runtime, Tonic, Axum, sqlx) and proficiency in Go (secondary)Testing Discipline: Commitment to test-driven design, writing highly testable code with mock interfaces for external service dependenciesCollaborative Architecture: Proven track record of writing crisp architectural specs and engaging in constructive, cross-functional design reviewsConcurrency & Reliability: Strong background in async concurrency models, lock-free patterns, distributed state handling, and mock-driven testing disciplineBachelor’s or Master’s degree in Computer Science, Computer Engineering, or equivalent practical experience5+ years of software engineering experience in cloud-native platforms or systems programmingAPIs & Serialization: gRPC/Protobuf contract design (via tonic/prost) and schema evolution (using zero-copy serialization and compile-time type generation)DPU & Fabric Ecosystem: Deep knowledge of NVIDIA DOCA, Host-Based Networking (HBN), BlueField DPU architectures, and the DPF (DOCA Platform Framework) operator model (BFB, DPUSet, DPUNode CRDs)Datacenter Protocols & OS: Expertise in BGP, MP-BGP, EVPN, VXLAN, L3VNI, route targets, and route server design. Hands-on experience with SONiC, Cumulus Linux, and NVUE (featuring a first-class NVUE client)Advanced SQL / PostgreSQL fluencyHigh-Performance Interconnects: Understanding of InfiniBand fabrics, NVLink / NMX-M partitioning, and RoCEv2Linux Kernel Networking: Advanced grasp of Linux netlink, network namespaces, routing tables, and internal DHCP/DNS service implementations (the project ships its own services)Cloud & Kubernetes Networking: Familiarity with CNI plugins (Calico, Cilium). OVS and DPDK experience is a plusHardware Management: Proficiency in Redfish, IPMI, and BMC abstractions across heterogeneous hardware platforms (Dell, Lenovo, NVIDIA reference hardware)Provisioning & Boot Infrastructure: Expertise in PXE/iPXE, UEFI, Secure Boot, measured boot, and TPM attestationIdentity & Security: Practical experience with PKI, X.509 certificates, TLS, SPIFFE/SVID, Vault, KMS, Keycloak (OAuth2/JWT), and RBACLinux Systems Internals: In-depth knowledge of boot chains, systemd, initramfs, BIOS configuration matricesVirtualization: Experience with KubeVirt / KVMKubernetes Control Planes: Experience writing custom controllers/operators using kube-rs or controller-runtime with CRD-driven reconciliation loop patternsDemonstrated participation or maintaining open-source systems projects
#J-18808-Ljbffr
Senior Software Engineer (Rust) in Berlin Arbeitgeber: Mirantis
Mirantis ist ein hervorragender Arbeitgeber, der Ihnen die Möglichkeit bietet, in einem dynamischen und innovativen Umfeld zu arbeiten, das sich auf fortschrittliche KI-Infrastrukturen konzentriert. Mit einem starken Fokus auf Mitarbeiterentwicklung und Teamarbeit fördern wir eine Kultur des Lernens und der Zusammenarbeit, während Sie an spannenden Projekten mit modernster Technologie wie NVIDIA GPUs und Kubernetes arbeiten. Unsere engagierte Infrastruktur- und Engineering-Abteilung bietet Ihnen nicht nur die Chance, Ihre Fähigkeiten im Bereich HPC-Netzwerke zu vertiefen, sondern auch einen klaren Karriereweg zum Senior HPC Network Engineer.