Auf einen Blick
- Aufgaben: Leite die Zuverlässigkeit und Resilienz unserer Cloud-Plattform und Produktionsdienste.
- Unternehmen: Innovatives Unternehmen im Bereich Plattformengineering mit globalem Einfluss.
- Vorteile: Attraktives Gehalt, flexible Arbeitszeiten und Möglichkeiten zur beruflichen Weiterentwicklung.
- Weitere Informationen: Dynamisches Umfeld mit hervorragenden Karrierechancen und einem starken Fokus auf Teamarbeit.
- Warum dieser Job: Gestalte die Zukunft der Cloud-Technologien und mache einen echten Unterschied.
- Qualifikationen: Über 15 Jahre Erfahrung in der Produktionstechnik und Führungskompetenz.
Das prognostizierte Gehalt liegt zwischen 60000 - 80000 € pro Jahr.
The Platform Engineering team builds, secures and operates scalable infrastructure supporting cloud-managed Saa S products with on-premises components deployed at customer sites.
As our Principal Site Reliability Engineer, you will provide the technical leadership for reliability across our global platform.
You will define the strategy, standards and operating model that ensure highly available, resilient and secure services, while remaining hands‑on with the design and operation of the technologies that underpin our production environments.
Working closely with Architecture, Dev Sec Ops, Cloud Operations and Product Development, you will drive a culture of observability, automation and continuous improvement, ensuring issues are identified and resolved before they impact customers.
- Define and lead the reliability strategy for Tier 1 and Tier 2 production services, including service-level objectives, error budgets, observability, resilience, disaster recovery and cloud security.
- Design and operate observability platforms, lead chaos engineering and disaster recovery initiatives, improve the reliability of Postgre SQL, Redis/Valkey, Kafka and Open Search, and deliver AI Ops capabilities including predictive monitoring, automated remediation and self‑healing.
- Lead major incident response, on‑call operations and global escalation, while mentoring engineers and establishing engineering standards adopted across the organisation.
- What you’ll be doing
You will own the reliability, resilience and operational excellence of our cloud platform and production services, embedding reliability principles into architectural decisions and platform design.
You'll establish observability standards covering metrics, logs, distributed tracing and profiling, while governing service-level objectives and error budgets across engineering teams.
You will lead resilience engineering through chaos testing, disaster recovery planning and validated failover exercises, co‑own cloud security posture with Dev Sec Ops, and ensure the reliability of our critical data and streaming platforms.
Working across multiple engineering disciplines, you will also optimise platform efficiency through automation, AI‑driven operations and continuous operational improvement.
- Own production reliability, observability, service-level objectives, error budgets, resilience, backup and disaster recovery, cloud security and platform governance across all critical services.
- Lead incident command, executive communications, post‑incident reviews, on‑call operations and global escalation, while delivering predictive monitoring, automated remediation and self‑healing capabilities.
- Partner with Architecture, Dev Sec Ops, Cloud Operations and Product Development to establish engineering standards, mentor engineers and deliver a shared reliability roadmap.
- What you’ll bring
You are an experienced Site Reliability Engineering leader with more than 15 years of production engineering experience and recent hands‑on responsibility for large‑scale, fault‑tolerant production systems running on AWS or GCP.
You have successfully designed and operated observability platforms, implemented service-level objective programmes and delivered measurable improvements in availability, reliability and mean time to recovery.
You have led high‑severity production incidents, planned and executed resilience testing and disaster recovery exercises, and influenced engineering practices across multiple teams through technical leadership and mentoring.
- 15+ years of production engineering experience with recent hands‑on responsibility for large‑scale, fault‑tolerant production systems on AWS or GCP, including ownership of observability, service-level objectives and error budgets.
- Proven experience leading resilience initiatives, validated disaster recovery and failover testing, together with command of high‑severity production incidents and measurable reliability improvements.
- Demonstrated technical leadership through architecture reviews, governance, mentoring, coaching and engineering standards adopted across multiple teams and services.
- You’ll be a great fit with
You have deep expertise across cloud infrastructure, platform engineering and cloud security, with practical experience securing distributed production environments through cloud security posture management, runtime vulnerability detection and automated policy enforcement.
You understand how to balance reliability, security, networking and operational efficiency at scale.
You have experience implementing AI‑powered operational capabilities, applying Fin Ops principles to optimise infrastructure, and designing highly available distributed systems that deliver exceptional resilience and performance.
- Proven experience with cloud security posture management, workload protection, policy‑as‑code, secure‑by‑default infrastructure and automated security enforcement across cloud‑native platforms and CI/CD pipelines.
- Hands‑on expertise with AI traffic management through LLM gateways, capacity planning, resource rightsising, Fin Ops‑based optimisation, autonomous operations, predictive alerting and self‑healing platforms.
- Deep knowledge of networking, routing, load balancing, connectivity resilience and distributed systems, with the ability to integrate networking, security and reliability into scalable platform architectures.
- #J-18808-Ljbffr
Principal Site Reliability Engineer Arbeitgeber: IonQ
Als Principal Site Reliability Engineer in unserem engagierten Platform Engineering Team haben Sie die Möglichkeit, eine Schlüsselrolle in der Gestaltung und Sicherstellung der Zuverlässigkeit unserer globalen Cloud-Plattform zu übernehmen. Wir bieten ein dynamisches Arbeitsumfeld, das von einer Kultur der kontinuierlichen Verbesserung und Zusammenarbeit geprägt ist, sowie umfangreiche Möglichkeiten zur beruflichen Weiterentwicklung und zum Mentoring. Unsere Mitarbeiter profitieren von flexiblen Arbeitszeiten, einem starken Fokus auf Work-Life-Balance und innovativen Projekten, die es Ihnen ermöglichen, Ihre technischen Fähigkeiten in einem unterstützenden und zukunftsorientierten Umfeld weiterzuentwickeln.