Senior Site Reliability Engineer (m/f/d)

Senior Site Reliability Engineer (m/f/d)

Kaiserslautern Vollzeit 63000 - 77000 € / Jahr (geschätzt) Homeoffice (teilweise)
TOPdesk

Auf einen Blick

  • Aufgaben: Übernehme die Zuverlässigkeit unserer Azure SaaS-Plattform und entwickle innovative Automatisierungslösungen.
  • Unternehmen: TOPdesk, ein führendes Unternehmen im Bereich Service-Management-Software mit globaler Präsenz.
  • Vorteile: Unbefristeter Vertrag, 30 Tage Urlaub, flexible Arbeitszeiten und Homeoffice-Möglichkeiten.
  • Weitere Informationen: Offene Unternehmenskultur mit flachen Hierarchien und tollen Teamevents.
  • Warum dieser Job: Gestalte die Zukunft der ITSM mit KI und verbessere die Zuverlässigkeit für Millionen von Nutzern.
  • Qualifikationen: Mindestens 5 Jahre Erfahrung in Site Reliability oder DevOps, starke Kenntnisse in Kubernetes und Automatisierung.

Das prognostizierte Gehalt liegt zwischen 63000 - 77000 € pro Jahr.

Company Description

TOPdesk builds service management software used across education, healthcare, government, and manufacturing.

We are 700+ colleagues in 8 offices worldwide.

Founded over 30 years ago, we serve more than 10 million users worldwide and have been helping organisations deliver better services ever since.

We are an open, collaborative organisation with little hierarchy — people own their work end to end and are trusted to make the decisions that matter.

We are reinventing ITSM and ESM for the agentic era, building AI agents our customers can trust, and this role is part of that.

Job Description

About the role

Our Azure Saa S estate keeps service management running for thousands of organisations worldwide, under SLA-backed 24/7 availability.

As a Senior Site Reliability Engineer, you own the reliability of that estate as an engineering problem — you set the SLOs, engineer out the toil behind them, and make the platform faster to change and cheaper to operate without trading away resilience.

You sit in the Saa S infrastructure function, working alongside cloud engineering and the product squads shipping to production.

You bring our AI-native ways of working into reliability: agents and bounded automation with observability, approvals, containment, and rollback — self-healing systems, not runbooks worked by hand.

What this is not A ticket-driven, break-fix ops role kept away from the code.

This is reliability as engineering — you own SLOs and error budgets, automate what you repeat, and design the platform to recover itself rather than reacting incident by incident.

The team We are a group of social technicians who value transparency, open feedback, and a healthy work-life balance — and who treat reliability as a shared, measurable objective, not a firefight.

  • What you'll own
  • SLOs and error budgets.

Define and own service-level objectives across the Azure (and potentially multi-cloud) estate, and use error budgets to steer the balance between shipping change and protecting reliability.

  • Toil elimination and self-healing automation.

Identify toil, classify it, and engineer it out — feeding self-healing automation and your findings into the reliability roadmap.

Standupanagent-basedsupportlayerthatownsrecurringtoilandcontinuouslyfeedsimprovementsbackintoreliability.

  • Observability consolidation.

Standardise metrics, alerting, and tracing across all datacenters, close coverage gaps on cloud workloads, and measurably reduce the alert-to-incident ratio from baseline.

  • Incident response and blameless postmortems. Lead incidents to resolution, run blameless postmortems, and turn every learning into a durable fix or an automation candidate.
  • Reliability of releases. Harden CI/CD and progressive delivery — canaries, safe rollouts, automated rollback — so change velocity and reliability rise together.
  • Capacity and performance. Model capacity, load-test critical paths, and keep the platform within its performance envelope as it scales across regions.
  • AI-native reliability. Bring agents and bounded automation — with observability, approvals, containment, and rollback — into detection, diagnosis, and remediation.
  • Runbooks that get used. Every alert links to a runbook; every runbook links to an automation candidate. You leave things more legible than you found them.
  • Capacity and cost forecasting.

Own capacity and cost planning across the multi-cloud estate, model usage and growth trends, and forecast short and long term infrastructure needs so spend and scaling decisions stay ahead of demand rather than reacting to it.

  • How you approach the work
  • Automate what you repeat — if you have done it manually twice, the third time is a design problem.
  • Measure before optimising: SLOs, baselines, and dashboards before opinions.
  • Design for failure — assume things break, and make recovery automatic and observable.
  • Consultative, not gatekeeping: you pair with product engineering teams and transfer knowledge as you go.
  • Treat cost and reliability as joint objectives, not a forced trade-off.
  • Pro-active collaboration with product teams.

You are involved in the early phases of product development, including design to help the teams make optimal choices and timely introduce appropriate SRE practices.

  • Technical environment
  • Scale: 10+ global datacenters; SLA-backed, 24/7 multi-tenant Saa S serving millions of end users.
  • Cloud: Azure across all production regions, with a mature landing-zone and networking architecture.
  • Compute: Kubernetes / Azure AKS alongside traditional VM infrastructure, all managed as code.
  • Infrastructure as code: Terraform via CI/CD and Git Ops workflows; configuration management with Puppet and Ansible across Linux and Windows.
  • Observability: metrics, alerting, and tracing across cloud-native and self-managed layers (e. g. Grafana, Prometheus, Victoria Metrics, Influx).
  • Automation: Python and automation tooling — and we expect you to take the reliability stack to the next level, not just operate today's.
  • Legacy: Java, MS SQL, heritage architecture — being decomposed. The SRE role is not responsible for the Java application code.
  • How we build: Claude Code as our primary AI-native SDLC tool; subagents and multi-agent workflows; MCP tool integrations; shared prompt, agent, and eval libraries.
  • Success in your first year
  • SLOs and error budgets are defined for the estate's critical services and actively used to steer delivery decisions.
  • The alert-to-incident ratio is measurably down, and runbooks you wrote are used by the on-call shift without escalation.
  • Toil you identified is automated — or has a credible, documented roadmap to be — and self-healing covers at least one high-frequency failure mode.
  • Postmortems produce durable fixes, not repeat incidents; recurring-incident rate is trending down against a documented baseline.
  • Product squads consult you during design, not only after incidents.

Qualifications

  • Required
  • Proven hands-on experience (5+ years) as a Site Reliability, Dev Ops, or Infrastructure Engineer running a production cloud environment at scale (Azure).
  • Fluent with SLOs, error budgets, and reliability engineering practice — you have set them, not just read about them.
  • Strong observability skills at scale — Grafana, Prometheus, Victoria Metrics, or equivalent — including alerting and tracing.
  • Kubernetes at operator level: Helm, namespace management, ingress controllers, RBAC, persistent volumes.
  • Coding for automation (Python or equivalent) and Terraform delivered via CI/CD.
  • Linux system administration — you understand what Puppet or Ansible is doing, not just whether it ran green.
  • Comfortable leading incidents in an on-call rotation with real SLA obligations, and the maturity to know when to elevate.
  • Strong written communication — your postmortems, runbooks, and architecture notes are unambiguous.
  • Nice to have
  • Experience with progressive delivery — canaries, feature flags, automated rollback.
  • Experience working within or migrating toward an Azure Cloud Adoption Framework or enterprise landing-zone structure.
  • Current, personal practice of AI-native software delivery (Claude Code or equivalent).
  • Experience with EU data residency / sovereign cloud requirements.
  • Additional Information
  • What's in it for you
  • Permanent employment contract and 30 days of annual vacation
  • Pleasant working atmosphere with flat hierarchies
  • Open working atmosphere in an international environment
  • Flexible working hours within a modern working environment
  • Possibility to work remotely
  • Well-founded onboarding by a buddy
  • Time for individual training opportunities to further develop your personal strengths
  • Joint employee events and team building measures
  • Employee subsidy for gym membership
  • Company health measures such as health days or fresh fruit
  • Free drinks (coffee, tea, water)
  • Gifts on special occasions, e. g. anniversary
  • Monthly tax-free payment in the form of a Mastercard
  • Possibility of time off (sabbatical)
  • Quality time together (table football, table tennis table, massage chair)
  • Corporate benefits
  • Vacation bonus

We welcome applications from all interested parties, regardless of their ethnic and social background, age, religion, gender, disability, sexual orientation, or identity.

#J-18808-Ljbffr

Senior Site Reliability Engineer (m/f/d) Arbeitgeber: TOPdesk

TOPdesk ist ein hervorragender Arbeitgeber, der seinen Mitarbeitern die Möglichkeit bietet, in einem dynamischen und internationalen Umfeld zu arbeiten. Mit flexiblen Arbeitszeiten, der Option auf bis zu 50% Homeoffice und einem umfassenden Onboarding-Programm fördert das Unternehmen eine offene und kollaborative Kultur, die individuelles Wachstum und Weiterbildung unterstützt. Die Mitarbeiter profitieren von attraktiven Zusatzleistungen wie einem Urlaubsgeld, speziellen Zahlungen und einer Vielzahl von Gesundheits- und Wellnessangeboten, die das Wohlbefinden am Arbeitsplatz fördern.

TOPdesk

Kontaktdaten:

TOPdesk Recruiting-Team

StudySmarter Expertenrat🤫

Wir sind der Meinung, dass du so Senior Site Reliability Engineer (m/f/d) erhalten könntest

Netzwerken in der IT-Community

In der IT-Consulting-Welt sollten wir regelmäßig auf Veranstaltungen wie Tech-Meetups oder Konferenzen gehen. Hier können wir nicht nur unser Netzwerk erweitern, sondern auch direkt mit potenziellen Arbeitgebern ins Gespräch kommen und unser Interesse an einer Vollzeitstelle zeigen.

Online-Foren und Gruppen nutzen

Sich in Online-Foren und Communities wie Stack Overflow oder LinkedIn-Gruppen umzusehen, kann uns helfen, Insider-Tipps zu erhalten und Informationen über offene Stellen in der IT-Beratung zu sammeln. Vergiss nicht, aktiv zu werden und Fragen zu stellen oder dein Wissen zu teilen – das erhöht unsere Sichtbarkeit!

Direkt bei TOPdesk bewerben

Viele Unternehmen, wie TOPdesk, stemmen ihre Vollzeitstellen bevorzugt über ihre eigenen Karriere-Webseiten. Also, lass uns regelmäßig auf deren Seite vorbeischauen und uns direkt bewerben, statt nur die üblichen Jobportale zu nutzen.

Überzeugende Projekte zeigen

Wir sollten unser Portfolio oder relevante Projekte gut sichtbar machen, egal ob das auf Github, persönlich oder auf LinkedIn ist. Bei IT-Consulting-Stellen kommt es oft auf praktische Erfahrungen an, also lass uns zeigen, was wir können!

Wir glauben, dass du diese Fähigkeiten brauchst, um Senior Site Reliability Engineer (m/f/d) mit Bravour zu bestehen

SLOs
Fehlerbudgets
Zuverlässigkeitsengineering
Observability (Grafana, Prometheus, VictoriaMetrics)
Kubernetes (Operator-Level)
Automatisierung (Python oder gleichwertig)
Terraform

Einige Tipps für deine Bewerbung 🫡

Zeige deine technischen Skills!:In der IT-Beratung zählen deine technischen Kenntnisse und Fähigkeiten. Achte darauf, relevante Programmiersprachen, Tools und Systeme in deinem Lebenslauf aufzulisten. Zeig auch, wenn du Zertifikate hast, die deine Kompetenz unterstützen – das könnte dir einen echten Vorteil verschaffen!

Verstehe die Branche!:Unterstreiche in deinem Anschreiben, dass du ein gutes Verständnis für aktuelle Trends und Herausforderungen in der IT-Branche hast. Zeig, dass du nicht nur die technischen Aspekte beherrschst, sondern auch die Bedürfnisse der Kunden erkennen und lösen kannst!

Deine Projekte zählen!:Falls du bereits an IT-Projekten gearbeitet hast, verlinke diese oder beschreibe sie in deinem Lebenslauf. Praktische Erfahrungen – sei es in Form von Praktika oder privaten Projekten – sind besonders wertvoll in der IT-Beratung. Zeige uns, was du kannst!

Individuelle Bewerbung ist der Schlüssel!:Jede Bewerbung sollte individuell auf TOPdesk und die ausgeschriebene Position Senior Site Reliability Engineer (m/f/d) zugeschnitten sein. Teile uns mit, warum gerade du eine gute Wahl für unser Team bist. Das zeigt dein Engagement und deine Motivation, die über eine Standardbewerbung hinausgeht.

Wie man sich auf ein Vorstellungsgespräch bei TOPdesk vorbereitet

Technische Vorbereitung ist alles!

Da du dich auf eine Vollzeitstelle in der IT-Beratung bewirbst, solltest du dir wirklich einen Überblick über die wichtigsten Tools und Technologien verschaffen, die in der Branche verwendet werden. Sei bereit, technische Fragen zu beantworten, die sich auf Software-Architektur oder Systemintegration beziehen könnten.

Praxisbeispiele parat haben

In der IT-Beratung ist es wichtig, konkrete Beispiele aus deiner bisherigen Erfahrung zu bringen. Überlege dir Projekte, bei denen du erfolgreich einen Kunden beraten hast oder Herausforderungen gelöst hast. Das zeigt, dass du nicht nur theoretisches Wissen hast, sondern auch in der Praxis erfolgreich sein kannst.

Soft Skills betonen

Ein großer Teil der IT-Beratung ist die Kommunikation mit Kunden und das Verständnis ihrer Bedürfnisse. Bereite dich darauf vor, über deine zwischenmenschlichen Fähigkeiten zu sprechen, wie du mit herausfordernden Kunden umgehst oder wie du in Teams arbeitest. Das wird den Interviewern zeigen, dass du mehr als nur technisches Wissen mitbringst!

Fragen zum Unternehmen vorbereiten

Schau dir spezifisch die Projekte von TOPdesk an und überlege dir, welche Fragen du dazu stellen möchtest. Zeig Interesse an den aktuellen Herausforderungen, vor denen das Unternehmen steht, und wie du dazu beitragen könntest. Das hebt dich von anderen Bewerbern ab und zeigt, dass du wirklich motiviert bist.