Auf einen Blick
- Aufgaben: Verantworte die Zuverlässigkeit unserer Azure SaaS-Plattform und automatisiere Prozesse.
- Unternehmen: Innovatives Unternehmen mit einem Fokus auf Cloud-Technologien und Teamarbeit.
- Vorteile: Attraktives Gehalt, flexible Arbeitszeiten und Möglichkeiten zur beruflichen Weiterentwicklung.
- Weitere Informationen: Wachstumsorientiertes Team mit einer Kultur der Transparenz und des offenen Feedbacks.
- Warum dieser Job: Gestalte die Zukunft der Zuverlässigkeit in einer dynamischen Umgebung mit modernster Technologie.
- Qualifikationen: Erfahrung in Site Reliability Engineering und Automatisierung von Prozessen.
Das prognostizierte Gehalt liegt zwischen 63000 - 77000 € pro Jahr.
About the role
Our Azure Saa S estate keeps service management running for thousands of organisations worldwide, under SLA-backed 24/7 availability.
As a Senior Site Reliability Engineer, you own the reliability of that estate as an engineering problem - you set the SLOs, engineer out the toil behind them, and make the platform faster to change and cheaper to operate without trading away resilience.
You sit in the Saa S infrastructure function, working alongside cloud engineering and the product squads shipping to production.
You bring our AI-native ways of working into reliability: agents and bounded automation with observability, approvals, containment, and rollback - self-healing systems, not runbooks worked by hand.
What this is not A ticket-driven, break-fix ops role kept away from the code.
This is reliability as engineering - you own SLOs and error budgets, automate what you repeat, and design the platform to recover itself rather than reacting incident by incident.
The team
We are a group of social technicians who value transparency, open feedback, and a healthy work-life balance - and who treat reliability as a shared, measurable objective, not a firefight.
- What you'll own
- SLOs and error budgets.
Define and own service-level objectives across the Azure (and potentially multi-cloud) estate, and use error budgets to steer the balance between shipping change and protecting reliability.
- Toil elimination and self-healing automation.
Identify toil, classify it, and engineer it out - feeding self-healing automation and your findings into the reliability roadmap.
Stand up an agent-based support layer that owns recurring toil and continuously feeds improvements back into reliability.
- Observability consolidation.
Standardise metrics, alerting, and tracing across all datacenters, close coverage gaps on cloud workloads, and measurably reduce the alert-to-incident ratio from baseline.
- Incident response and blameless postmortems. Lead incidents to resolution, run blameless postmortems, and turn every learning into a durable fix or an automation candidate.
- Reliability of releases. Harden CI/CD and progressive delivery - canaries, safe rollouts, automated rollback - so change velocity and reliability rise together.
- Capacity and performance. Model capacity, load-test critical paths, and keep the platform within its performance envelope as it scales across regions.
- AI-native reliability. Bring agents and bounded automation - with observability, approvals, containment, and rollback - into detection, diagnosis, and remediation.
- Runbooks that get used. Every alert links to a runbook; every runbook links to an automation candidate. You leave things more legible than you found them.
- Capacity and cost forecasting.
Own capacity and cost planning across the multi-cloud estate, model usage and growth trends, and forecast short and long term infrastructure needs so spend and scaling decisions stay ahead of demand rather than reacting to it.
- How you approach the work
- Automate what you repeat - if you have done it manually twice, the third time is a design problem.
- Measure before optimising: SLOs, baselines, and dashboards before opinions.
- Design for failure - assume things break, and make recovery automatic and observable.
- Consultative, not gatekeeping: you pair with product engineering teams and transfer knowledge as you go.
- Treat cost and reliability as joint objectives, not a forced trade-off.
- Pro-active collaboration with product teams.
You are involved in the early phases of product development, including design to help the teams make optimal choices and timely introduce appropriate SRE practices.
Technical environment
Scale: 10+ global datacenters; SLA-backed, 24/7 multi-tenant Saa S serving millions of end users.
Cloud: Azure across all production regions, with a mature landing-zone and networking architecture.
Compute: Kubernetes / Azure AKS alongside traditional VM infrastructure, all managed as code.
Infrastructure as code: Terraform via CI/CD and Git Ops workflows; configuration management with Puppet and Ansible across Linux and Windows.
Observability: metrics, alerting, and tracing across cloud-native and self-managed layers (e. g. Grafana, Prometheus, Victoria Metrics, Influx).
Automation: Python and automation tooling - and we expect you to take the reliability stack to the next level, not just operate today
Legacy: Java, MS SQL, heritage architecture - being decomposed. The SRE role is not responsible for the Java application code.
How we build: Claude Code as our primary AI-native SDLC tool; subagents and multi-agent workflows; MCP tool integrations; shared prompt, agent, and eval libraries.
Success in your first year
SLOs and error budgets are defined for the estate's critical services and actively used to steer delivery decisions.
The alert-to-incident ratio is measurably down, and runbooks you wrote are used by the on-call shift without escalation.
- Toil you identified is automated - or has a credible, document
- #J-18808-Ljbffr
Senior Site Reliability Engineer (m/f/d) Arbeitgeber: PVH (Tommy Hilfiger/Calvin Klein)
Lyreco Coffee Solutions bietet eine herausragende Arbeitsumgebung für den Corporate Account Sales Manager in Dietikon, wo ein motiviertes Team und eine menschenzentrierte Unternehmenskultur auf dich warten. Mit viel Gestaltungsspielraum und einer strategischen Führungsposition hast du die Möglichkeit, deine Karriere aktiv zu gestalten und von attraktiven Sozialleistungen sowie umfangreichen Weiterentwicklungsmöglichkeiten zu profitieren. Hier wird Eigenverantwortung großgeschrieben, und du kannst deine Erfolge durch unser Very Lyreco People-Programm anerkannt sehen.
Kontaktdaten:
PVH (Tommy Hilfiger/Calvin Klein) Recruiting-Team
StudySmarter Expertenrat🤫
Wir sind der Meinung, dass du so Senior Site Reliability Engineer (m/f/d) erhalten könntest
✨Engagier dich in Entwickler-Communities!
Lass uns mal ehrlich sein: In der Software-Entwicklung sind Netzwerke Gold wert! Tummel dich in GitHub-Projekten, nehme an lokalen Meetups oder Hackathons teil und vernetze dich mit anderen Entwicklern. So steigerst du nicht nur deine Sichtbarkeit, sondern lernst auch die neuesten Trends und Technologien kennen.
✨Zeig deine Fähigkeiten!
Erstelle ein Portfolio, das deine besten Projekte und Code-Examples zeigt. Nichts überzeugt mehr als ein praktischer Beweis deiner Skills. Das kann auch helfen, bei PVH (Tommy Hilfiger/Calvin Klein) anzuklopfen, wenn du dich auf die Stelle als Senior Site Reliability Engineer (m/f/d) bewirbst – so wissen sie gleich, was sie von dir erwarten können!
✨Nutze Jobplattformen speziell für Tech-Jobs!
Plattformen wie Stack Overflow Jobs oder AngelsList sind perfekte Orte, um Vollzeitstellen in der Software-Entwicklung zu finden. Hier sind viele tolle Unternehmen auf der Suche nach Talenten wie uns, also schau regelmäßig vorbei und bewirb dich direkt über die Website.
✨Such dir Mentoren und Feedback!
Hol dir Feedback von erfahrenen Entwicklern, die dir Tipps geben können, was Recruiter wirklich suchen. Ob über LinkedIn oder persönliche Kontakte: Menschen, die sich in der Branche auskennen, können enorm wertvoll sein, um dir zu helfen, dich optimal auf deine Bewerbung bei PVH (Tommy Hilfiger/Calvin Klein) vorzubereiten!
Wir glauben, dass du diese Fähigkeiten brauchst, um Senior Site Reliability Engineer (m/f/d) mit Bravour zu bestehen
Einige Tipps für deine Bewerbung 🫡
Highlights deiner Coding-Skills:In der Software-Entwicklung kommt es auf konkrete Fähigkeiten an. Vergiss nicht, relevante Programmiersprachen und Frameworks in deinen Lebenslauf aufzunehmen. Zeig uns, was du kannst – vielleicht mit einem Link zu deinem GitHub-Profil oder einer Übersicht deiner Side Projects, die deine Programmierkenntnisse illustrieren.
Dokumentation deiner Erfolge:Gerade bei einer Vollzeitstelle in der Software-Entwicklung sind konkrete Ergebnisse Gold wert. Nenn uns Zahlen und Ergebnisse aus deinen vorherigen Projekten. Hast du den Code optimiert oder Systemfehler behoben? Solche Erfolge zeigen, dass du die Sprache der Entwickler sprichst und einen echten Mehrwert bringst.
Attraktive Projektbeschreibungen:Wenn du an Projekten gearbeitet hast, die hervorstechen, beschreibe sie ausführlich in deinem Lebenslauf. Was war das Problem, das du gelöst hast? Welche Technologien hast du eingesetzt? Das gibt uns einen klaren Einblick in deine Herangehensweise und Problemlösungsfähigkeiten.
Motivation zeigen:In deinem Anschreiben solltest du deine Motivation für die Stelle im Bereich Software-Entwicklung bei PVH (Tommy Hilfiger/Calvin Klein) klar herausstellen. Warum sprichst gerade du die Anforderungen für diese Vollzeitrolle an? Mach deutlich, was dich an der Arbeit bei uns reizt und wie du über das rein Technische hinaus wachsen möchtest.
Wie man sich auf ein Vorstellungsgespräch bei PVH (Tommy Hilfiger/Calvin Klein) vorbereitet
✨Technische Vorbereitung auf die Coding-Challenges
In der Software-Entwicklung sind technische Fragen oft ein zentraler Teil des Interviews. Macht euch mit Plattformen wie LeetCode oder HackerRank vertraut, um eure Problemlösungsfähigkeiten zu trainieren. Zeigt im Interview viel Selbstbewusstsein beim Erklären eurer Ansätze!
✨Das eigene Portfolio im besten Licht präsentieren
Stellt sicher, dass ihr ein aussagekräftiges Portfolio habt, das einige eurer besten Projekte zeigt. Seid bereit, darüber zu sprechen, was eure Rolle war, welche Technologien ihr verwendet habt und welche Herausforderungen es gab. Das gibt den Interviewern einen Einblick in eure praktische Erfahrung.
✨Teamfähigkeit und Kommunikation betonen
In einer Vollzeit-Position wird Kommunikation im Team sehr wichtig sein. Seid bereit, Beispiele aus der Vergangenheit zu teilen, in denen ihr effektiv im Team gearbeitet habt. Dies zeigt, dass ihr nicht nur technische Fähigkeiten habt, sondern auch gut ins Team passt.
✨Vorbereitung auf Fragen zur Software-Architektur
Bereitet euch darauf vor, Fragen zur Software-Architektur zu beantworten. Themen wie RESTful APIs, Microservices und Cloud-Architekturen können Teil eures Interviews sein. Zeigt euer Verständnis durch Diskussionen und Beispiele aus eurer bisherigen Arbeit oder Projekte.