Senior Site Reliability Engineer

Senior Site Reliability Engineer

Berlin Vollzeit 49500 - 60500 € / Jahr (geschätzt) Homeoffice (teilweise)
C

Auf einen Blick

  • Aufgaben: Sichere und optimiere unsere Plattform für KI-Anwendungen und sorge für reibungslose Abläufe.
  • Unternehmen: CloudFactory, ein mission-driven Unternehmen mit globaler Gemeinschaft.
  • Vorteile: Wettbewerbsfähiges Gehalt, Weiterbildungsmöglichkeiten und eine sinnvolle Arbeit.
  • Weitere Informationen: Dynamisches Umfeld mit globaler Zusammenarbeit und hervorragenden Entwicklungsmöglichkeiten.
  • Warum dieser Job: Gestalte die Zukunft der KI und arbeite an innovativen Projekten mit echtem Einfluss.
  • Qualifikationen: Mindestens 5 Jahre Erfahrung in Infrastrukturengineering oder SRE, Kenntnisse in Kubernetes.

Das prognostizierte Gehalt liegt zwischen 49500 - 60500 € pro Jahr.

At Cloud Factory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world.

By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale.

More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives.

Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following.

Our Culture

At Cloud Factory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:

  • Mission-Driven: We focus on creating economic and social impact.
  • People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging.
  • Innovative: We embrace change and find better ways to do things together.
  • Globally Connected: We foster collaboration between diverse cultures and perspectives.

If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!

Role Summary

As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly.

You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security.

The SRE team owns the foundation of AI Platform’s Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers and that powers many ML and LLM powered features.

We give every other engineering team a reliable base to build on, and we own the software delivery lifecycle end to end: the tooling, patterns, and automation that reduce friction for the whole org.

This is an exciting opportunity to grow professionally while contributing to a mission-driven organization.

Responsibilities

  • What you’ll own
  • Reliability of platform(includes ML and LLM workloads) - model serving and inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services
  • Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks we run everywhere else
  • Company-wide technical direction: shaping the roadmap and building golden paths that raise the baseline for every team
  • Developer tooling and automation that compounds - reusable Git Hub Actions, Git Ops workflows, Terraform modules - so every engineer ships faster
  • Reusable components packaging common open-source tools (Grafana, Istio, Cloud Native stack, and ML tooling such as model registries and feature stores) for teams to deploy in any environment
  • Secure-by-default infrastructure - baking security, compliance audits, cost governance, and audit trails into the platform in close partnership with our lead/backend/staff engineers.

Who you are (must-haves)

  • 5+ years in infrastructure engineering, Dev Ops, or SRE, operating large-scale, high-availability production systems using Kubernetes
  • Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or Cloudformation, on at least one major cloud (AWS preferred).
  • Good proficiency in Python or Go or general scripting for automation and tooling(automation with higher language preferred)
  • AI is already in your daily loop - Agentic tooling (Claude Code, Codex, Droid, internal skills) is part of how you ship and not what you are experimenting with.

We believe AI tools can be great with human judgement and we want the SRE team to bring the next wave day to day operations.

  • First-principles reasoning - Reasoning from constraints and failure modes naming the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off)
  • At least one infrastructure build you owned end to end - with the outcome metric attached (deploy time, MTTR, cost, adoption, availability).
  • Cross-functional strength. Track record working with product, backend/frontend teams to pull through collective initiative.
  • ML & AI platform (strongly preferred)
  • Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management
  • Model serving and inference at production scale (eg KServe, Ray Serve, Triton, v LLM, or similar) with real latency and cost constraints(preferred Ray Serve)
  • MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents)
  • LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request)
  • Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails
  • Any other General requirements
  • Global Collaboration: Ability to work across global teams and different cultures across various time zones with strong communication skills.
  • Problem Solving: Ability to break down complex problems into simple, actionable solutions.
  • Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
  • Availability: Willingness to support processes for 24x7 operational support.

At Cloud Factory, we believe that work should be more than just a job—it should be a platform for growth, impact, and community.

Here, you’ll earn with purpose, learn every day, and serve a mission that truly matters.

If you’re looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we’d love to have you on this journey!

#J-18808-Ljbffr

Senior Site Reliability Engineer Arbeitgeber: CloudFactory

CloudFactory ist ein hervorragender Arbeitgeber, der eine mission-driven Kultur fördert und seinen Mitarbeitern die Möglichkeit bietet, in einem dynamischen Umfeld zu wachsen und einen echten Einfluss auf die Welt zu haben. Mit einem starken Fokus auf persönliche Entwicklung, innovativen Arbeitsmethoden und einer globalen Gemeinschaft, die Vielfalt schätzt, bietet CloudFactory nicht nur ein Arbeitsumfeld, sondern auch eine Plattform für bedeutungsvolle Arbeit und Zusammenarbeit. Hier können Sie Ihre Fähigkeiten als Senior Site Reliability Engineer weiterentwickeln und Teil einer Organisation werden, die sich leidenschaftlich für den Einsatz von KI zur Transformation der Gesellschaft einsetzt.

C

Kontaktdaten:

CloudFactory Recruiting-Team

StudySmarter Expertenrat🤫

Wir sind der Meinung, dass du so Senior Site Reliability Engineer erhalten könntest

Netzwerken in der IT-Community

In der IT-Consulting-Welt sollten wir regelmäßig auf Veranstaltungen wie Tech-Meetups oder Konferenzen gehen. Hier können wir nicht nur unser Netzwerk erweitern, sondern auch direkt mit potenziellen Arbeitgebern ins Gespräch kommen und unser Interesse an einer Vollzeitstelle zeigen.

Online-Foren und Gruppen nutzen

Sich in Online-Foren und Communities wie Stack Overflow oder LinkedIn-Gruppen umzusehen, kann uns helfen, Insider-Tipps zu erhalten und Informationen über offene Stellen in der IT-Beratung zu sammeln. Vergiss nicht, aktiv zu werden und Fragen zu stellen oder dein Wissen zu teilen – das erhöht unsere Sichtbarkeit!

Direkt bei CloudFactory bewerben

Viele Unternehmen, wie CloudFactory, stemmen ihre Vollzeitstellen bevorzugt über ihre eigenen Karriere-Webseiten. Also, lass uns regelmäßig auf deren Seite vorbeischauen und uns direkt bewerben, statt nur die üblichen Jobportale zu nutzen.

Überzeugende Projekte zeigen

Wir sollten unser Portfolio oder relevante Projekte gut sichtbar machen, egal ob das auf Github, persönlich oder auf LinkedIn ist. Bei IT-Consulting-Stellen kommt es oft auf praktische Erfahrungen an, also lass uns zeigen, was wir können!

Wir glauben, dass du diese Fähigkeiten brauchst, um Senior Site Reliability Engineer mit Bravour zu bestehen

Infrastruktur Engineering
DevOps
Site Reliability Engineering (SRE)
Kubernetes
Helm
Terraform
Cloudformation

Einige Tipps für deine Bewerbung 🫡

Zeige deine technischen Skills!:In der IT-Beratung zählen deine technischen Kenntnisse und Fähigkeiten. Achte darauf, relevante Programmiersprachen, Tools und Systeme in deinem Lebenslauf aufzulisten. Zeig auch, wenn du Zertifikate hast, die deine Kompetenz unterstützen – das könnte dir einen echten Vorteil verschaffen!

Verstehe die Branche!:Unterstreiche in deinem Anschreiben, dass du ein gutes Verständnis für aktuelle Trends und Herausforderungen in der IT-Branche hast. Zeig, dass du nicht nur die technischen Aspekte beherrschst, sondern auch die Bedürfnisse der Kunden erkennen und lösen kannst!

Deine Projekte zählen!:Falls du bereits an IT-Projekten gearbeitet hast, verlinke diese oder beschreibe sie in deinem Lebenslauf. Praktische Erfahrungen – sei es in Form von Praktika oder privaten Projekten – sind besonders wertvoll in der IT-Beratung. Zeige uns, was du kannst!

Individuelle Bewerbung ist der Schlüssel!:Jede Bewerbung sollte individuell auf CloudFactory und die ausgeschriebene Position Senior Site Reliability Engineer zugeschnitten sein. Teile uns mit, warum gerade du eine gute Wahl für unser Team bist. Das zeigt dein Engagement und deine Motivation, die über eine Standardbewerbung hinausgeht.

Wie man sich auf ein Vorstellungsgespräch bei CloudFactory vorbereitet

Technische Vorbereitung ist alles!

Da du dich auf eine Vollzeitstelle in der IT-Beratung bewirbst, solltest du dir wirklich einen Überblick über die wichtigsten Tools und Technologien verschaffen, die in der Branche verwendet werden. Sei bereit, technische Fragen zu beantworten, die sich auf Software-Architektur oder Systemintegration beziehen könnten.

Praxisbeispiele parat haben

In der IT-Beratung ist es wichtig, konkrete Beispiele aus deiner bisherigen Erfahrung zu bringen. Überlege dir Projekte, bei denen du erfolgreich einen Kunden beraten hast oder Herausforderungen gelöst hast. Das zeigt, dass du nicht nur theoretisches Wissen hast, sondern auch in der Praxis erfolgreich sein kannst.

Soft Skills betonen

Ein großer Teil der IT-Beratung ist die Kommunikation mit Kunden und das Verständnis ihrer Bedürfnisse. Bereite dich darauf vor, über deine zwischenmenschlichen Fähigkeiten zu sprechen, wie du mit herausfordernden Kunden umgehst oder wie du in Teams arbeitest. Das wird den Interviewern zeigen, dass du mehr als nur technisches Wissen mitbringst!

Fragen zum Unternehmen vorbereiten

Schau dir spezifisch die Projekte von CloudFactory an und überlege dir, welche Fragen du dazu stellen möchtest. Zeig Interesse an den aktuellen Herausforderungen, vor denen das Unternehmen steht, und wie du dazu beitragen könntest. Das hebt dich von anderen Bewerbern ab und zeigt, dass du wirklich motiviert bist.