Senior Site Reliability Engineer (all genders)

Senior Site Reliability Engineer (all genders)

Berlin Vollzeit 63000 - 77000 € / Jahr (geschätzt) Homeoffice (teilweise)
F

Auf einen Blick

  • Aufgaben: Gestalte die Zuverlässigkeit unserer Systeme und arbeite an spannenden Cloud-Projekten.
  • Unternehmen: Führendes Unternehmen im Bereich Produktentdeckung für den europäischen E-Commerce.
  • Vorteile: Flexibles Arbeiten, modernes Tech-Stack und Raum für persönliche Entwicklung.
  • Weitere Informationen: Dynamisches Team mit offener Feedbackkultur und echten Wachstumschancen.
  • Warum dieser Job: Beeinflusse direkt den Erfolg führender E-Commerce-Marken in Europa.
  • Qualifikationen: Erfahrung als SRE oder starkes Interesse an der Weiterentwicklung in diese Rolle.

Das prognostizierte Gehalt liegt zwischen 63000 - 77000 € pro Jahr.

Introduction

FACT-Finder builds product discovery technology for e Commerce and is trusted by leading online shops across Europe with its two products Next Generation and Infinity.

Both products are moving toward a modern, hybrid platform based on Kubernetes and Harvester – with the option to scale fully into the cloud in the mid-term.

As a Senior Site Reliability Engineer (SRE), you make sure our systems stay fast, available, and scalable throughout this transformation.

You work closely with the Hosting team and experienced engineers, and actively shape our journey toward a modern Saa S company.

  • Your mission
  • You define and own SLOs, SLIs, and error budgets across both products and make data-driven decisions on reliability and performance.
  • You drive incident response: fast detection, clear communication, blameless postmortems, and meaningful follow-through.
  • You consistently reduce manual work through automation and Git Ops (e. g. Argo CD / Flux) and build out self-healing and self-service capabilities.
  • You support the development of an NG Search Operator (custom Kubernetes operator / CRDs) and the rollout of auto-scaling (HPA, VPA, KEDA, cluster autoscaler).
  • You evolve our observability – metrics, logs, traces, alerting, and runbooks that actually help on call.
  • You plan capacity and cost across on-premise (Frankfurt, Stockholm) and cloud – including burst scenarios into the public cloud.
  • You leverage AI tools to noticeably accelerate diagnosis, alerting, and operational workflows.

Your profile

  • Experience as an SRE, infrastructure, or production engineer in a Saa S or platform environment – or a strong software/operations background with a clear drive to grow into an SRE role.
  • Solid understanding of SLOs, error budgets, incident management, and observability.
  • Hands-on experience with Kubernetes and interest in cluster lifecycle, upgrades, and operator patterns.
  • Experience with or strong interest in Harvester or comparable HCI/virtualization platforms (Kube Virt, v Sphere/ESXi, Open Stack).
  • Familiarity with Git Ops (Argo CD / Flux), container storage (Longhorn, Ceph), and Kubernetes networking (load balancing, ingress).
  • Knowledge of auto-scaling primitives (HPA, VPA, cluster autoscaler, KEDA) and capacity planning on-prem and in the cloud.
  • Understanding of networking in production-grade datacenters (incl. VLAN).
  • A strong automation instinct and a mindset to structurally eliminate toil.
  • Practical experience using AI tools in day-to-day operations.
  • Fluent English; German is a plus.
  • THE JOY OF WORKING WITH US
  • Impact from day one: Your work directly influences the revenue of leading e Commerce brands across Europe.
  • Modern tech stack: Kubernetes, Harvester, Git Ops, auto-scaling, and an exciting path toward the cloud – with room to build things right.
  • AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.
  • Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.
  • Flexible work: Hybrid work model with a focus on outcomes.
  • Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.
  • Job Location
  • Berlin, Munich, Pforzheim or Stockholm (Hybrid)

About us

We are one of the leading

Product Discovery Platforms for European e Commerce . Our search, navigation and recommendations handle billions of shopper queries every year — driving a combined

GMV of more than €160 billion across our customers. These include

Intersport, Spar, Douglas and more than 2,000 further B2B and B2C e Commerce companies across Europe.

For over 20 years we have been an essential part of the complex e Commerce landscape, and with our AI-powered Product Discovery Experience (PDX) we are now entering the next phase — as a

Private-Equity-owned company (Genui) , under a new Mc Kinsey-shaped management team, with a clear growth and EBITDA plan. Offices in Pforzheim, Berlin, Munich and Stockholm.

Senior Site Reliability Engineer (all genders) Arbeitgeber: FACT-Finder Holding GmbH

Als Arbeitgeber bietet FactFinder eine dynamische und unterstützende Arbeitsumgebung, in der Sie als Account Executive für den italienischen Markt die Möglichkeit haben, bedeutende Geschäfte zu tätigen und langfristige Kundenbeziehungen aufzubauen. Mit einem transparenten Provisionsmodell, modernen Tools und einer Kultur, die auf Autonomie und Vertrauen setzt, fördern wir Ihr persönliches Wachstum und Ihre Karrierechancen in einem innovativen Umfeld. Arbeiten Sie mit führenden eCommerce-Unternehmen zusammen und gestalten Sie die Zukunft des digitalen Handels in Italien.

F

Kontaktdaten:

FACT-Finder Holding GmbH Recruiting-Team

Wir glauben, dass du diese Fähigkeiten brauchst, um Senior Site Reliability Engineer (all genders) mit Bravour zu bestehen

SRE Erfahrung
Kubernetes
GitOps (Argo CD / Flux)
Automatisierung
Fehlerbudget-Management
Incident Management
Observability