Tasks
Our Object Storage platform runs on Ceph, spans multiple data centers, and holds double-digit petabytes of customer data. It is growing fast. We are looking for a staff-level engineer who knows Ceph, Linux and the network underneath it well enough to keep it healthy as it scales.
We follow a "you build it, you run it" model. You will deploy, operate and improve the platform with the Network, SRE, Data Center and infrastructure support teams. A separate development team works on the product, and you will work closely with them.
Your tasks
- Own the reliability, performance and capacity of our production Ceph clusters across locations
- Solid understanding of server hardware (disks, controllers, NICs) and how it affects storage performance, so you can diagnose hardware-related issues remotely
- Diagnose the hardest issues across the full stack, from disk and server through the Linux network stack to the Ceph service, and lead incident response and postmortems
- Automate deployments, upgrades and routine operations (Ansible) so they need less manual work
- Build monitoring and observability that catch problems before customers do
- Plan capacity and growth with the Data Center and Network teams
- Take part in the on-call rotation
- Follow Ceph releases and the community, and bring useful practices back to the team
Your profile
- 5+ years as an SRE, Linux or storage engineer, with proven experience running Ceph in production (the more petabytes the better; 20 PB+ clusters are ideal)
- Strong knowledge of Linux and its network stack, plus good networking fundamentals
- Solid understanding of server hardware (disks, controllers, NICs) and how it affects storage performance
- Experience with object, block and file storage concepts
- Automation experience, ideally Ansible, and scripting in Python or Bash (you don't need to be a software engineer)
- Monitoring and observability experience (e.g., Prometheus, Grafana)
- Calm, structured troubleshooting when things break, and clear communication in English (German is a plus)
- Willingness to undergo extended security vetting (SÜ2)
Benefits
- Hybrid working model.
- Flexible working hours through trust-based working hours.
- At some locations a subsidized canteen and various free drinks.
- Modern office space with very good transport connections.
- Various employee discounts for activities and products.
- Employee events such as summer and winter parties, as well as workshops.
- Numerous training and development opportunities.
- Various health offers, such as sports and health courses.
Qualifications
#J-18808-Ljbffr
Staff Reliability Engineer (f/m/d) in Berlin Arbeitgeber: 1&1 IONOS SE
Strato, als Teil der IONOS Gruppe, bietet eine dynamische und innovative Arbeitsumgebung, die auf Wachstum und Erfolg ausgerichtet ist. Mit einem hybriden Arbeitsmodell, flexiblen Arbeitszeiten und umfangreichen Weiterbildungsangeboten fördert das Unternehmen nicht nur die berufliche Entwicklung seiner Mitarbeiter, sondern auch ein positives und kollaboratives Arbeitsklima. Die Möglichkeit, in einem führenden Technologieunternehmen zu arbeiten, das sich für Nachhaltigkeit und Kundenorientierung einsetzt, macht Strato zu einem attraktiven Arbeitgeber im DACH-Raum.