Auf einen Blick
- Aufgaben: Stärke die Zuverlässigkeit und Leistung unserer Systeme durch Automatisierung und Monitoring.
- Unternehmen: Playon, ein innovatives Unternehmen im Bereich Software Engineering.
- Vorteile: Vielfältige Gesundheitspläne, Aktienoptionen, flexibles Arbeiten und offene Urlaubspolitik.
- Weitere Informationen: Dynamisches Team mit großartigen Entwicklungsmöglichkeiten.
- Warum dieser Job: Gestalte die Zukunft der Software-Zuverlässigkeit und arbeite an spannenden Projekten.
- Qualifikationen: Erfahrung in Python, Cloud-Infrastruktur und CI/CD-Pipelines erforderlich.
Das prognostizierte Gehalt liegt zwischen 60000 - 80000 € pro Jahr.
Playon is looking for an experienced Senior Site Reliability Engineer to help us strengthen the reliability, performance, and scalability of our systems.
This role sits at the intersection of software engineering and operations — focused on building the tools, automation, and visibility that enable our teams to deliver resilient software at scale.
You’ll work closely with application engineers, Dev Ops, and QA teams to evolve our infrastructure, CI/CD pipelines, observability frameworks, and reliability practices.
This is a hands‑on engineering role with a strong emphasis on automation, performance analysis, and continuous improvement.
- The Outcomes You’ll Deliver
- Assess and improve visibility: Work with engineering teams to review our current dashboards, metrics, and logs, identify the biggest gaps, and make targeted improvements that help us better understand system health.
- Tighten monitoring and alerting: Refine alerts and dashboards for the most critical services so we can catch issues earlier and respond faster.
- Build observability into delivery: Add instrumentation and telemetry into existing build and deploy processes to make reliability checks part of our normal release workflow.
- Clarify what “reliable” means: Help define initial SLIs and SLOs for a few core user flows, aligning the team on what good performance and availability look like.
- Streamline incident response: Partner with the Event Commander/on‑call rotation to improve how we communicate, coordinate, and follow up during incidents.
- Reduce manual effort: Automate routine checks and monitoring tasks to free up engineers for more impactful work.
Over time, you'll take on a larger role shaping how we measure, monitor, and improve reliability across all services — setting standards, mentoring others, and helping engineering teams make data‑driven decisions about performance and stability.
Key Responsibilities
- Contribute to system observability i. e implementing, improving metrics, alerting, and dashboards for better insight and faster recovery.
- Develop automation, tooling, and monitoring solutions to support high service availability.
- Partner with application and quality engineering teams to implement best practices in reliability, release automation, and testing.
- Drive operational excellence through proactive incident prevention, blameless postmortems, and capacity planning.
- Participate in on‑call rotations to support critical services and ensure rapid response to incidents.
Qualifications
- Solid experience in Python, especially for automation, tooling, and data‑driven operational tasks.
- Proficiency in at least one of Java, C++, or Go.
- Strong understanding of Linux systems, cloud infrastructure (AWS, GCP, or Azure), and modern deployment practices (Docker, Kubernetes, Terraform).
- Experience with CI/CD pipelines, version control, and automated testing frameworks.
- Experience with observability tools (Prometheus, Grafana, ELK, Datadog, etc.) and log/metric analysis for diagnosing issues.
- Proven experience facilitating and documenting Critical User Journeys translating them to actionable SLA/SLO for automation.
- Demonstrated ability to collaborate with cross‑functional teams and communicate clearly in high‑impact situations.
- A problem‑solver who approaches reliability as a shared responsibility across engineering.
- Familiarity with AI‑augmented development tools (Claude, Codex) as part of a modern engineering workflow.
- Nice to Have
- Experience writing or maintaining end‑to‑end or integration tests for distributed systems.
- Background in performance testing, capacity planning, or chaos engineering.
- Contributions to internal developer tooling or reliability‑focused frameworks.
- Exposure to security, compliance, or change management processes in production environments.
- Relevant certifications.
- The Benefits We Offer
- Multiple medical insurance plans to choose from
- Dental, vision, life and disability insurance
- Employee Emergency Fund
- Company equity (stock options)
- Open PTO policy
- 401K plan with company match
- Hybrid/flexible work environment
Note: Must be a full‑time employee to participate in the company’s employee health benefit plan. Part‑time employees and interns are not eligible to participate.
#J-18808-Ljbffr
Senior Site Reliability Engineer Arbeitgeber: Playonsports
Playon ist ein hervorragender Arbeitgeber, der seinen Mitarbeitern die Möglichkeit bietet, in einem dynamischen und innovativen Umfeld zu arbeiten. Mit einem starken Fokus auf Automatisierung und kontinuierliche Verbesserung fördert das Unternehmen eine Kultur des Wachstums und der Zusammenarbeit, während es gleichzeitig attraktive Vorteile wie flexible Arbeitszeiten, umfassende Gesundheitsleistungen und Unternehmensanteile bietet. Die Position des Senior Site Reliability Engineer ermöglicht es Ihnen, aktiv an der Weiterentwicklung der Systemzuverlässigkeit und -leistung mitzuwirken und dabei Ihre Fähigkeiten in einem unterstützenden Team weiter auszubauen.