Site Reliability Engineer

Site Reliability Engineer

Vollzeit Kein Homeoffice möglich
A

Overview

We are hiring a skilled Site Reliability Engineer (SRE) to strengthen our 24x7 operational support team. In this role, you’ll ensure platform stability, reliability, and security through observability, automation, and proactive monitoring. You will manage Kubernetes, CI/CD pipelines, and IaC solutions, while scripting in Python, Go, or Bash to drive efficiency. You’ll configure and optimize Elasticsearch/Prometheus platforms, ensuring secure logging, scalable monitoring, and insightful dashboards. As part of a global support team, you’ll participate in incident response, troubleshooting, and major incident management, ensuring rapid recovery and minimal downtime. You’ll also uphold security standards, compliance requirements, and contribute to the continuous improvement of SOPs. Collaboration with engineers and stakeholders will be key in delivering resilient, reliable, and high-performing systems.

Responsibilities

  • Manage Kubernetes, CI/CD pipelines, and IaC solutions
  • Scripting in Python, Go, or Bash to drive efficiency
  • Configure and optimize Elasticsearch/Prometheus platforms, ensuring secure logging, scalable monitoring, and insightful dashboards
  • Participate in incident response, troubleshooting, and major incident management
  • Uphold security standards, compliance requirements, and contribute to the continuous improvement of SOPs
  • Collaborate with engineers and stakeholders to deliver resilient, reliable, and high-performing systems

Qualifications

  • Experience managing Kubernetes, CI/CD pipelines, and IaC solutions
  • Proficiency scripting in Python, Go, or Bash
  • Experience configuring and optimizing Elasticsearch and Prometheus; secure logging and dashboards
  • Experience in incident response, troubleshooting, and major incident management
  • Knowledge of security standards and compliance; ability to contribute to SOP improvements
  • Strong collaboration with engineers and stakeholders

#J-18808-Ljbffr

Site Reliability Engineer Arbeitgeber: Apprize Technology Solutions

Als Arbeitgeber im Bereich der KI- und ML-Entwicklung bieten wir Ihnen die Möglichkeit, an innovativen Projekten zu arbeiten, die die Zukunft der Bildverarbeitung gestalten. Unsere offene und kollaborative Unternehmenskultur fördert kreatives Denken und kontinuierliches Lernen, während wir Ihnen durch gezielte Schulungen und Entwicklungsmöglichkeiten helfen, Ihre Fähigkeiten weiter auszubauen. Zudem profitieren Sie von flexiblen Arbeitszeiten und einem modernen Arbeitsplatz in einer dynamischen Umgebung, die den Austausch von Ideen und die Zusammenarbeit im Team unterstützt.

A

Kontaktdaten:

Apprize Technology Solutions Recruiting-Team