- Design and operate highly available, scalable, and secure cloud platforms on AWS
- Build and maintain Kubernetes-based infrastructure to support applications, data, and AI workloads
- Improve platform reliability through automation, Infrastructure as Code (IaC), and self-service capabilities
- Implement and enhance observability solutions using Datadog, including monitoring, logging, tracing, alerting, dashboards, and SLO management
- Support and optimize large-scale data processing environments using Airflow, Amazon EMR, S3, and other AWS data services
- Work with Data and AI teams to improve the reliability, scalability, and operational maturity of Machine Learning and Artificial Intelligence platforms
- Lead incident response activities, root cause analyses, and post-incident reviews
- Define and measure SLIs, SLOs, and error budgets
- Improve deployment processes, CI/CD pipelines, and release reliability
- Optimize cloud infrastructure utilization, performance, and costs
- Mentor team members and promote SRE best practices across the engineering organization
- Increase platform availability and reliability
- Improve observability and reduce incident resolution time
- Increase automation and reduce repetitive manual operational effort
- Deliver reliable, scalable, and cost-efficient Data and AI platforms
- Collaborate with Engineering teams to deliver resilient production systems
Requirements
- Currently pursuing or have completed a bachelor's degree
- Solid experience in Site Reliability Engineering, Platform Engineering, Cloud Engineering, or DevOps roles
- Strong hands-on experience with AWS services and cloud-native architectures
- Deep knowledge of Kubernetes and containerized workloads in production environments
- Experience managing and troubleshooting large-scale distributed systems
- Solid experience with observability platforms, preferably Datadog
- Experience supporting data platforms and pipelines using technologies such as Airflow, EMR, Spark, and S3
- Expertise in Infrastructure as Code (IaC) using Terraform or similar tools
- Experience building and maintaining CI/CD pipelines and platform automation
- Strong knowledge of Linux, networking, and system performance troubleshooting
- Proficiency in scripting and automation using Python, Bash, or similar languages
- Experience supporting large-scale cloud-native platforms in AWS environments
- Experience operating Kubernetes platforms and managing cluster lifecycles
- Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, error budgets, and operational excellence practices
- Experience implementing observability solutions using tools such as Datadog, Prometheus, Grafana, OpenTelemetry, or similar technologies
- Familiarity with data processing and workflow orchestration platforms such as Airflow, Spark, or EMR
- Experience with Infrastructure as Code (IaC) and platform automation practices
- AWS, Kubernetes, Terraform, or Datadog certifications
- Experience working in large-scale, highly available, or mission-critical enterprise environments
- Intermediate technical English
Core Competencies
Demonstrates expertise in designing and operating scalable, secure cloud platforms on AWS, with a strong focus on Kubernetes, Infrastructure as Code, and observability solutions. Proficient in enhancing platform reliability and automation while mentoring team members in Site Reliability Engineering best practices.
Highest-signal resume keywords
- AWS Services
- Kubernetes Management
- Infrastructure as Code (IaC)
- Observability Solutions (Datadog)
- CI/CD Pipeline Development
Hard Skills
- Site Reliability Engineering
- Platform Engineering
- Cloud Engineering
- DevOps
- Data Processing (Airflow, EMR, Spark)
- Scripting (Python, Bash)
- Linux System Performance
- Networking Troubleshooting
- Automation
- Cloud-Native Architectures
Soft Skills
- Mentoring
- Collaboration
- Incident Response Leadership
Certifications & Qualifications
- AWS Certification
- Kubernetes Certification
- Terraform Certification
- Datadog Certification
Industry Keywords
- Cloud Platforms
- Distributed Systems
- Operational Excellence
- SLOs
- SLIs
- Error Budgets
Tools & Technologies
- Datadog
- Terraform
- Prometheus
- Grafana
- OpenTelemetry
#J-18808-Ljbffr
Mid-Level SRE Analyst Arbeitgeber: Jobtailor
Als Front Office Supervisor in unserem dynamischen Team bieten wir Ihnen die Möglichkeit, in einem unterstützenden und freundlichen Arbeitsumfeld zu wachsen. Wir legen großen Wert auf die berufliche Entwicklung unserer Mitarbeiter und bieten regelmäßige Schulungen sowie die Chance, Verantwortung zu übernehmen. Unsere Lage ermöglicht es Ihnen, Teil einer lebendigen Gemeinschaft zu sein, während Sie gleichzeitig die Standards unseres Franchise-Partners einhalten und unseren Gästen einen unvergesslichen Aufenthalt bieten.