Cloud Performance Engineering - Site Reliability Engineer

Cloud Performance Engineering - Site Reliability Engineer

Vollzeit 60000 - 80000 € / Jahr (geschätzt) Kein Homeoffice möglich
S

Auf einen Blick

  • Aufgaben: Gestalte die Zuverlässigkeit und Leistung von Cloud-Diensten für ein besseres Gesundheitswesen.
  • Unternehmen: Smile Digital Health - Innovatives Unternehmen mit Fokus auf globale Gesundheit.
  • Vorteile: Flexibles Arbeiten, wettbewerbsfähiges Gehalt und umfassende Gesundheitsleistungen.
  • Weitere Informationen: Dynamisches Team, das Vielfalt und Inklusion schätzt.
  • Warum dieser Job: Trage zur Verbesserung der globalen Gesundheit bei und arbeite mit modernster Technologie.
  • Qualifikationen: Erfahrung mit Cloud-Anbietern und Performance Engineering erforderlich.

Das prognostizierte Gehalt liegt zwischen 60000 - 80000 € pro Jahr.

Working for a company like Smile Digital Health means supporting our mandate for

#Better Global Health .

We strive towards this goal every day, and the results can be seen in the impact of our innovative health data platform and data management solutions, which are used in over 20 countries.

We were #19 on Deloitte's Technology Fast 50 Ranking for 2024!

Smile Digital Health makes it easy for healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform.

At its heart, the Smile platform enables people and organizations to better manage healthcare data.

We help generate and liberate structured healthcare data to ensure effective delivery across care teams and health systems bringing

#Better Global Health to patients everyday!

The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health, its clients, and partners.

This role designs and automates performance testing frameworks, integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks.

Working closely with engineering, product, and security teams, the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.

Responsibilities

  • Collaborate with our Security Operations teams to define and implement best practices around Cloud Service Provider configuration for Azure and other cloud providers.
  • Develop, implement, and coordinate a multi-tenant approach around service offerings for databases, container platforms, authentication, certificates, and product registries.
  • Design, develop, and maintain cloud performance testing strategies, frameworks, and environments to validate application scalability, reliability, and resiliency.
  • Develop and automate load, stress, spike, and endurance (soak) testing as part of CI/CD pipelines.
  • Analyze application and infrastructure performance to identify bottlenecks and recommend performance optimizations across cloud-native services.
  • Develop and maintain cost and utilization tracking and attribution processes across Cloud Service Providers.
  • Create documentation detailing Cloud Service Provider offerings, implementation patterns, and best practices.
  • Develop and maintain technical relationships with our core Cloud Service Providers.
  • Implement and maintain secure, scalable infrastructure platforms for delivering cloud services.
  • Ensure internal and external SLAs are consistently met or exceeded, while continuously monitoring and improving system performance, reliability, and availability.
  • Create tools for automating deployment, monitoring, and platform operations.
  • Implement and manage observability solutions (logging, metrics, tracing) using Open Telemetry, Prometheus, Grafana, Azure Monitor, and related technologies to provide actionable performance insights.
  • Plan and execute chaos engineering experiments to evaluate and improve application resiliency and fault tolerance.

Requirements

  • 5+ years of experience with Cloud Service Providers and best practices around implementation and configuration, preferably managing Azure environments supporting Saa S products.
  • Experience working across multiple cloud providers (Azure required; AWS and/or Google Cloud Platform considered an asset).
  • Strong experience in Cloud Performance Engineering, including performance analysis, capacity planning, scalability testing, and optimization of distributed cloud-native applications.
  • Proven experience working with microservices architecture, with a strong focus on Java-based services.
  • Experience applying Chaos Engineering practices to evaluate and improve system resiliency.
  • Strong experience designing and executing performance testing strategies, including load, stress, spike, and endurance (soak) testing, to validate application scalability and defined latency and error-rate thresholds.
  • Hands-on experience with performance testing tools such as JMeter, Gatling, Azure Load Testing, or k6.
  • Experience validating application services sustaining 500+ transactions per second (TPS) while meeting defined performance objectives.
  • Hands-on experience deploying and managing containerized applications using Docker and Kubernetes, including autoscaling and performance optimization.
  • Experience using Terraform to provision and manage cloud infrastructure using Infrastructure as Code (Ia C).
  • Experience tuning Kafka (partitioning, consumer group sizing, throughput/latency trade-offs) and other messaging/queueing platforms to sustain target transaction rates.
  • Hands-on experience implementing and using observability platforms including Open Telemetry, Prometheus, Grafana, Azure Monitor, Application Insights, and Log Analytics.
  • Proven experience with Security and Compliance (SOC 2, HIPAA, ISO 27001) best practices and implementing controls that support high-velocity software delivery teams.

Benefits

  • Remote Work Environment
  • Flexible Time Away From Work Policy including PTO, Personal and Sick Days
  • Competitive Salary and Health/Medical Benefits
  • RRSP/TFSA/401K Employee Contribution
  • Life and Disability
  • Employee Assistance Program
  • FHIR Study Program and Skillsoft Learning
  • Super HAPI Fun Club

Smile's core values include respect, inclusion, embracing our differences, and celebrating shared values because our people are the foundation of our success.

We are big on creating a sense of belonging and empowering each other to bring our authentic selves to work.

We are dedicated to fostering a workplace that values diversity, equity, and inclusion.

We welcome and encourage candidates of all backgrounds to apply.

Candidates are encouraged to inform us if they wish to discuss or require accommodations during interviews or while working at Smile.

#J-18808-Ljbffr

Cloud Performance Engineering - Site Reliability Engineer Arbeitgeber: Smile Digital Health

Smile Digital Health ist ein hervorragender Arbeitgeber, der sich für #BetterGlobalHealth einsetzt und innovative Lösungen im Gesundheitsdatenmanagement bietet. Mit einem flexiblen Arbeitsumfeld, umfangreichen Weiterbildungsmöglichkeiten und einem starken Fokus auf Diversität und Inklusion fördert das Unternehmen eine Kultur des Respekts und der Zugehörigkeit. Die Mitarbeiter profitieren von wettbewerbsfähigen Gehältern, umfassenden Gesundheitsleistungen und einem unterstützenden Team, das gemeinsam an der Verbesserung der globalen Gesundheitsversorgung arbeitet.

S

Kontaktdaten:

Smile Digital Health Recruiting-Team

Wir glauben, dass du diese Fähigkeiten brauchst, um Cloud Performance Engineering - Site Reliability Engineer mit Bravour zu bestehen

Cloud Service Providers
Azure
AWS
Google Cloud Platform
Cloud Performance Engineering
Performance Analysis
Capacity Planning