Role overview
Staff-level DevOps opening to lead the design and evolution of a Kubernetes-based platform engineering practice for a globally available, redundant microservices environment. The role combines deep technical ownership of multi-cluster Kubernetes infrastructure with broad influence across engineering, including paved-road tooling for AI/ML workloads. It is an individual-contributor technical leadership role, not a people-management position.
Responsibilities
- Own the architecture and roadmap for the multi-cluster Kubernetes platform, including scaling, upgrades, and multi-tenancy.
- Establish GitOps-based deployment workflows using tools such as Argo CD or Flux.
- Design and evolve an internal developer platform that gives engineering teams self-service access to compute, environments, and observability.
- Architect service mesh, networking, and ingress strategies for reliable, secure service-to-service communication.
- Define platform standards for GPU and ML workload scheduling and partner with AI/ML teams on training and inference infrastructure needs.
- Drive capacity planning and cost optimization across Kubernetes and cloud infrastructure, and own platform reliability through SLOs, incident response, and postmortems.
- Mentor experienced engineers and represent platform engineering in cross-org technical decisions.
Requirements
- Bachelor's degree in Computer Science or a related field, or equivalent work experience.
- 8+ years of DevOps, platform, or infrastructure engineering experience.
- 5+ years of hands-on Kubernetes experience in production, including cluster architecture, upgrades, and multi-tenant environments.
- Strong experience operating cloud infrastructure across AWS and GCP.
- Strong GitOps experience with Argo CD or Flux, and Infrastructure-as-Code experience with Terraform or Pulumi.
- Deep understanding of container networking, ingress, and service mesh (e.g., Istio, Linkerd, Cilium), plus middleware experience with Nginx, Kafka, and Redis at scale.
- Experience with GPU scheduling and resource quotas for ML/AI workloads on GKE and EKS, and operating observability platforms such as Prometheus, Grafana, Datadog, or OpenTelemetry.
Nice to have
- Experience designing internal developer platforms or paved-road tooling.
- Strong Linux, networking, storage, and security fundamentals, plus excellent cross-team communication skills.
Benefits and work setup
- Remote-eligible across Mexico; team members within roughly 80 km of the Guadalajara office are expected onsite to support collaboration.
- Package includes major medical insurance for the employee and dependents, vision and dental coverage, life insurance, paid time off (10 personal days before the first anniversary, then vacation plus additional personal days), a 30-day Christmas bonus, 50% vacation premium, company-matched food vouchers, and a 13% matched savings fund.
- Employee Assistance Program and ongoing learning and development opportunities.
#J-18808-Ljbffr
Staff DevOps Engineer Arbeitgeber: Nextiva
Nextiva ist ein hervorragender Arbeitgeber, der eine dynamische und innovative Arbeitsumgebung bietet, in der Mitarbeiter die Möglichkeit haben, bedeutende Beiträge zu leisten und ihre Karriere voranzutreiben. Mit einem starken Fokus auf persönliche Entwicklung, flexiblen Arbeitszeiten und umfassenden Gesundheitsleistungen fördert Nextiva eine Kultur des Wachstums und der Zusammenarbeit. Die Position des Territory Partner Managers ermöglicht es Ihnen, strategische Partnerschaften aufzubauen und gleichzeitig von einem unterstützenden Team und einer positiven Unternehmenskultur zu profitieren.