Quant DevOps Engineer

Quant DevOps Engineer

Vollzeit Vor Ort
Devopsroles

pThe bQuant DevOps Engineer /b owns and operates the application engineering platform, cloud infrastructure, and data execution environment that application and modeling workloads depend on. /ppThe role is accountable for the end‑to‑end reliability, scalability, security, performance, and cost efficiency of the platform, supporting large‑scale, data‑intensive and compute‑heavy workloads such as parallel Databricks clusters and high‑memory systems. /ppThis position requires broad and deep technical expertise, strong operational judgment, and the ability to act as a central technical interface between engineering teams, IT operations, security/SecOps, and data users. It is a senior ownership role with significant impact on delivery speed, platform stability, and infrastructure cost. /ph3Platform Infrastructure Ownership /h3ulliEnd-to-end ownership and delivery accountability for the application execution platform and production model runs (Databricks, Servers). /liliDay-to-day operational responsibility: availability, incident handling, runtime management, and delivery continuity for business-critical runs (incl. peak periods like YE). /liliManage Azure cloud infrastructure, including:ulliVirtual machines, storage, identity and access management (RBAC) /liliNetworking components such as firewalls, peering, and cross‑subscription connectivity /li /ul /liliLead standardization with central teams across observability, security controls, platform services, and “golden paths,” while keeping delivery running /liliEnsure platform reliability, scalability, and long‑term sustainability /li /ulh3CI/CD Automation /h3ulliDesign, maintain, and continuously improve CI/CD pipelines using Azure DevOps and related tooling /liliBuild and evolve automation using scripting and build tools /liliOptimize pipeline performance, reliability, and parallel execution to support large‑scale workloads /li /ulh3Container Runtime Management /h3ulliOwn Docker image creation, lifecycle management, and governance /liliOptimize build processes, caching strategies, and container security /liliSupport containerized execution environments for compute‑heavy workloads and services /li /ulh3Infrastructure as Code Configuration Management /h3ulliMaintain and evolve Infrastructure as Code using Terraform /liliOperate and improve configuration management systems (e.g. SaltStack) /liliReduce configuration drift and improve reproducibility across environments /li /ulh3Observability, Security Secrets /h3ulliOwn and operate observability platforms (e.g. ELK, Prometheus, Grafana) /liliEnsure meaningful metrics, logs, dashboards, and alerting are in place /liliManage secrets and platform security tooling (e.g. Wiz, Snyk) /liliCollaborate closely with Security and SecOps teams on controls, findings, and improvements /li /ulh3Data Platform Capacity Planning /h3ulliConfigure/support Databricks usage for application workloads; manage workspace‑level configuration/permissions as delegated; partner with central Databricks lead for global administration and optimization. /liliSupport large‑scale data and compute workloads /liliLead capacity planning for:ulliHighly parallel Databricks clusters (e.g. up to ~10 × 80‑node clusters) /liliMemory‑intensive systems (multi‑terabyte RAM) /liliData pipelines producing terabytes of data /li /ul /liliBalance performance, reliability, and cost across platform decisions /li /ulh3Operations Incident Response /h3ulliAct as senior escalation point for platform and infrastructure incidents /liliParticipate in a limited on‑call rotation /liliInvestigate incidents and execute or coordinate remediation /liliPerform manual interventions when automation is insufficient /liliDrive post‑incident reviews and platform improvements /li /ulh3Cross‑Team Organizational Coordination /h3ulliServe as primary technical contact for:ulliIT Operations /liliSecurity and SecOps /liliArchitecture and governance bodies /liliService management processes /li /ul /liliCoordinate platform‑related work across teams /liliSupport customer‑facing technical discussions related to platform capabilities and constraints /li /ulh3Experience /h3ulliBackground in bInsurance, Finance, or Scientific / High‑Performance Computing environments /b /liliStrong experience in bplatform or DevOps engineering /b within production environments /liliSolid expertise in bcloud infrastructure /b, preferably Microsoft Azure /liliHands‑on experience with bCI/CD, Infrastructure as Code, and container platforms /b /liliProven experience operating bdata‑intensive and compute‑heavy systems /b /li /ulh3Technical Competencies /h3ulliStrong troubleshooting and operational mindset /liliAbility to manage and balance competing constraints:ulliCost /liliPerformance /liliSecurity /liliReliability /li /ul /liliDeep understanding of platform stability, scalability, and automation /li /ulh3Professional Competencies /h3ulliSenior‑level autonomy and decision‑making capability /liliOwnership mindset; accountable for outcomes rather than tasks /liliAbility to operate as a trusted senior technical interface across teams /li /ul #J-18808-Ljbffr

Devopsroles

Kontaktdaten:

Devopsroles Recruiting-Team