- Deploy and operate parallel/distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre
- Design storage architectures optimized for AI workload patterns such as checkpoint I/O bursts, sequential dataset reads, and KV cache for inference
- Implement multi-tenant storage isolation with per-tenant QoS, quotas, and access controls
- Configure and optimize GPU Direct Storage for direct GPU-to-storage data paths
- Deploy and manage storage networking including NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX for cluster-wide storage orchestration
- Diagnose and tune storage performance using IOPS, throughput, latency profiling, fio, IOR, and mdtest
- Own the runbook for common failure modes
- Plan storage capacity aligned with GPU cluster growth and customer workload projections
- Manage firmware, data migration, and disaster recovery procedures
- Instrument storage telemetry including IO tail latency, checkpoint durations, NVMe SMART, filesystem health, and RDMA counters
- Feed telemetry into the platform team's metrics, logs, and traces store
- Partner with the platform team to define the storage-fault predictor, including signals, labels, and false-positive tolerances
- Convert novel incidents into automation, progressing from SOPs to runbook-as-code and agent-executable remediation
- Deliver observability and a baseline predictor for the top three storage-fault classes
- Reduce storage-incident MTTR
- Design storage for Nvidia GB200-class clusters
Requirements
- 5+ years in enterprise or HPC storage operations, with at least 2 years supporting AI/ML workloads
- Hands-on deployment and operations experience with at least two of: WEKA, VAST Data, Ceph, DDN/Lustre
- Strong understanding of AI training I/O patterns: checkpoint frequency, dataset loading, shuffle buffers
- Experience with high-performance storage networking (NFS over RDMA, NVMe-oF)
- Knowledge of GPU Direct Storage and RDMA-based data transfer
- Proficiency in storage performance benchmarking and tuning (fio, IOR, mdtest)
- Experience implementing multi-tenant storage with isolation and QoS
- Strong Linux systems knowledge (kernel tuning, filesystem internals, block device management)
- Experience shipping an anomaly detector for storage/IO telemetry or ability to articulate the labels and features needed
- Runbook-as-code mindset, with every SOP executable by a machine within a quarter
Core Competencies
Demonstrates expertise in deploying and managing parallel and distributed storage systems optimized for AI workloads, with a strong focus on performance tuning, multi-tenant isolation, and telemetry instrumentation. Proficient in high-performance storage networking and capable of implementing automation for storage operations.
Highest-signal resume keywords
- WEKA Deployment
- Ceph Operations
- GPU Direct Storage
- Storage Performance Benchmarking
- Multi-Tenant Storage Isolation
ATS Optimization Keywords
Hard Skills
- Storage Architecture Design
- AI Workload Optimization
- Storage Performance Tuning
- Linux Systems Knowledge
- Anomaly Detection for Telemetry
Soft Skills
- Problem-Solving
- Collaboration
Industry Keywords
- Enterprise Storage Operations
- HPC Storage
- AI/ML Workloads
- Telemetry Instrumentation
- Disaster Recovery Procedures
Tools & Technologies
- NFS over RDMA
- NVMe-oF
- Fio
- IOR
- Mdtest
#J-18808-Ljbffr
Senior GPU Cloud Storage Solutions Expert – SRE SME Arbeitgeber: Jobtailor
Als Front Office Supervisor in unserem dynamischen Team bieten wir Ihnen die Möglichkeit, in einem unterstützenden und freundlichen Arbeitsumfeld zu wachsen. Wir legen großen Wert auf die berufliche Entwicklung unserer Mitarbeiter und bieten regelmäßige Schulungen sowie die Chance, Verantwortung zu übernehmen. Unsere Lage ermöglicht es Ihnen, Teil einer lebendigen Gemeinschaft zu sein, während Sie gleichzeitig die Standards unseres Franchise-Partners einhalten und unseren Gästen einen unvergesslichen Aufenthalt bieten.