- Contribute to Bitdeer's NeoCloud SRE platform for observing, protecting, and operating a multi-region GPU rental fleet
- Build collection agents, metrics/logs/traces/profiles stores, enrichment services, and collection monitors
- Write ingestion, query, and storage-path code
- Contribute to alerting, correlation, and SLO frameworks; implement and tune default alert rules
- Contribute to topology services, cluster-health rollups, and OSS-SRE-tool collection plugins for Kubernetes, Slurm, Ray, Volcano, Kueue, and KubeRay
- Help build remediation actuators, orchestration/workflow components, inspection probes, and job schedulers
- Instrument services with metrics, logs, and traces using OpenTelemetry
- Build dashboards and write actionable on-call runbooks
- Write unit, integration, and contract tests for shipped components
- Participate in chaos and soak tests led by senior engineers
- Develop components from design through production using GitOps and the CI/CD release pipeline
- Meet declared SLOs and maintain drift-free systems
- Operate what you build under senior-engineer guidance
- Participate in on-call as a shadow before taking primary responsibility
- Progress toward independently delivering components and owning a sub-context within 12 months
Requirements
- 0–2 years of software engineering experience; new graduates with strong projects or internships welcome
- Solid fundamentals in Go, Python, Java, or Rust; Go preferred
- Ability to write clean, tested, readable code and explain design choices
- Data structures, algorithms, concurrency, TCP/HTTP networking, and operating-system fundamentals
- Understanding of distributed-systems concepts including idempotency, retries, back-pressure, caching, and eventual consistency
- Hands-on exposure to Prometheus, Grafana, Loki, or similar observability tools
- Ability to write basic PromQL queries and instrument services
- Familiarity with Linux, the shell, system logs, and standard debugging tools
- Kubernetes basics, including Pods, Services, and Deployments; experience running something on Kubernetes
- Git, branching, pull requests, and CI pipelines such as GitHub Actions or GitLab CI
- Unit and integration testing discipline
- Clear written and verbal English
- Curiosity and willingness to learn GPU/AI infrastructure, AIOps, distributed systems, and observability
- Nice-to-have: internship or project experience in monitoring, observability, telemetry pipelines, or platform/SRE tooling
- Nice-to-have: exposure to GPU/AI infrastructure such as DCGM, InfiniBand/RoCE, Kubernetes GPU Operator, Slurm, or Ray
- Nice-to-have: exposure to AIOps/ML-adjacent tooling
- Nice-to-have: contributions to open-source observability or cloud-native projects
Core Competencies
Demonstrates strong software engineering fundamentals with proficiency in Go, Python, or Java, and a solid understanding of distributed systems and observability tools. Capable of writing clean, tested code and instrumenting services for performance monitoring and reliability.
Highest-signal resume keywords
- Proficiency In Go, Python, Java, Or Rust
- Experience With Kubernetes
- Familiarity With Prometheus And Grafana
- Understanding Of Distributed Systems Concepts
- Unit And Integration Testing Discipline
ATS Optimization Keywords
Hard Skills
- Software Engineering
- Clean Code Writing
- Data Structures
- Algorithms
- Concurrency
- TCP/HTTP Networking
- Operating System Fundamentals
- PromQL Queries
- Unit Testing
- Integration Testing
Soft Skills
- Clear Written And Verbal English
- Curiosity
- Willingness To Learn
Industry Keywords
- Observability
- Telemetry Pipelines
- Platform/SRE Tooling
- GPU/AI Infrastructure
- AIOps
Tools & Technologies
- Kubernetes
- Prometheus
- Grafana
- Loki
- Git
- GitHub Actions
- GitLab CI
- OpenTelemetry
- CI/CD Release Pipeline
- Linux
#J-18808-Ljbffr
SRE Monitoring Platform Software Engineer, Entry Level Arbeitgeber: Jobtailor
Als Front Office Supervisor in unserem dynamischen Team bieten wir Ihnen die Möglichkeit, in einem unterstützenden und freundlichen Arbeitsumfeld zu wachsen. Wir legen großen Wert auf die berufliche Entwicklung unserer Mitarbeiter und bieten regelmäßige Schulungen sowie die Chance, Verantwortung zu übernehmen. Unsere Lage ermöglicht es Ihnen, Teil einer lebendigen Gemeinschaft zu sein, während Sie gleichzeitig die Standards unseres Franchise-Partners einhalten und unseren Gästen einen unvergesslichen Aufenthalt bieten.