Senior Platform Support Engineer (Remote)

Senior Platform Support Engineer (Remote)

Vollzeit 63000 - 77000 € / Jahr (geschätzt) Homeoffice (teilweise)
R

Auf einen Blick

  • Aufgaben: Biete technischen Support für Kunden, die GPU-Cloud-Plattformen nutzen.
  • Unternehmen: Radian Arc, ein innovatives Unternehmen im Bereich Cloud Gaming und KI.
  • Vorteile: Attraktives Gehalt, flexibles Arbeiten und internationale Teamkultur.
  • Weitere Informationen: Dynamisches Arbeitsumfeld mit großartigen Entwicklungsmöglichkeiten.
  • Warum dieser Job: Gestalte die Zukunft der Technologie mit und arbeite an spannenden Projekten.
  • Qualifikationen: Mindestens 5 Jahre Erfahrung in Cloud-Support oder Systemadministration.

Das prognostizierte Gehalt liegt zwischen 63000 - 77000 € pro Jahr.

About Radian Arc.

Radian Arc provides an infrastructure-as-a-service (Iaa S) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks.

Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.

What impact you will have

Mission: Provide advanced technical support for customers running workloads on the GPU cloud platform, ensuring reliable operation of edge- and large-scale GPU clusters and infrastructure services.

The Senior Cloud Support Engineer acts as a technical escalation point for complex incidents, helping diagnose and resolve issues across compute, networking, storage, and orchestration layers.

This role works closely with engineering and operations teams to improve platform reliability, reduceincident frequency, and enhance the overall customer experience.

  • What you’ll do
  • Customer Support & Incident Management
  • Provide advanced technical support for customers operating workloads on bare-metal and virtualized GPU infrastructure.
  • Diagnose and resolve complex customer issues affecting GPU clusters, compute nodes, networking, and storage.
  • Investigate incidents across multiple layers of the stack including firmware, drivers, operating systems, and platform services.
  • Perform root cause analysis (RCA) for major incidents and contribute to long-term remediation efforts.
  • Serve as a technical escalation point for complex or high-priority support cases.
  • Infrastructure Troubleshooting

Troubleshoot issues affecting

  • GPU compute nodes
  • Kubernetes clusters
  • Networking infrastructure
  • Local NVMe, hyperconverged and distributed storage systems

Analyze logs, telemetry, and monitoring signals to identify underlying causes of platform instability.

Investigate issues related to GPU drivers, firmware, networking, and system performance.

Security Monitoring & Incident Triage

Monitor and investigate security alerts generated by the platform security stack.

  • Analyze and triage alerts generated by
  • Wazuh
  • The Hive
  • Cortex

Validate alerts, determine impact, and elevate potential security incidents to the security engineering team.

Assist in collecting system telemetry, logs, and forensic data required for incident investigations.

Improve alert runbooks and operational procedures to reduce false positives and improve response time.

GPU HPC Workload Support

Provide advanced support for large-scale GPU workloads running distributed training and inference jobs.

Diagnose failures affecting multi-GPU and multi-node workloads.

  • Investigate performance issues impacting distributed workloads, including:
  • GPU utilization
  • Communication latency
  • Storage bottlenecks
  • networking congestion
  • Support scheduling systems used for GPU workloads, including troubleshooting:
  • Job queue failures
  • Scheduling constraints
  • Cluster resource fragmentation
  • High-Performance Networking Troubleshooting

Diagnose issues affecting high-performance networking fabrics used by distributed workloads.

Support environments using

  • RDMA
  • Ro CE networking

Investigate performance issues affecting GPU-to-GPU communication and distributed training pipelines.

  • Data Center Coordination
  • Coordinate with customer’s data center technicians and infrastructure teams to perform remote diagnostics and hardware interventions when required.
  • Assist in validating on-premise installations and deployments of GPU infrastructure, ensuring hardware, networking, and platform components are correctly installed and operational.
  • Support hardware troubleshooting and identify faulty components across GPU nodes, networking equipment, and storage systems.
  • Coordinate and track hardware replacements and RMA processes with vendors and data center staff.
  • Validate hardware health after replacements, including GPU nodes, NICs, DPUs, storage devices, and power components.
  • Work closely with deployment and infrastructure teams to verify service readiness after installations, expansions, or hardware maintenance activities.
  • Operational Excellence
  • Participate in on-call 24/7 rotations to ensure production platform availability.
  • Respond to monitoring alerts and resolve operational incidents in accordance with defined SLAs.
  • Improve operational runbooks, troubleshooting guides, and support documentation.
  • Contribute to improving incident response processes and operational tooling.
  • Cross-Team Collaboration
  • Work closely with platform engineering, infrastructure engineering, and networking teams to resolve systemic issues.
  • Provide feedback to engineering teams on recurring operational problems affecting customers.
  • Help translate customer issues into actionable improvements for the platform.
  • Automation & Tooling
  • Develop automation scripts and tools to streamline support workflows.
  • Improve observability dashboards and alerts to enable faster issue detection and resolution.
  • Contribute to automation initiatives that reduce manual intervention in operational processes.
  • Knowledge Sharing & Mentorship
  • Mentor medior support engineers and share troubleshooting expertise across the team.
  • Lead the creation of knowledge base articles, troubleshooting guides, and operational documentation.
  • Contribute to training initiatives that improve the team’s technical capabilities.
  • Technical Stack
  • Operating Systems
  • Linux (Ubuntu)
  • GPU Infrastructure
  • NVIDIA GPU platforms
  • CUDA drivers
  • GPU monitoring tools (nvidia-smi)
  • Platform Infrastructure
  • Kubernetes
  • Container runtimes
  • Kube Virt
  • Distributed compute environments
  • Networking
  • NVIDIA Cumulus
  • OOB, north-south, and east-west fabric topologies
  • TCP/IP
  • VLAN, VXLAN, OVS/OVN
  • Routing fundamentals (BGP, VRFs)
  • DNS / DHCP
  • High-performance networking (RDMA/NVLink/NCCL)
  • Observability
  • Grafana
  • Zabbix
  • Security Monitoring
  • Wazuh
  • The Hive
  • Cortex
  • Automation
  • Python
  • Bash
  • Ansible, Terraform
  • Collaboration & Documentation
  • Jira
  • Confluence
  • Zendesk
  • Pager Duty
  • Slack
  • What you’ll need
  • Core Experience
  • 5+ years of experience in cloud support, infrastructure operations, or systems administration.
  • Experience supporting large-scale infrastructure environments or GPU clusters.
  • Systems Expertise
  • Strong Linux systems administration skills.
  • Experience troubleshooting issues across compute, networking, and storage layers.
  • Familiarity with Kubernetes platforms and containerized workloads.
  • GPU Infrastructure
  • Experience working with GPU hardware platforms or HPC environments.
  • Familiarity with GPU monitoring tools and debugging GPU-related issues.
  • Networking
  • Solid understanding of networking fundamentals including L2/L3 concepts, routing, and load balancing.
  • Ability to diagnose connectivity issues affecting distributed workloads.
  • Operational Mindset
  • Strong troubleshooting and incident response skills.
  • Experience participating in on-call rotations and handling production incidents.
  • Ability to perform root cause analysis and drive operational improvements.
  • Communication & Collaboration
  • Excellent written and verbal communication skills.
  • Ability to explain complex technical concepts to both technical and non-technical stakeholders.
  • Proven ability to collaborate effectively with cross-functional engineering teams.

Location & work modality

Malaysia or comparable time zone

Start

August 2026

Type of Contract

Contractor

Average 40 hours per week, 9x5 business hour support with after hour on-call response/resolution for category 1 incidents

What we offer

Attractive compensation package reflecting your expertise and experience.

A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid-friendly approach.

You’ll be part of a fast-growing scale-up with a mission to make a positive impact, offering an exciting career evolution.

Our job titles may span more than one job level.

The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.

Our inclusive responsibility

Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer.

All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.

#J-18808-Ljbffr

Senior Platform Support Engineer (Remote) Arbeitgeber: Radian Arc

Radian Arc ist ein hervorragender Arbeitgeber, der eine dynamische und inklusive Arbeitsumgebung bietet, in der Vielfalt geschätzt wird. Mit einem attraktiven Vergütungspaket und flexiblen Arbeitsmodellen ermöglicht das Unternehmen seinen Mitarbeitern, sich in einer schnell wachsenden Scale-up-Umgebung weiterzuentwickeln und einen positiven Einfluss auf die Technologiebranche auszuüben. Die Möglichkeit zur Zusammenarbeit mit internationalen Teams und die Förderung von beruflichem Wachstum machen Radian Arc zu einem idealen Arbeitsplatz für talentierte Fachkräfte.

R

Kontaktdaten:

Radian Arc Recruiting-Team

StudySmarter Expertenrat🤫

Wir sind der Meinung, dass du so Senior Platform Support Engineer (Remote) erhalten könntest

Netzwerken in der IT-Community

In der IT-Consulting-Welt sollten wir regelmäßig auf Veranstaltungen wie Tech-Meetups oder Konferenzen gehen. Hier können wir nicht nur unser Netzwerk erweitern, sondern auch direkt mit potenziellen Arbeitgebern ins Gespräch kommen und unser Interesse an einer Vollzeitstelle zeigen.

Online-Foren und Gruppen nutzen

Sich in Online-Foren und Communities wie Stack Overflow oder LinkedIn-Gruppen umzusehen, kann uns helfen, Insider-Tipps zu erhalten und Informationen über offene Stellen in der IT-Beratung zu sammeln. Vergiss nicht, aktiv zu werden und Fragen zu stellen oder dein Wissen zu teilen – das erhöht unsere Sichtbarkeit!

Direkt bei Radian Arc bewerben

Viele Unternehmen, wie Radian Arc, stemmen ihre Vollzeitstellen bevorzugt über ihre eigenen Karriere-Webseiten. Also, lass uns regelmäßig auf deren Seite vorbeischauen und uns direkt bewerben, statt nur die üblichen Jobportale zu nutzen.

Überzeugende Projekte zeigen

Wir sollten unser Portfolio oder relevante Projekte gut sichtbar machen, egal ob das auf Github, persönlich oder auf LinkedIn ist. Bei IT-Consulting-Stellen kommt es oft auf praktische Erfahrungen an, also lass uns zeigen, was wir können!

Wir glauben, dass du diese Fähigkeiten brauchst, um Senior Platform Support Engineer (Remote) mit Bravour zu bestehen

Technische Unterstützung
Fehlerdiagnose
GPU-Cluster Management
Kubernetes
Netzwerkinfrastruktur
Speichersysteme
Root Cause Analysis (RCA)

Einige Tipps für deine Bewerbung 🫡

Zeige deine technischen Skills!:In der IT-Beratung zählen deine technischen Kenntnisse und Fähigkeiten. Achte darauf, relevante Programmiersprachen, Tools und Systeme in deinem Lebenslauf aufzulisten. Zeig auch, wenn du Zertifikate hast, die deine Kompetenz unterstützen – das könnte dir einen echten Vorteil verschaffen!

Verstehe die Branche!:Unterstreiche in deinem Anschreiben, dass du ein gutes Verständnis für aktuelle Trends und Herausforderungen in der IT-Branche hast. Zeig, dass du nicht nur die technischen Aspekte beherrschst, sondern auch die Bedürfnisse der Kunden erkennen und lösen kannst!

Deine Projekte zählen!:Falls du bereits an IT-Projekten gearbeitet hast, verlinke diese oder beschreibe sie in deinem Lebenslauf. Praktische Erfahrungen – sei es in Form von Praktika oder privaten Projekten – sind besonders wertvoll in der IT-Beratung. Zeige uns, was du kannst!

Individuelle Bewerbung ist der Schlüssel!:Jede Bewerbung sollte individuell auf Radian Arc und die ausgeschriebene Position Senior Platform Support Engineer (Remote) zugeschnitten sein. Teile uns mit, warum gerade du eine gute Wahl für unser Team bist. Das zeigt dein Engagement und deine Motivation, die über eine Standardbewerbung hinausgeht.

Wie man sich auf ein Vorstellungsgespräch bei Radian Arc vorbereitet

Technische Vorbereitung ist alles!

Da du dich auf eine Vollzeitstelle in der IT-Beratung bewirbst, solltest du dir wirklich einen Überblick über die wichtigsten Tools und Technologien verschaffen, die in der Branche verwendet werden. Sei bereit, technische Fragen zu beantworten, die sich auf Software-Architektur oder Systemintegration beziehen könnten.

Praxisbeispiele parat haben

In der IT-Beratung ist es wichtig, konkrete Beispiele aus deiner bisherigen Erfahrung zu bringen. Überlege dir Projekte, bei denen du erfolgreich einen Kunden beraten hast oder Herausforderungen gelöst hast. Das zeigt, dass du nicht nur theoretisches Wissen hast, sondern auch in der Praxis erfolgreich sein kannst.

Soft Skills betonen

Ein großer Teil der IT-Beratung ist die Kommunikation mit Kunden und das Verständnis ihrer Bedürfnisse. Bereite dich darauf vor, über deine zwischenmenschlichen Fähigkeiten zu sprechen, wie du mit herausfordernden Kunden umgehst oder wie du in Teams arbeitest. Das wird den Interviewern zeigen, dass du mehr als nur technisches Wissen mitbringst!

Fragen zum Unternehmen vorbereiten

Schau dir spezifisch die Projekte von Radian Arc an und überlege dir, welche Fragen du dazu stellen möchtest. Zeig Interesse an den aktuellen Herausforderungen, vor denen das Unternehmen steht, und wie du dazu beitragen könntest. Das hebt dich von anderen Bewerbern ab und zeigt, dass du wirklich motiviert bist.