Auf einen Blick
- Aufgaben: Entwickle und betreibe Kubernetes-Infrastruktur für KI-Plattformen und sorge für Zuverlässigkeit.
- Unternehmen: Innovatives Unternehmen, das mit großen Marken zusammenarbeitet und eine unterstützende Kultur bietet.
- Vorteile: 35 Tage Urlaub, private Krankenversicherung, finanzielle Beratung und 40 Stunden bezahlte Weiterbildung.
- Weitere Informationen: Flexibles Arbeiten, inklusive hybrider und remote Optionen, in einem unterstützenden Team.
- Warum dieser Job: Gestalte die Zukunft der KI mit und arbeite an spannenden Projekten in einem dynamischen Umfeld.
- Qualifikationen: Mindestens 5 Jahre Erfahrung in DevOps oder SRE, starke Kenntnisse in Kubernetes und Terraform.
Das prognostizierte Gehalt liegt zwischen 58500 - 71500 € pro Jahr.
- Working at Create Future
- Create Future is Version 1 company and an
AI-native consulting partner where people do work that matters and are supported to do it well.
We work alongside organisations such as Pay Pal, adidas, Nat West, Fan Duel and Money Saving Expert, building digital products and services that make a difference while always putting people first.
We’re a team of creators.
We write code, shape delivery, build go-to-market strategies, develop AI solutions and create the practices that support our people.
We work side by side with our clients, challenging what’s not working and helping them to build the future.
Our commitment to craft, quality, and culture has helped us scale to over 600 people in just a few years.
- Our UK Benefits
- 35 days leave (including bank holidays).
- Private medical insurance.
- Enhanced parental and adoption leave.
- Financial coaching + 5% pension match.
- 40 hours of paid learning and development.
View our full list of UK benefits.
Create Future is a Great Place to Work-Certified™ company and has won Best Workplaces UK multiple years in a row.
Join us on our journey. Let’s create tomorrow, together, today.
About the role
AI platforms are still mostly run like prototypes.
There's usually a Kubernetes cluster somebody set up in a hurry, no SLOs, no runbooks, and a cost line nobody can explain.
We're engaging a Senior Cloud Engineer to bring proper SRE discipline to a major AI platform programme with a client in a high-traffic, heavily regulated consumer sector.
You'll own how the platform runs.
That means the Kubernetes infrastructure behind model serving, agent orchestration and batch inference.
It means the CI/CD pipelines that ship models, agents, tools and prompt changes.
It means the SLOs and on-call practice that make reliability a stated commitment rather than a hope, and the observability that makes AI-specific failure modes visible, including drift, silent quality regression, cost blowouts and agent loops.
You'll be the operational conscience of the programme.
When a new capability is about to ship without limits, monitoring or a runbook, you're the person who says so.
This is a role for someone who finds that work satisfying rather than thankless, and who wants to do it on a platform where the SRE patterns are still being written.
- What you'll be doing
- Technical Delivery & Implementation
- Kubernetes for AI workloads: Design, build and operate the Kubernetes infrastructure behind model serving, agent orchestration and batch inference, defined in Terraform and deployed through Git Ops.
- CI/CD for non-deterministic systems: Build deployment pipelines suited to AI workloads, covering model rollouts, agent and tool updates, and prompt and configuration changes, with automated eval and regression checks before promotion.
- Reliability engineering: Define and track SLOs and SLAs for platform services, run incident response and root cause analysis, and write post-mortems people will read.
- On-call and runbooks: Take part in the programme's on-call rotation, and build runbooks clear enough that someone else can use them at 3am.
- AI-specific observability: Instrument latency, token usage, cost per request, model and agent error rates, retrieval quality and drift, with dashboards and alerting that surface problems before users do.
- Fin
Ops: Embed cost estimation, anomaly detection and resource optimisation into the pipeline as a default rather than a monthly review.
- Security posture: Apply zero-trust networking, least-privilege access and secrets discipline to a platform handling sensitive data across multiple third-party model providers.
- Client Delivery & Stakeholder Management
- Operability from day one: Work with the inference control plane, evaluation and knowledge platform workstreams so new capabilities arrive with monitoring, limits and a runbook, instead of acquiring them after the first incident.
- Technical advisory: Guide the client's architecture and AWS service choices on this platform, and build the working relationships that let that guidance land.
- Cost accountability: Understand the programme’s budget context, build the cost‑benefit case for infrastructure decisions, and be able to justify spend to a finance audience.
- Risk and delivery: Spot emerging operational bottlenecks in the client's estate, raise them early, and plan and estimate your own stream accurately.
- Stakeholder communication: Explain reliability and cost trade‑offs to technical and non‑technical audiences, and pivot your framing to whoever is in the room.
- Fitting in fast: Work within the client's existing platform standards and tooling where they’re sound, and make the case for change where they aren’t.
- Documentation & Handover
- Runbooks and decision records: Leave operational documentation, architecture decision records and Terraform the client's own engineers can maintain confidently.
- Knowledge transfer: Bring the client's engineers along in SRE practice as you go, so the reliability discipline holds after the engagement closes.
- Exit readiness: Treat a clean handover as part of the definition of done from the first sprint, not something arranged in the final fortnight.
Skills & Experience
- Core Technical Capabilities: We're looking for a mix of platform and SRE (50%), software engineering (30%), AI systems operations (20%).
- Experience: 5+ years in Dev Ops, SRE or platform engineering, including real production ownership of Kubernetes rather than consuming a cluster someone else runs.
- Infrastructure as code: Strong Terraform, and production experience on AWS. GCP or Azure alongside it is welcome.
- Python: Solid enough for real automation and operational tooling, with sound engineering practice around it.
- Pipelines: Built and maintained CI/CD with Git Hub Actions, Git Lab CI, Argo CD, Jenkins or similar.
- Observability: Hands‑on with Datadog, Prometheus, Grafana or Open Telemetry, and opinionated about what’s worth alerting on.
- Incident management: Calm, methodical instincts under pressure, and a habit of fixing the class of problem rather than the instance.
- AI and ML operations: Experience operating model serving, LLM inference or MLOps pipelines in production is a strong plus.
Familiarity with model APIs, vector databases, MCP or orchestration frameworks helps.
- Networking and security: VPC design, service mesh and zero-trust patterns.
- Domain & Sector Experience
- Regulated industries: Experience running platforms under strict compliance and data privacy controls, in i Gaming, financial services or banking, is highly advantageous.
- Scale: Comfort with high‑traffic, latency‑sensitive systems where peak load is a sporting fixture rather than a gradual curve.
- Contract and consulting delivery: A track record of taking on operational ownership of an unfamiliar estate quickly, and being trusted with it.
- Useful credentials
- Certified Kubernetes Administrator (CKA).
- AWS Solutions Architect Professional or Dev Ops Engineer Professional.
- Fin Ops Certified Practitioner (FOCP), which is a strong differentiator for this role.
- Terraform Associate.
- What we’ll offer you
- We trust people to do their best work.
- That means flexibility over rigid rules, impact over activity, and real investment in your growth both professionally and personally.
- You’ll be part of a supportive, and friendly culture, surrounded by smart, curious people who care deeply about what they do.
- We offer flexible working, including hybrid and remote options.
Our office hubs are located in Edinburgh, Leeds, Manchester, London and Bulgaria, with occasional travel to client sites or Create Future offices when needed.
- We trust you to manage your time balancing collaboration with client time and focused work. What matters is the impact you have, not how busy you look.
- Our hiring process
We try to keep our hiring process clear, fair and respectful of your time. We aim to get back to everyone who applies and we will be upfront about where you are in the process.
- Call with our Talent Acquisition Team
- Role specific capability interview
Depending on the role, we might also ask you to do a short presentation, a practical or technical task or have a values focused conversation.
We will explain what is involved before anything happens.
Inclusion at Create Future
We believe diverse teams build better workplaces and better products. We want Create Future to be a place where people feel able to be themselves and do their best work.
If you need any adjustments or support during the application process, just. We will do what we can to help.
We look forward to your application!
#J-18808-Ljbffr
Senior Cloud Engineer, AI Platform SRE Arbeitgeber: xdesign
CreateFuture ist ein hervorragender Arbeitgeber, der seinen Mitarbeitern die Möglichkeit bietet, an bedeutungsvollen Projekten zu arbeiten und dabei in einer unterstützenden und freundlichen Kultur zu gedeihen. Mit flexiblen Arbeitsmodellen, umfangreichen Weiterbildungsmöglichkeiten und einem klaren Fokus auf persönliche sowie berufliche Entwicklung, ist CreateFuture ein Ort, an dem Talente geschätzt werden und echte Auswirkungen erzielt werden können. Die Auszeichnung als Great Place to Work-Certified™ und die zahlreichen Benefits, wie 35 Tage Urlaub und private Krankenversicherung, machen das Unternehmen zu einem attraktiven Arbeitsplatz in der dynamischen Tech-Branche.
StudySmarter Expertenrat🤫
Wir sind der Meinung, dass du so Senior Cloud Engineer, AI Platform SRE erhalten könntest
✨Netzwerken in der IT-Community
In der IT-Consulting-Welt sollten wir regelmäßig auf Veranstaltungen wie Tech-Meetups oder Konferenzen gehen. Hier können wir nicht nur unser Netzwerk erweitern, sondern auch direkt mit potenziellen Arbeitgebern ins Gespräch kommen und unser Interesse an einer Vollzeitstelle zeigen.
✨Online-Foren und Gruppen nutzen
Sich in Online-Foren und Communities wie Stack Overflow oder LinkedIn-Gruppen umzusehen, kann uns helfen, Insider-Tipps zu erhalten und Informationen über offene Stellen in der IT-Beratung zu sammeln. Vergiss nicht, aktiv zu werden und Fragen zu stellen oder dein Wissen zu teilen – das erhöht unsere Sichtbarkeit!
✨Direkt bei xdesign bewerben
Viele Unternehmen, wie xdesign, stemmen ihre Vollzeitstellen bevorzugt über ihre eigenen Karriere-Webseiten. Also, lass uns regelmäßig auf deren Seite vorbeischauen und uns direkt bewerben, statt nur die üblichen Jobportale zu nutzen.
✨Überzeugende Projekte zeigen
Wir sollten unser Portfolio oder relevante Projekte gut sichtbar machen, egal ob das auf Github, persönlich oder auf LinkedIn ist. Bei IT-Consulting-Stellen kommt es oft auf praktische Erfahrungen an, also lass uns zeigen, was wir können!
Wir glauben, dass du diese Fähigkeiten brauchst, um Senior Cloud Engineer, AI Platform SRE mit Bravour zu bestehen
Einige Tipps für deine Bewerbung 🫡
Zeige deine technischen Skills!:In der IT-Beratung zählen deine technischen Kenntnisse und Fähigkeiten. Achte darauf, relevante Programmiersprachen, Tools und Systeme in deinem Lebenslauf aufzulisten. Zeig auch, wenn du Zertifikate hast, die deine Kompetenz unterstützen – das könnte dir einen echten Vorteil verschaffen!
Verstehe die Branche!:Unterstreiche in deinem Anschreiben, dass du ein gutes Verständnis für aktuelle Trends und Herausforderungen in der IT-Branche hast. Zeig, dass du nicht nur die technischen Aspekte beherrschst, sondern auch die Bedürfnisse der Kunden erkennen und lösen kannst!
Deine Projekte zählen!:Falls du bereits an IT-Projekten gearbeitet hast, verlinke diese oder beschreibe sie in deinem Lebenslauf. Praktische Erfahrungen – sei es in Form von Praktika oder privaten Projekten – sind besonders wertvoll in der IT-Beratung. Zeige uns, was du kannst!
Individuelle Bewerbung ist der Schlüssel!:Jede Bewerbung sollte individuell auf xdesign und die ausgeschriebene Position Senior Cloud Engineer, AI Platform SRE zugeschnitten sein. Teile uns mit, warum gerade du eine gute Wahl für unser Team bist. Das zeigt dein Engagement und deine Motivation, die über eine Standardbewerbung hinausgeht.
Wie man sich auf ein Vorstellungsgespräch bei xdesign vorbereitet
✨Technische Vorbereitung ist alles!
Da du dich auf eine Vollzeitstelle in der IT-Beratung bewirbst, solltest du dir wirklich einen Überblick über die wichtigsten Tools und Technologien verschaffen, die in der Branche verwendet werden. Sei bereit, technische Fragen zu beantworten, die sich auf Software-Architektur oder Systemintegration beziehen könnten.
✨Praxisbeispiele parat haben
In der IT-Beratung ist es wichtig, konkrete Beispiele aus deiner bisherigen Erfahrung zu bringen. Überlege dir Projekte, bei denen du erfolgreich einen Kunden beraten hast oder Herausforderungen gelöst hast. Das zeigt, dass du nicht nur theoretisches Wissen hast, sondern auch in der Praxis erfolgreich sein kannst.
✨Soft Skills betonen
Ein großer Teil der IT-Beratung ist die Kommunikation mit Kunden und das Verständnis ihrer Bedürfnisse. Bereite dich darauf vor, über deine zwischenmenschlichen Fähigkeiten zu sprechen, wie du mit herausfordernden Kunden umgehst oder wie du in Teams arbeitest. Das wird den Interviewern zeigen, dass du mehr als nur technisches Wissen mitbringst!
✨Fragen zum Unternehmen vorbereiten
Schau dir spezifisch die Projekte von xdesign an und überlege dir, welche Fragen du dazu stellen möchtest. Zeig Interesse an den aktuellen Herausforderungen, vor denen das Unternehmen steht, und wie du dazu beitragen könntest. Das hebt dich von anderen Bewerbern ab und zeigt, dass du wirklich motiviert bist.