Key Notes
A founding Cloud SRE remit covers Wayve’s Model Development Platform and multi-tenant GPU Compute environment for training, inference and on-road experimentation. Production Kubernetes, major-cloud, distributed-systems, Linux, automation and observability experience are essential. The London role has two office days each week and a 24/7 first-line on-call rotation, with responsibility for establishing reliability practices as the function develops.
What You'll Work On
- Own reliability, availability and performance for the Model Development Platform and multi-tenant GPU Compute services used for training and inference.
- Define SLOs, SLIs and error budgets, then improve capacity planning, scaling and efficiency across GPU-backed clusters.
- Participate in the 24/7 on-call rotation, leading incident triage, escalation and root-cause analysis, turning post-incident findings into architectural or automation changes.
- Design observability and deployment-safety systems while automating cluster operations, training workflows, remediation and scaling for production cloud services.
Why This Role Matters
- Wayve’s Model Development Platform carries work from raw data to on-road experimentation; its availability and performance determine whether engineers can run that lifecycle predictably.
- Wayve’s multi-tenant GPU Compute platform needs capacity planning and recovery automation to run large training and inference workloads efficiently as demand scales.
What They Are Looking For
- Experience: Proven SRE, Production Engineer or Cloud Reliability experience supporting large-scale cloud systems is essential.
- Technical: Operating production Kubernetes clusters and workloads in AWS, GCP or Azure is essential.
- Technical: Production distributed-systems experience, Linux and Python, Go, C++ or another automation-oriented language are essential.
- Technical: Troubleshooting networking, storage and performance, plus observability stacks such as Datadog, Prometheus, Grafana or OpenTelemetry, is essential.
- Preferred: GPU-backed environments, large-scale ML infrastructure, production MLOps or a founding SRE remit is desirable.
- Experience: Working with large compute clusters is essential; exposure to AI/ML training or inference workloads is strongly preferred.