SmartMobilityTalent

Senior SRE, AI Infrastructure

Wayve London, United KingdomHybridFull Time

Key Notes

A founding Cloud SRE remit covers Wayve’s Model Development Platform and multi-tenant GPU Compute environment for training, inference and on-road experimentation. Production Kubernetes, major-cloud, distributed-systems, Linux, automation and observability experience are essential. The London role has two office days each week and a 24/7 first-line on-call rotation, with responsibility for establishing reliability practices as the function develops.

What You'll Work On

  • Own reliability, availability and performance for the Model Development Platform and multi-tenant GPU Compute services used for training and inference.
  • Define SLOs, SLIs and error budgets, then improve capacity planning, scaling and efficiency across GPU-backed clusters.
  • Participate in the 24/7 on-call rotation, leading incident triage, escalation and root-cause analysis, turning post-incident findings into architectural or automation changes.
  • Design observability and deployment-safety systems while automating cluster operations, training workflows, remediation and scaling for production cloud services.

Why This Role Matters

  • Wayve’s Model Development Platform carries work from raw data to on-road experimentation; its availability and performance determine whether engineers can run that lifecycle predictably.
  • Wayve’s multi-tenant GPU Compute platform needs capacity planning and recovery automation to run large training and inference workloads efficiently as demand scales.

What They Are Looking For

  • Experience: Proven SRE, Production Engineer or Cloud Reliability experience supporting large-scale cloud systems is essential.
  • Technical: Operating production Kubernetes clusters and workloads in AWS, GCP or Azure is essential.
  • Technical: Production distributed-systems experience, Linux and Python, Go, C++ or another automation-oriented language are essential.
  • Technical: Troubleshooting networking, storage and performance, plus observability stacks such as Datadog, Prometheus, Grafana or OpenTelemetry, is essential.
  • Preferred: GPU-backed environments, large-scale ML infrastructure, production MLOps or a founding SRE remit is desirable.
  • Experience: Working with large compute clusters is essential; exposure to AI/ML training or inference workloads is strongly preferred.

Wayve

60 open jobs

Wayve develops end-to-end AI for self-driving vehicles, training neural networks to drive from camera data rather than detailed pre-mapped rules. The London company works with automakers including Nissan.

Sector
Connected & Automated Mobility
Headquarters
United Kingdom
Founded
2017
Operating Regions
Europe, North America, Asia-Pacific, Middle East & North Africa
Employees
1,001-5,000

Frequently Asked Questions

  • This role is Hybrid, based in London, United Kingdom.

  • Yes, this is a Full-time position.

Explore More Jobs

Discover more jobs from this company.

View All
View All