SmartMobilityTalent

Staff SRE, AI Infrastructure

Wayve London, United KingdomHybridFull Time

Key Notes

A founding Staff SRE remit establishes reliability foundations for Wayve’s Model Development Platform and multi-tenant GPU Compute environment. Essential experience includes GPU-backed environments or large-scale ML infrastructure, production MLOps, Kubernetes, a major cloud provider, distributed systems, Linux, automation and observability. The hybrid London role requires two office days each week and participation in a 24/7 on-call rotation.

What You'll Work On

  • Establish reliability, availability and performance standards for the Model Development Platform and GPU Compute services as a founding Staff SRE.
  • Set SLOs, SLIs and error budgets, improving capacity planning, cluster scaling and efficiency across multi-tenant GPU fleets.
  • Direct incident triage, escalation and root-cause analysis during the 24/7 rotation, converting findings into durable platform changes.
  • Architect observability, production-readiness, deployment-safety and automation practices with ML, platform and software teams across large-scale cloud services.

Why This Role Matters

  • Wayve’s Model Development Platform supports the route from raw data to on-road experimentation; Staff-level reliability standards shape how consistently that route operates.
  • Wayve’s GPU Compute platform uses capacity planning, automation and incident recovery to keep shared training and inference workloads efficient as large compute clusters scale.

What They Are Looking For

  • Experience: The Staff role requires proven SRE, Production Engineer or Cloud Reliability experience supporting large-scale cloud systems.
  • Technical: Operating GPU-backed environments or large-scale ML infrastructure, plus production training or inference pipelines, is essential.
  • Technical: Production Kubernetes and workloads in AWS, GCP or Azure are essential.
  • Technical: Linux, proficiency in a scripting or systems language such as Python, Go or C++, distributed-systems troubleshooting and observability are essential.
  • Preferred: Infrastructure as code, multi-team SLO programmes or founding an SRE function are desirable.

Wayve

60 open jobs

Wayve develops end-to-end AI for self-driving vehicles, training neural networks to drive from camera data rather than detailed pre-mapped rules. The London company works with automakers including Nissan.

Sector
Connected & Automated Mobility
Headquarters
United Kingdom
Founded
2017
Operating Regions
Europe, North America, Asia-Pacific, Middle East & North Africa
Employees
1,001-5,000

Frequently Asked Questions

  • This role is Hybrid, based in London, United Kingdom.

  • Yes, this is a Full-time position.

Explore More Jobs

Discover more jobs from this company.

View All
View All