Key Notes
Large-scale ML infrastructure ownership covers multi-billion-parameter foundation models, accelerator profiling and distributed training for Waymo simulation. The role lists five years of software engineering, including three years in machine-learning infrastructure, plus a relevant bachelor’s degree or equivalent practical experience. The London position is on-site and publishes a base salary of £155,000–£163,000.
What You'll Work On
- Define architecture and technical direction for large autonomous-vehicle foundation models spanning data engineering, model development and deployment.
- Design large distributed systems for planet-scale dataset generation and model training across the complete machine-learning lifecycle.
- Translate Simulator product goals into measurable performance and system requirements, aligning components across large machine-learning platforms.
- Collaborate with Google DeepMind, Waymo Realism Modeling and Oxford teams while mentoring engineers working on simulation infrastructure.
Why This Role Matters
- Waymo’s Simulator relies on scalable training infrastructure to model realistic agents, roads, traffic controls and weather with multi-billion-parameter foundation models.
- Waymo’s Simulator engineers use distributed data-generation and training systems to accelerate model development, while agreed performance requirements keep infrastructure work aligned with product goals.
What They Are Looking For
- Qualification: Relevant qualifications include a bachelor’s degree in computer science, robotics or a similar field, or equivalent practical experience.
- Experience: Five or more years of software engineering, including at least three years in machine-learning infrastructure, is sought.
- Preferred: Experience with DeepSpeed, PyTorch, TensorFlow or comparable machine-learning infrastructure tools is preferred.
- Preferred: Distributed training, gradient sharding and accelerator profiling expertise is preferred.
- Preferred: Autonomous-driving, simulation or machine-learning accelerator familiarity is preferred.
- Preferred: A master’s degree or equivalent experience, and ten or more years in software engineering including at least five in ML infrastructure, are preferred.