Key Notes
ML compiler, runtime and kernel optimisation sits at the centre of this Staff role, with targets such as NVIDIA and Qualcomm edge platforms. Constrained production-system experience and proficiency in at least one relevant toolchain are essential. The work spans model graphs, memory movement and in-vehicle compute, combining hands-on optimisation with benchmarking and technical direction for Wayve’s first driving product.
What You'll Work On
- Implement and validate compiler, runtime and kernel optimisations including operator fusion, execution scheduling, quantisation-aware performance and custom kernels.
- Profile model graphs, compilers, runtimes, kernel execution and memory movement to isolate measurable inference bottlenecks across the full stack.
- Build benchmarks and regression tests that verify performance gains across models, target devices and software releases.
- Optimise NVIDIA Orin or Thor and Qualcomm deployments while influencing model architecture, training and on-device deployment decisions.
Why This Role Matters
- Wayve’s first driving product needs transformer models to run efficiently on low-cost, low-power devices; compiler and kernel optimisation helps fit inference within in-vehicle compute constraints.
- Performance improvements must hold across models, devices and software releases; benchmarking and regression tests expose where an optimisation succeeds or introduces a loss on another target.
What They Are Looking For
- Experience: Improving performance in production systems under latency, memory, bandwidth, power, thermal or cost constraints is essential.
- Technical: Proficiency with TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, MLIR, ONNX or a comparable relevant toolchain is essential.
- Technical: Debugging, profiling, testing and maintainable code, with work spanning model behaviour and kernel or runtime execution are essential.
- Stakeholder: Aligning multiple stakeholders around performance trade-offs and priorities is essential.
- Preferred: Multi-target graph execution, embedded deployment, NVIDIA or Qualcomm SoCs, Python, C++ or mentoring are desirable.