Key Notes
End-to-end inference optimisation helps Wayve’s transformer models run on low-power edge accelerators and GPUs. Production-performance experience under latency, memory, bandwidth, power, thermal or cost constraints is essential, alongside proficiency in at least one relevant toolchain. The full-time London role spans model graphs, runtimes, kernels and in-vehicle compute under Wayve’s hybrid working policy.
What You'll Work On
- Profile model graphs, compilers, runtimes, kernel execution and memory movement to locate inference bottlenecks across the complete on-device stack.
- Implement and validate operator fusion, execution scheduling, quantisation-aware performance and custom-kernel optimisations for supported inference targets.
- Build benchmarking and regression tests across models, target devices and software releases for inference performance.
- Tune inference for NVIDIA Orin or Thor and Qualcomm targets, advising model developers on architecture and deployment choices.
Why This Role Matters
- Wayve’s first driving product must execute large transformer models on low-power edge hardware; inference optimisation connects model capability to feasible in-vehicle execution.
- Wayve’s NVIDIA and Qualcomm deployments are checked with cross-device benchmarks to detect when model, runtime or kernel changes fail to preserve performance gains.
What They Are Looking For
- Experience: Improving production-system performance under latency, memory, bandwidth, power, thermal or cost constraints is essential.
- Technical: Proficiency with TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL or a comparable relevant toolchain is essential.
- Technical: Debugging, profiling, testing and maintainable code, with work spanning model behaviour and kernel or runtime execution are essential.
- Stakeholder: Aligning multiple stakeholders around performance trade-offs and priorities is essential.
- Preferred: Edge deployment, NVIDIA or Qualcomm SoCs, Python, C++ and technical mentoring are desirable.