Key Notes
Production reliability across Trainline’s AWS platform is the focus of this £55,000–£63,000 hybrid role. The engineer joins incident response and on-call rotations, builds observability and improves alerts and post-incident actions. SRE concepts, Linux troubleshooting and scripting experience are listed, with Python preferred. The office requirement is at least 60% of working time over twelve weeks.
What You'll Work On
- Participate in Trainline production incidents, supporting investigation, mitigation, communication and coordinated service restoration across affected platform teams.
- Build observability from metrics, logs, events and traces using New Relic, ELK, Grafana and related ReliabilityOps tooling.
- Refine monitoring and alerting signals around business and customer impact, including alert noise and mean time to detection.
- Support AWS-hosted infrastructure and shared platform services through Terraform, GitHub Actions, Docker, ECS and well-structured reliability scripts.
- Contribute to post-incident reviews and follow-up actions and join the SRE on-call rotation for Trainline’s production platform.
Why This Role Matters
- Trainline’s production platform needs prompt detection and coordinated restoration during incidents; observability and on-call response make operational evidence available to the teams restoring affected services.
- Product engineering teams need services that are ready to operate safely; collaboration on resilience practices and AWS infrastructure brings reliability considerations into how those services are deployed.
What They Are Looking For
- Domain: Relevant domain knowledge includes SRE concepts including service-level indicators, service-level objectives and error budgets.
- Technical: Relevant technical skills include observability tools such as New Relic, Elastic, Influx or Grafana for detection and diagnosis.
- Technical: Relevant technical skills include cloud-provider experience, preferably AWS, plus Linux troubleshooting and scripting, preferably in Python.
- Domain: Relevant domain knowledge includes load balancing, reverse proxies, upstream health checks and application resilience patterns such as timeouts and circuit breakers.
- Technical: Relevant technical skills include time-series data management and build, deployment or configuration tooling such as GitHub Actions and Terraform.