Site Reliability Engineer (AWS)
This role involves ensuring the reliability and resilience of large-scale, cloud-native platforms running on AWS and Kubernetes, supporting critical national services. You'll respond to incidents across a 24/7 on-call rota, diagnose complex production issues, and implement automation to reduce operational toil. The focus is on proactive system improvement, observability, and collaboration with engineering teams to maintain high availability and security in a regulated environment.