Site Reliability Engineer (GCP & Kubernetes)
London (Hybrid, 2 days onsite per week)
6-Month Contract | Inside IR35 | Competitive Day Rate
We are looking for an experienced Site Reliability Engineer (SRE) to join a high-performing technology team responsible for building, securing, and operating cloud-native platforms that support innovative AI-driven applications.
This is a key role focused on cloud infrastructure development, platform reliability, security hardening, and operational excellence. You will work closely with software engineers, platform teams, and technical stakeholders to design, deploy, and maintain resilient production systems capable of supporting complex and scalable workloads.
You will be responsible for the reliability, performance, scalability, and security of cloud-based applications and infrastructure, ensuring services are delivered efficiently and operate effectively in production environments.
Key Responsibilities:
Design, build, and maintain cloud infrastructure within Google Cloud Platform (GCP).
Deploy, manage, and optimise Kubernetes-based environments and containerised applications.
Develop and support backend services and operational tooling using Python.
Create and maintain Infrastructure as Code using Terraform.
Implement security hardening, monitoring, observability, and reliability best practices across cloud platforms.
Troubleshoot and resolve complex production issues, ensuring minimal service disruption.
Develop and improve CI/CD pipelines to enable efficient and reliable software delivery.
Work closely with development teams to improve platform resilience, scalability, and operational performance.
Implement monitoring, alerting, and incident response processes to support production systems.
Drive automation initiatives to reduce operational overhead and improve system reliability.We're Looking For:
Proven experience as a Site Reliability Engineer, Platform Engineer, Cloud Engineer, or DevOps Engineer.
Strong hands-on experience with Google Cloud Platform (GCP).
Experience deploying, operating, and supporting production cloud applications at scale.
Strong knowledge of Kubernetes and container orchestration technologies.
Commercial experience developing with Python.
Experience building and supporting backend services and APIs, ideally using FastAPI.
Strong knowledge of Terraform and Infrastructure as Code principles.
Experience designing highly available, resilient, and secure cloud environments.
Strong troubleshooting and incident management skills within production environments.
Experience with CI/CD pipelines, Git, GitHub, and DevOps best practices.
Strong understanding of cloud networking concepts, with particular emphasis on GCP networking.
Ability to take ownership of production systems and drive continuous improvement initiatives.Desirable:
Microsoft Azure experience.
Experience deploying and supporting AI/ML applications in production environments.
Exposure to Large Language Models (LLMs), NLP, or agent-based systems.
Experience with observability and monitoring tools such as Sentry, Prometheus, Grafana, or similar platforms.
Strong Docker and multi-container application architecture experience.
Experience working within scientific, pharmaceutical, genomics, or bioinformatics environments.
Start-up or high-growth technology company experience.
Open-source software contributions.
Technical leadership, mentoring, or coaching experience.Contract Details:
6-Month Initial Contract
Inside IR35
London-based
Hybrid Working (2 days onsite per week)
Competitive Day Rate
Potential extension opportunitiesThe Opportunity
This role offers the chance to work on modern cloud-native systems supporting innovative AI-enabled technologies. You'll be part of a collaborative engineering environment where reliability, automation, security, and scalability are key priorities, with the opportunity to make a significant impact on critical production platforms.
If you're a Site Reliability Engineer or Platform Engineer with strong GCP, Kubernetes, Terraform, and Python experience, we'd love to hear from you
Read Less