Key Responsibilities
- Own the reliability, availability, performance, and operational health of Kubernetes/OpenShift platforms.
- Participate in incident response, on-call support, troubleshooting, RCA, and post-incident improvements.
- Support platform upgrades, component releases, infrastructure changes, and production deployments.
- Drive automation to reduce operational toil and improve scalability and resilience.
- Develop and maintain scripts, tooling, dashboards, runbooks, and operational documentation.
- Manage cluster lifecycle activities, including upgrades, capacity planning, performance optimisation, and production readiness.
- Lead vulnerability management, CVE assessment, prioritisation, patching, and remediation across Kubernetes, OpenShift, Linux, and supporting components.
- Improve platform observability, monitoring, metrics, logging, and alerting.
- Collaborate with global SRE, Infrastructure, Security, and L2 Operations teams.
- Contribute to platform modernisation, self-service, automation, reliability, and customer experience initiatives.
Required Skills & Experience
- Experience in SRE, Platform Engineering, DevOps, Systems Engineering, or Infrastructure Operations.
- Strong hands-on experience with Red Hat OpenShift or enterprise Kubernetes in production.
- Strong Linux administration and troubleshooting skills.
- Deep understanding of Kubernetes architecture, networking, scheduling, storage, container orchestration, and cluster lifecycle management.
- Strong understanding of SRE principles, incident management, observability, automation, capacity planning, reliability, and continuous improvement.
- Excellent communication, collaboration, and problem-solving skills





