Overview
A cloud platform supporting business critical services is reaching a scale where reliability can no longer depend on individual teams solving operational problems independently. The consultant will help introduce stronger reliability engineering practices across the platform, using service behaviour and production evidence to decide where engineering effort will have the greatest effect.
Responsibilities
- Define reliability objectives and operational standards for critical services.
- Analyse production incidents, latency patterns and capacity constraints.
- Improve observability and establish useful service level indicators.
- Work with development teams on resilience patterns, failure handling and capacity planning.
- Lead technical reviews following significant incidents and track resulting engineering actions.
Requirements
- 10+ years of software or infrastructure engineering experience.
- Strong background in distributed systems and cloud platforms.
- Proven SRE experience with observability, incident management and reliability engineering.
- Strong Kubernetes and infrastructure automation knowledge.
- Comfortable analysing production behaviour using metrics, logs and traces.
Expertise
Cloud Infrastructure
Distributed Systems
Incident Management
Kubernetes
Observability
SRE
