Define service level indicators and set realistic service level objectives for critical user journeys.
Site Reliability Engineering and Error Budget Management
Builds service level objectives, error budgets and incident response practice so engineering teams balance reliability against delivery speed.
Course Overview
Modern digital services fail in ways that simple uptime checks never catch: a page that loads but times out for one request in twenty, or a payment API that stays reachable while silently dropping transactions. This course builds practical site reliability engineering skills around one core discipline: error budget management. Participants learn to define service level indicators and objectives for the user journeys that matter, convert those objectives into error budgets that give engineering and product teams a shared numeric language for risk, and build the alerting, incident command and postmortem practices that keep those budgets meaningful rather than theoretical. Teaching combines worked examples from real service architectures with applied exercises: calculating burn rates from sample traffic data, running a simulated major incident, and drafting a production readiness review for a fictional release. Participants leave with templates for error budget policies, postmortem reports and reliability scorecards, along with the judgement to negotiate reliability trade-offs directly with engineering and product leadership.
Expected Learning Outcomes
Calculate error budgets and design burn-rate alerts that trigger before customers notice degradation.
Run blameless postmortems that convert incidents into tracked, verifiable corrective actions.
Prioritise a toil reduction backlog and automate the highest-cost manual operational tasks.
Design chaos engineering experiments and game days that expose hidden failure dependencies.
Build production readiness review checklists that gate risky releases before launch.
Report reliability performance to leadership using error budget and DORA metric dashboards.
Who Should Attend
Site reliability engineers moving from reactive firefighting to structured reliability practice.
Platform and DevOps engineers responsible for production infrastructure uptime.
Software engineering managers who must balance feature velocity against reliability targets.
Incident commanders and on-call leads coordinating major incident response.
Cloud infrastructure architects designing resilient, observable distributed systems.
IT operations leaders introducing site reliability practice into a traditional operations team.
Course Modules
Select any module to see its sessions and points.
01Service Level Foundations and Error Budget Design
2 sessions · 8 points
Session 1Defining Service Level Indicators and Objectives
- Map critical user journeys to the technical signals that most affect customer-perceived reliability.
- Select service level indicators for availability, latency, throughput and correctness per service tier.
- Set service level objective targets using historical performance data rather than arbitrary round numbers.
- Differentiate service level objectives from contractual service level agreements and their penalty clauses.
Session 2Constructing and Governing Error Budgets
- Calculate error budgets from agreed service level objectives over rolling measurement windows.
- Write error budget policies that freeze non-essential releases once the budget is exhausted.
- Configure multi-window, multi-burn-rate alerts that distinguish fast outages from slow degradation.
- Negotiate error budget trade-offs between product feature teams and reliability stakeholders.
02Observability and Incident Response Engineering
2 sessions · 8 points
Session 1Building Multi-Signal Observability Pipelines
- Instrument services with distributed tracing to follow requests across microservice boundaries.
- Correlate structured logs, metrics and traces using shared identifiers for faster root cause analysis.
- Control metric cardinality and retention costs while preserving diagnostic value in dashboards.
- Design service-tier dashboards that surface the signals on-call engineers need within seconds.
Session 2Incident Command and Blameless Postmortems
- Classify incident severity using impact and urgency criteria that trigger the right response tier.
- Assign incident commander, communications and operations roles during a live major incident.
- Facilitate blameless postmortems that separate contributing factors from individual blame.
- Track postmortem corrective actions to closure and verify they prevent repeat incidents.
03Reducing Toil and Engineering for Resilience
2 sessions · 8 points
Session 1Toil Identification and Automation Roadmaps
- Audit recurring operational tasks to separate genuine engineering work from manual toil.
- Score toil items by frequency and effort to build a prioritised automation backlog.
- Convert frequently used manual runbooks into self-service tools and automated remediation.
- Measure the time reclaimed by automation and reinvest it in reliability engineering work.
Session 2Chaos Engineering and Capacity Planning
- Design controlled fault-injection experiments that test resilience without harming customers.
- Plan and facilitate game days that rehearse team response to simulated infrastructure failure.
- Forecast capacity needs from historical load growth ahead of predictable demand peaks.
- Map upstream and downstream service dependencies to expose single points of failure.
04Scaling Reliability Practice Across the Organisation
2 sessions · 8 points
Session 1Production Readiness Reviews and Release Engineering
- Apply a production readiness review checklist covering monitoring, rollback and capacity before launch.
- Compare canary, blue-green and progressive rollout strategies for risk-controlled releases.
- Automate rollback triggers tied to error budget consumption and anomaly detection.
- Gate deployment pipelines so releases pause automatically when reliability thresholds are breached.
Session 2Reliability Team Structure and Executive Reporting
- Compare embedded, centralised and platform-based models for organising reliability engineers.
- Build reliability scorecards that translate error budgets into language executives act on.
- Track deployment frequency, lead time, change failure rate and recovery time as delivery health signals.
- Present a reliability investment case that ties automation spend to reduced incident cost.
What the participant receives
4 course modules
A structured syllabus
8 training sessions
across 5 days
32 detailed points
Applied, detailed content
Accredited attendance certificate
On completing the programme
Complete your registration
We will contact you within one business day to confirm.
Ready to start?
Reserve your seat and start building the skill.
