Classify incident severity accurately using impact and urgency criteria rather than instinct or pressure from stakeholders.
Major Incident Management and Problem Post-Mortems
Lead major incidents from detection to resolution and turn each one into a rigorous, blameless post-mortem that produces problem records and durable fixes rather than repeat outages.
Course Overview
A major incident tests an organisation twice: once during the outage itself, and again in the days afterward when the temptation is to close the ticket and move on without understanding why it happened. This course trains you to run both halves well. You will classify severity accurately, take command of a major incident bridge, and communicate status to stakeholders without either understating impact or triggering unnecessary panic. Once service is restored, you will run a structured, blameless post-mortem using techniques such as the five whys and timeline reconstruction to separate the triggering event from the underlying systemic weakness. Sessions include live-simulated incident scenarios with deliberately incomplete information, plus practice writing problem records that lead to funded, tracked corrective actions rather than a report nobody implements. You finish able to run incidents that end in genuine prevention, not just apology.
Expected Learning Outcomes
Command a major incident bridge, coordinating technical responders while keeping the focus on restoration.
Communicate incident status to executives and customers in language calibrated to actual, not assumed, impact.
Reconstruct an accurate incident timeline from logs, chat transcripts and responder recollection.
Facilitate a blameless post-mortem that surfaces systemic contributing factors rather than individual fault.
Apply root cause techniques, including the five whys and contributing factor analysis, to separate trigger from cause.
Write problem records with corrective actions that are owned, resourced and tracked to closure.
Who Should Attend
IT operations staff who act or could act as incident commanders
Site reliability engineers responsible for service restoration during outages
Problem managers running root cause analysis after major incidents
Service desk and support leads coordinating stakeholder communication during incidents
Engineering managers accountable for implementing post-mortem corrective actions
IT service continuity staff building incident response playbooks
Course Modules
Select any module to see its sessions and points.
01Classifying and Declaring Major Incidents
2 sessions · 8 points
Session 1Severity Classification and Escalation
- Apply consistent impact and urgency criteria to classify an incident's severity within minutes of detection.
- Distinguish a major incident requiring formal command from a routine incident handled through standard channels.
- Escalate promptly when early signals suggest severity may increase, rather than waiting for confirmation.
- Avoid both under-declaration, which delays resources, and over-declaration, which erodes trust in future alerts.
Session 2Standing Up the Incident Bridge
- Assemble the right technical responders quickly using a pre-agreed on-call and escalation structure.
- Assign clear roles, including incident commander, technical lead and communications lead, at the outset.
- Establish a single source of truth for incident status to prevent conflicting updates across channels.
- Set a cadence for status updates that keeps responders focused without constant interruption.
02Commanding the Response
2 sessions · 8 points
Session 1Directing Technical Restoration
- Keep the bridge focused on restoring service first, deferring full root cause investigation until stability returns.
- Manage parallel workstreams when multiple hypotheses require simultaneous investigation.
- Decide when to escalate to vendor support or invoke a disaster recovery plan rather than continue troubleshooting.
- Recognise when a workaround is preferable to a full fix under time pressure, and document that decision.
Session 2Communicating with Stakeholders
- Calibrate executive and customer communications to actual measured impact rather than worst-case assumptions.
- Prepare holding statements that acknowledge an incident honestly without committing to an unverified resolution time.
- Coordinate with customer-facing teams so external messaging remains consistent with internal status.
- Close the incident with a clear all-resolved communication that includes a commitment to the post-mortem timeline.
03Investigating Root Cause
2 sessions · 8 points
Session 1Reconstructing the Incident Timeline
- Assemble an accurate timeline from monitoring data, chat logs and responder input before memory fades.
- Identify the precise trigger event and distinguish it from earlier warning signs that were missed or ignored.
- Note detection and response delays separately, since each points to a different improvement opportunity.
- Cross-check the timeline against automated system logs to correct gaps in human recollection.
Session 2Applying Root Cause Techniques
- Apply the five whys technique to move from the immediate trigger to the underlying systemic weakness.
- Use contributing factor analysis to capture the several conditions that combined to allow the incident to occur.
- Distinguish a genuine root cause from a superficial explanation that would not prevent recurrence if fixed alone.
- Validate proposed root causes against the reconstructed timeline before finalising the post-mortem findings.
04Running Blameless Post-Mortems and Driving Fixes
2 sessions · 8 points
Session 1Facilitating the Post-Mortem Meeting
- Set explicit blameless ground rules at the start of the meeting to encourage honest disclosure of mistakes.
- Separate discussion of what happened from discussion of who was involved to keep the focus systemic.
- Draw out contributing factors from quieter participants who may hold key information they hesitate to share.
- Close the meeting with a documented, prioritised list of corrective actions rather than open-ended discussion.
Session 2Writing Problem Records and Tracking Closure
- Convert post-mortem findings into a formal problem record with a clearly assigned owner and target date.
- Distinguish immediate corrective actions from longer-term structural fixes requiring separate project funding.
- Track problem records to closure through existing governance forums rather than letting them lapse silently.
- Review closed problem records periodically to confirm the corrective action actually prevented recurrence.
What the participant receives
4 course modules
A structured syllabus
8 training sessions
across 5 days
32 detailed points
Applied, detailed content
Accredited attendance certificate
On completing the programme
Complete your registration
We will contact you within one business day to confirm.
Ready to start?
Reserve your seat and start building the skill.
