Information & Communications Technology

Major Incident Management and Problem Post-Mortems

Lead major incidents from detection to resolution and turn each one into a rigorous, blameless post-mortem that produces problem records and durable fixes rather than repeat outages.

Duration5 training days
Content4 modules · 8 sessions
On completionAccredited attendance certificate
About the programme

Course Overview

A major incident tests an organisation twice: once during the outage itself, and again in the days afterward when the temptation is to close the ticket and move on without understanding why it happened. This course trains you to run both halves well. You will classify severity accurately, take command of a major incident bridge, and communicate status to stakeholders without either understating impact or triggering unnecessary panic. Once service is restored, you will run a structured, blameless post-mortem using techniques such as the five whys and timeline reconstruction to separate the triggering event from the underlying systemic weakness. Sessions include live-simulated incident scenarios with deliberately incomplete information, plus practice writing problem records that lead to funded, tracked corrective actions rather than a report nobody implements. You finish able to run incidents that end in genuine prevention, not just apology.

Expected Learning Outcomes

01

Classify incident severity accurately using impact and urgency criteria rather than instinct or pressure from stakeholders.

02

Command a major incident bridge, coordinating technical responders while keeping the focus on restoration.

03

Communicate incident status to executives and customers in language calibrated to actual, not assumed, impact.

04

Reconstruct an accurate incident timeline from logs, chat transcripts and responder recollection.

05

Facilitate a blameless post-mortem that surfaces systemic contributing factors rather than individual fault.

06

Apply root cause techniques, including the five whys and contributing factor analysis, to separate trigger from cause.

07

Write problem records with corrective actions that are owned, resourced and tracked to closure.

Who Should Attend

01

IT operations staff who act or could act as incident commanders

02

Site reliability engineers responsible for service restoration during outages

03

Problem managers running root cause analysis after major incidents

04

Service desk and support leads coordinating stakeholder communication during incidents

05

Engineering managers accountable for implementing post-mortem corrective actions

06

IT service continuity staff building incident response playbooks

Course Modules

Select any module to see its sessions and points.

01

Classifying and Declaring Major Incidents

2 sessions · 8 points

Session 1Severity Classification and Escalation

  • Apply consistent impact and urgency criteria to classify an incident's severity within minutes of detection.
  • Distinguish a major incident requiring formal command from a routine incident handled through standard channels.
  • Escalate promptly when early signals suggest severity may increase, rather than waiting for confirmation.
  • Avoid both under-declaration, which delays resources, and over-declaration, which erodes trust in future alerts.

Session 2Standing Up the Incident Bridge

  • Assemble the right technical responders quickly using a pre-agreed on-call and escalation structure.
  • Assign clear roles, including incident commander, technical lead and communications lead, at the outset.
  • Establish a single source of truth for incident status to prevent conflicting updates across channels.
  • Set a cadence for status updates that keeps responders focused without constant interruption.
02

Commanding the Response

2 sessions · 8 points

Session 1Directing Technical Restoration

  • Keep the bridge focused on restoring service first, deferring full root cause investigation until stability returns.
  • Manage parallel workstreams when multiple hypotheses require simultaneous investigation.
  • Decide when to escalate to vendor support or invoke a disaster recovery plan rather than continue troubleshooting.
  • Recognise when a workaround is preferable to a full fix under time pressure, and document that decision.

Session 2Communicating with Stakeholders

  • Calibrate executive and customer communications to actual measured impact rather than worst-case assumptions.
  • Prepare holding statements that acknowledge an incident honestly without committing to an unverified resolution time.
  • Coordinate with customer-facing teams so external messaging remains consistent with internal status.
  • Close the incident with a clear all-resolved communication that includes a commitment to the post-mortem timeline.
03

Investigating Root Cause

2 sessions · 8 points

Session 1Reconstructing the Incident Timeline

  • Assemble an accurate timeline from monitoring data, chat logs and responder input before memory fades.
  • Identify the precise trigger event and distinguish it from earlier warning signs that were missed or ignored.
  • Note detection and response delays separately, since each points to a different improvement opportunity.
  • Cross-check the timeline against automated system logs to correct gaps in human recollection.

Session 2Applying Root Cause Techniques

  • Apply the five whys technique to move from the immediate trigger to the underlying systemic weakness.
  • Use contributing factor analysis to capture the several conditions that combined to allow the incident to occur.
  • Distinguish a genuine root cause from a superficial explanation that would not prevent recurrence if fixed alone.
  • Validate proposed root causes against the reconstructed timeline before finalising the post-mortem findings.
04

Running Blameless Post-Mortems and Driving Fixes

2 sessions · 8 points

Session 1Facilitating the Post-Mortem Meeting

  • Set explicit blameless ground rules at the start of the meeting to encourage honest disclosure of mistakes.
  • Separate discussion of what happened from discussion of who was involved to keep the focus systemic.
  • Draw out contributing factors from quieter participants who may hold key information they hesitate to share.
  • Close the meeting with a documented, prioritised list of corrective actions rather than open-ended discussion.

Session 2Writing Problem Records and Tracking Closure

  • Convert post-mortem findings into a formal problem record with a clearly assigned owner and target date.
  • Distinguish immediate corrective actions from longer-term structural fixes requiring separate project funding.
  • Track problem records to closure through existing governance forums rather than letting them lapse silently.
  • Review closed problem records periodically to confirm the corrective action actually prevented recurrence.

What the participant receives

4 course modules

A structured syllabus

8 training sessions

across 5 days

32 detailed points

Applied, detailed content

Accredited attendance certificate

On completing the programme

Complete your registration

We will contact you within one business day to confirm.

Ready to start?

Reserve your seat and start building the skill.

Enroll now

Share this course