Information & Communications Technology

Cloud Disaster Recovery Runbooks and Failover Testing

Builds recovery tiering, orchestrated runbooks and tested failover procedures so cloud applications recover within agreed time and data-loss limits.

Duration5 training days
Content4 modules · 8 sessions
On completionAccredited attendance certificate
About the programme

Course Overview

A disaster recovery plan that has never been tested is a guess, and cloud outages tend to arrive exactly when that guess is wrong. This course teaches disaster recovery design and failover testing for cloud applications: setting recovery time and recovery point objectives by business impact, writing orchestrated runbooks, and proving recovery capability through scheduled live tests rather than paper exercises. Participants learn to select a recovery strategy per application tier, automate failover steps such as DNS rerouting and database promotion, and design test scenarios that validate real recovery time without unnecessary customer impact. Exercises include building a tiered recovery objective model, drafting an orchestrated runbook for a multi-tier application, and planning a live failover test with defined success criteria and rollback triggers. Participants finish with runbook templates, a failover test plan format and an audit evidence structure aligned to business continuity requirements.

Expected Learning Outcomes

01

Classify applications into recovery tiers and set recovery time and recovery point objectives.

02

Select a disaster recovery strategy, from backup-restore to active-active, matched to business impact.

03

Write orchestrated recovery runbooks that sequence infrastructure, data and application restoration.

04

Automate failover steps including DNS routing and database promotion with safeguards against false triggers.

05

Design and execute live failover tests that measure actual recovery time against target objectives.

06

Validate data integrity and application function after failover before declaring recovery complete.

07

Produce audit evidence mapping disaster recovery testing to business continuity requirements.

Who Should Attend

01

Cloud infrastructure engineers responsible for backup, replication and failover configuration.

02

Business continuity managers who must evidence recovery capability to auditors and regulators.

03

Site reliability engineers writing and maintaining disaster recovery runbooks.

04

IT leaders deciding recovery tier and standby infrastructure investment per application.

05

Database administrators managing cross-region replication and recovery point objectives.

06

Incident and crisis managers coordinating communication during a live failover event.

Course Modules

Select any module to see its sessions and points.

01

Disaster Recovery Strategy and Objectives

2 sessions · 8 points

Session 1Setting Recovery Objectives by Application Tier

  • Classify applications into criticality tiers linked to acceptable recovery time and recovery point objectives.
  • Calculate recovery time and recovery point objectives from business impact analysis, not technical convenience.
  • Select a disaster recovery strategy, from backup-restore to active-active, appropriate to each tier.
  • Cost disaster recovery options against the business impact of extended downtime for that tier.

Session 2Aligning Backup Design to Recovery Requirements

  • Design backup schedules and retention periods that meet each tier's recovery point objective.
  • Implement immutable or air-gapped backup copies resilient to ransomware encryption attempts.
  • Validate that database replication lag stays within the agreed recovery point objective.
  • Document data dependencies so restores happen in the correct order during recovery.
02

Writing Disaster Recovery Runbooks

2 sessions · 8 points

Session 1Structuring Runbooks for Orchestrated Recovery

  • Sequence recovery steps by dependency so infrastructure, data and application layers restore in order.
  • Write decision points into runbooks for scenarios involving partial rather than total outage.
  • Assign clear roles and escalation contacts to each step within the recovery runbook.
  • Version runbooks alongside the infrastructure and application changes they describe.

Session 2Automating Recovery Steps

  • Convert manual runbook steps into scripts or infrastructure as code that redeploy environments on demand.
  • Automate DNS failover using health checks and predefined failover routing policies.
  • Integrate database promotion and replica failover into the automated recovery sequence.
  • Build safeguards that prevent automated failover from triggering on false positive signals.
03

Failover Testing and Validation

2 sessions · 8 points

Session 1Designing Failover Test Scenarios

  • Plan tabletop exercises that rehearse decision-making before committing to a live failover test.
  • Design live failover tests that validate actual recovery time against the target objective.
  • Scope tests to limit customer impact while still exercising genuine production infrastructure.
  • Define success criteria and rollback triggers in writing before any test begins.

Session 2Executing and Learning from Recovery Tests

  • Run scheduled failover tests and record the actual recovery time and recovery point achieved.
  • Validate application function and data integrity after failover before declaring success.
  • Capture gaps between documented runbooks and what actually happened during the test.
  • Update runbooks and retest until recovery time consistently meets the stated objective.
04

Governance, Compliance and Continuous Improvement

2 sessions · 8 points

Session 1Meeting Business Continuity Requirements

  • Map disaster recovery test evidence to business continuity standards such as ISO 22301 audit requirements.
  • Maintain a test history log that demonstrates recovery capability to auditors and regulators.
  • Review third-party and cloud provider dependencies for their own disaster recovery commitments.
  • Report residual recovery risk to leadership in business rather than purely technical terms.

Session 2Sustaining Readiness Across a Changing Estate

  • Trigger runbook reviews whenever architecture, dependencies or system ownership change materially.
  • Schedule recurring failover tests so recovery readiness does not decay between major incidents.
  • Track disaster recovery cost against standby infrastructure spend to justify tier decisions.
  • Build a communication plan that keeps stakeholders informed during an actual disaster.

What the participant receives

4 course modules

A structured syllabus

8 training sessions

across 5 days

32 detailed points

Applied, detailed content

Accredited attendance certificate

On completing the programme

Complete your registration

We will contact you within one business day to confirm.

Ready to start?

Reserve your seat and start building the skill.

Enroll now

Share this course