Technical Disaster Recovery Plan Template
Five documents and five sheets, including a verification log keyed on restores attempted rather than on backup jobs that reported success.
Free download · No account needed
Every backup job is green. That is a fact about the jobs: a process ran, read some bytes and wrote them somewhere, and nothing in it ever tried to bring a system back. In the worked example here, 34 of 34 jobs were green, 11 of those 34 systems had ever had a restore attempted, and 5 of the 11 attempts failed. Not one failed on a bad backup. They failed on broken replication, a slow storage tier, a key held inside the blast radius, and two restores that came up on default settings.
Two of those five are documented behaviour rather than misconfiguration. AWS states that you cannot restore a snapshot into an existing instance and that the new one loads its data in the background while already reporting itself available. Recovery points drift the same way: point-in-time recovery ships its transaction logs every five minutes, so an objective of zero against it is unachievable rather than ambitious.
So the pack measures four things instead of restating a plan. What has actually been restored. What each mechanism's recovery floor really is, beside the number somebody wrote down. How long a recovery takes from detection rather than from the restore command, which at Halstow was 185 minutes against a stated 60. And whether the recovery order can be executed at all, because a plan sorted by business criticality starts with the service that needs five other things running first.
What's in the pack
Backup Verification Log
One row per restore attempt rather than per backup job, with the elapsed time, what failed, and whether the failure was in the data or around it.
Recovery Objective by Service
Stated recovery point, the mechanism behind it, and what that mechanism can actually deliver, as three separate columns that never overwrite each other.
Recovery Time Breakdown
Nine steps from detection to verification with minutes against each, so the restore stops standing in for the recovery it is a fifth of.
Recovery Order
The estate sorted by what each service needs at startup, beside the criticality order, with the disagreement expressed in positions moved.
Test Result History
Every exercise with its class named, because a tabletop and a real restore are both called tested and only one of them finds a broken replica.
Recovery Plan and Test Procedure
The plan organised by what is lost rather than what caused it, and the procedure for running a functional test that has not been quietly made easier.
Communication Plan
Who is told what during a recovery measured in hours, including the fallback channel that does not authenticate through the system being recovered.
How to use it
- 1
Open in River, or take it blank
Send the pack your service inventory and backup configuration inside River, or download the five documents and five sheets and work through them without an account.
- 2
Count restores, not backups
Four states per system: restored recently, restored once long ago, never restored, or no backup at all. The last two are where the work is.
- 3
Check every objective against its mechanism
The stated figure stays. The mechanism's floor goes next to it, and the gap between them is the finding rather than an error to correct.
- 4
Time one recovery end to end
From the moment a monitor would have fired, not from the restore command, with minutes against each step so the fixes have somewhere to land.
Frequently asked questions
Is this template free?
Yes, and the download needs no account and no card. Edit with AI is the optional half: it reads the inventory and backup configuration you send, works out what has never been restored, and checks each objective against its mechanism. Every other pack sits in the template library.
What format are the downloaded files?
Word (.docx) for the five documents and CSV (.csv) for the five sheets, zipped together, with nothing to convert on either side. The columns that matter are already there: each mechanism's floor beside the objective somebody stated, minutes against every recovery step, and the class of each test rather than just its date.
Our backups have been green for years. What is the risk?
That the green is describing the job rather than the data. Halstow's five restore failures were a replica that had stopped replicating, objects in a slow retrieval tier, a key inside the blast radius, and two default configurations. A backup job cannot see any of those.
How is this different from a business continuity plan?
Scope. A continuity plan covers processes, people and manual workarounds, and belongs with the business continuity pack. This one covers whether the systems come back: restore evidence, recovery-point floors, measured recovery times and an executable order.
We run an annual DR exercise. Does that count?
It depends which class it was. NIST SP 800-34 describes a tabletop as discussion-based only, not involving the deployment of equipment. Halstow's exercise covered all 41 services and four had ever had a real restore.
Why sort recovery by dependency instead of importance?
Because importance is not an order you can execute. Nine of the first twelve steps of Halstow's criticality-ordered plan blocked on something further down the list. Checkout sat at position one and could not start until the identity, secrets and DNS services at positions 16, 18 and 29 were up.
What else belongs alongside this?
The two things that decide whether you notice and whether you can act. A failure mode review finds what breaks and what would go undetected, and the SLO pack sets what counts as unavailable in the first place. Telling customers while it happens is the incident communication pack.
Find out which systems have ever come back
Take the Word documents and CSV sheets blank, or open this exact pack in River and send it the inventory and backup configuration you already have.
Edit with AI