← IncidentReady
When did you last test your backups? The 60-minute drill
For companies whose entire IT department is one person. A proactive drill, not a crisis response — run it before ransomware forces you to find out live. Print this.
Why this matters
A backup job that says "success" every night tells you a copy was written somewhere. It does not tell you the copy opens, restores in a usable timeframe, or survives an attacker who's also hunting for it. The only way to know is to actually restore something. Run this quarterly at minimum, and immediately after any backup software, storage target, or retention change.
Before you start
- Confirm your backup admin console login works, with restore rights (not just backup/read) — if you can't log in, that's already a finding.
- Pick a non-production location to restore into — a scratch VM, isolated folder, or spare machine. Never restore test data over live production.
- Have last night's backup job log open, and note the current time — you're about to measure restore duration.
Test 1: random-file restore (~15 min)
- Without looking at a list, name a real file from memory — an invoice, a contract, a spreadsheet someone actually uses. Don't cherry-pick an easy one.
- In your backup admin console, go to Restore (or Browse Backups), pick the source system, and select yesterday's or last week's restore point.
- Locate that file and restore it to the scratch location, not its original path. Open it — confirm no corruption error, content matches, modified date is sane.
- Repeat with one file from a different system — a mailbox item or SaaS file (OneDrive/SharePoint/Google Drive), not just local disk again.
Fail conditions: file not found, restore errors out, file is corrupted or truncated, or the wrong version comes back.
Test 2: full server/VM restore (~30–40 min)
This tests whether you can rebuild a whole box, not just grab a file — the scenario you'll actually be in during ransomware or a hardware failure.
- Pick one production server or VM that matters — a file server, domain controller, or line-of-business app server.
- Start a full restore of its latest clean backup point into an isolated test environment — a separate VLAN or air-gapped host with no route to production. Never restore an unverified image directly onto your network.
- Boot it. Confirm it powers on, logs in, and the expected service is running (SQL, file shares, the app), with a test client able to connect on the isolated segment.
- Confirm the data matches the expected point in time — check a timestamp or database row you know should be there. Tear down the test instance when done.
Fail conditions: image won't boot, restore hangs or errors, restored server is missing services or data, or required manual steps nobody had documented.
Measure your real restore time
Write down what your business can actually tolerate, then measure against it — a "successful" restore that takes 14 hours when the business can only absorb 4 is still a finding.
- RTO (Recovery Time Objective): how long from "we need to restore" to "system is back and usable" — e.g., target 4 hours.
- RPO (Recovery Point Objective): how old the restore point is, i.e. how much data you can afford to lose — e.g., target 15 minutes or nightly.
If measured time exceeds target, that's a serious finding minimum. Root causes are usually bandwidth to the backup target, the tier's restore throughput, or too much manual setup between "image restored" and "usable." Fix the bottleneck, don't just note it.
3-2-1 audit, plus the immutable copy
Confirm the rule is actually true today, not just configured once and forgotten:
- 3 copies — production data plus at least 2 backup copies. Count them; "the backup job says success" is not counting them.
- 2 different media/storage types — e.g., local NAS snapshot and cloud backup target, not two copies on the same storage array.
- 1 copy offsite — a different building or cloud region/account from production, and at least one copy immutable or offline/air-gapped — the copy ransomware can't touch. Confirm in the vendor console directly ("immutability," "object lock," "WORM," or "air-gap" status); don't assume it's on because you enabled it once.
- The credentials that manage production backups are not the same credentials that can delete the immutable copy — one login able to nuke both means you have one copy, not two.
- Every system that should be in scope actually is — cross-check against your current asset list; new servers and SaaS tenants are the most common gap. Retention meets policy and old restore points are confirmed still recoverable, not just still listed. If using tape or offline rotation, confirm the current copy was actually rotated recently — a "cold" copy nobody's touched in 8 months is a critical finding.
What a failing grade means
Critical findings (can't restore at all, or the only copy is one an attacker could also encrypt/delete) get treated like an open incident — assign an owner, fix within 5 business days, re-run the specific failed test to confirm. Serious findings (RTO broken, or a system was never in scope) get 30 days. Minor findings (works, but process is sloppy) wait for the next scheduled drill. File the completed scorecard where your cyber insurer's application can find it — "we test restores quarterly" is a question on most renewal applications, and this drill is your evidence.
This checklist is 1 of 10 runbooks.
The full IncidentReady pack includes the complete drill scorecard, drill-schedule template, escalation email to your backup vendor, and the failure-to-finding severity table — plus ransomware, phishing/BEC, compromised accounts, lost devices, and a tabletop kit to rehearse it all — written for businesses of 5–50 people. Editable, $29.
Get the pack — $29 Read the free phishing runbook →
Preparedness templates, not legal advice. © 2026 IncidentReady