Disaster recovery planning made practical: set RTO and RPO targets, match recovery strategies, plan for ransomware and test until it works.
Outages happen — from hardware failures and cloud region incidents to ransomware and human error. Disaster recovery (DR) planning makes sure the business can restore critical services within agreed limits.
Start with the business
A business impact analysis identifies which services matter most and what downtime costs. For each critical service, agree two targets with its business owner:
- Recovery time objective (RTO): how long the service can be unavailable.
- Recovery point objective (RPO): how much data loss, measured in time, is acceptable.
Match strategies to targets
| Strategy | Typical fit |
|---|---|
| Backup and restore | Longer RTO/RPO, lowest cost |
| Pilot light / warm standby | Moderate RTO, core systems pre-provisioned |
| Active-active across sites or regions | Near-zero downtime, highest cost |
Plan for ransomware
Traditional DR assumes the backup is clean. Ransomware breaks that assumption, so keep immutable or offline copies and plan for rebuilding in an isolated environment. Our whitepaper Building Ransomware-Resilient Backup covers this.
Test, then test again
- Tabletop exercises to walk through roles and decisions.
- Technical restore tests for individual systems.
- Full failover tests for the most critical services.
- Update runbooks after every test and every major change.
6 best practices for effective disaster recovery planning
- Start with a business impact analysis. Interview business owners to understand which processes matter most and the cost of downtime and data loss for each.
- Tier your applications. Group systems into recovery tiers with consistent RTO and RPO targets, so investment matches business value.
- Document dependencies. Applications rely on identity services, DNS, networks and databases. Recovery order must reflect these dependencies.
- Keep runbooks current. Step-by-step recovery instructions should be stored where they remain accessible during an outage and updated after every change.
- Test realistically. Combine tabletop exercises with technical failover tests, and measure actual recovery times against targets.
- Plan communications. Define how staff, customers, regulators and partners will be informed during a disaster.
Disaster recovery versus business continuity
Disaster recovery focuses on restoring IT systems and data. Business continuity covers how the whole organisation keeps operating, including people, facilities, suppliers and manual workarounds. Both plans should be aligned and tested together.
Common mistakes to avoid
- Setting aggressive RTO targets without funding the technology to meet them.
- Storing the recovery plan only on systems that may be unavailable.
- Testing only easy components rather than full application recovery.
- Overlooking SaaS applications and third-party dependencies.
Frequently asked questions
What is the difference between RTO and RPO?
RTO is how quickly a service must be restored. RPO is how much data loss, measured in time, is acceptable.
Is cloud-based disaster recovery cheaper?
Disaster recovery as a service can reduce costs by avoiding a second data centre, but test performance and data transfer costs carefully.
A 90-day action plan
Days 1 to 30: run business impact interviews for the top ten processes, agree recovery targets with business owners and map the systems and suppliers behind each process.
Days 31 to 60: compare current capabilities with those targets, close the largest gaps first and write runbooks for the most critical applications.
Days 61 to 90: hold a tabletop exercise with IT and business leaders, then perform a technical failover test for one critical service and record lessons learned.
Questions to ask recovery service providers
- What recovery times have customers similar to us achieved in tests?
- How often can we test, and is testing included in the price?
- How are recovery environments isolated from compromised production systems?
- Where are recovery sites located, and how are they secured?
- What support is provided during a real declared disaster?
Key terms explained
- Business impact analysis: a study of how disruption affects operations, finances and customers.
- Failover: switching a service to a secondary system or site.
- Failback: returning service to the primary site after recovery.
- Runbook: step-by-step instructions for a specific recovery task.
- Warm standby: a partly running secondary environment that can be scaled up quickly.
The bottom line
Recovery capability is proven only when it is tested. Understanding business impact, setting realistic targets, documenting dependencies, maintaining runbooks and rehearsing regularly turn a document into a dependable capability. Include ransomware scenarios, SaaS dependencies and communication plans, and fund the technology needed to meet agreed targets. Regular exercises with both IT and business teams build the confidence and muscle memory needed when a real disruption occurs.
Further reading on disaster recovery planning
For authoritative, vendor-neutral guidance on disaster recovery planning, see the NIST Cybersecurity Framework. You can also browse our free whitepapers.

