Template

Disaster Recovery Plan Template: Example, Format and Recovery Runbook

A deep DR plan template with architecture, dependencies, backup and cyber recovery, a 12-step runbook, 30 review questions, technical and business validation, failback and test evidence.

A disaster recovery plan (DRP) turns technology recovery requirements into a controlled sequence for restoring infrastructure, applications, data, integrations and business service. It should not be a list of server names. A credible DRP connects the business RTO/RPO to architecture, recovery dependencies, cyber/integrity decisions, runbooks, technical validation, business validation, communications, evidence and failback.

DR plan structure: 16 core sections

#SectionRequired outcome
1Scope and service mappingKnow which business services, applications, environments and locations are covered.
2Recovery objectivesRecord business-approved RTO/RPO and any stricter infrastructure sequencing targets.
3Architecture and dependency mapExpose identity, DNS, network, storage, secrets, databases, middleware, APIs, SaaS and supplier dependencies.
4Roles and authorityDefine incident, cyber, infrastructure, application, database and business-validation roles.
5Activation and recovery modeDefine when to restore, fail over, rebuild cleanly or wait for containment evidence.
6Backup and replication designKnow what is protected, frequency, immutability/isolation, retention, monitoring and restoration evidence.
7Recovery environment readinessCapacity, licences, secrets, network routes, security controls, images and infrastructure-as-code.
8Ordered recovery runbookExecutable steps with owners, expected duration, prerequisites, validation and rollback.
9Data reconciliationIdentify lost/duplicate/queued transactions and reconcile authoritative sources.
10Cyber clean recoverySeparate availability recovery from integrity recovery; define trusted recovery point and clean-room process.
11Technical validationHealth, security, performance, interfaces, job scheduling and monitoring checks.
12Business validationBusiness owner verifies outcomes with known test cases before general release.
13CommunicationsStakeholder status, vendor escalation and decision cadence.
14Failback / return to primaryStability criteria, sync direction, change freeze, rollback and approval.
15Testing and evidenceMeasure actual recovery time/point and retain defects, screenshots/logs and business sign-off.
16MaintenanceUpdate after architecture, supplier, application, RTO/RPO or security changes.

Build an RTO budget, not just an RTO label

If a business service has a two-hour RTO, every dependency cannot also have a two-hour RTO. The end-to-end sequence needs a time budget. The example below is illustrative; each organisation should measure its own recovery sequence and parallel activities.

Elapsed targetRecovery activityGate
0–10 minDeclare recovery path; confirm incident/cyber containment stateAuthority and target recovery point agreed
10–30 minRecover/confirm identity, DNS, network, secrets and storage prerequisitesCore platform dependencies healthy
20–55 minDatabase restore/failover and integrity checksDatabase consistent and access controlled
40–75 minApplication and middleware recoveryTechnical health checks pass
65–90 minIntegrations, queues, scheduled jobs and downstream connectionsNo uncontrolled duplicate/replay risk
85–105 minBusiness validation using known transactions/casesBusiness owner accepts minimum service
105–120 minControlled traffic release / priority usersMonitoring, rollback and communications active

System and application recovery inventory

FieldExample / question
Business serviceWhich customer or internal outcome depends on this system?
Application / componentApplication, database, queue, API gateway, identity, DNS, network, batch scheduler, storage, endpoint or SaaS
OwnerWho makes technical decisions and who validates business outcome?
RTO / RPOWhat approved requirement applies and where is the source BIA?
Recovery methodRestore, active-passive failover, active-active, rebuild from code, SaaS vendor recovery, manual substitute
PrerequisitesNetwork, credentials, secrets, storage, licences, certificates, external API, supplier
Recovery locationRegion/site/account/subscription/tenant
Runbook ID / automationControlled procedure and tested version
Last demonstrated resultMeasured recovery time, achieved recovery point, defects, date, test scope

Backup and recovery evidence

Backup success is not the same as recoverability. Record whether backups can be restored into an isolated environment, whether encryption keys and configuration are available, whether the recovery point is old enough to pre-date a cyber compromise, and whether the business can reconcile transactions created after that point.

  • Define backup scope for databases, files, virtual machines or images, application configuration, infrastructure-as-code, secrets/keys according to secure key-management practices, network/security configuration and critical SaaS exports where needed.
  • Monitor failed jobs and capacity. A green daily job status is only one control.
  • Test restoration at a frequency and depth proportional to criticality. Capture elapsed time and validation evidence.
  • Use separation/immutability/offline patterns where appropriate so the same credentials or destructive event cannot erase production and recovery copies.
  • Protect recovery administration: emergency access and privileged credentials are dependencies of the DR plan.
  • Reconcile RPO with transaction behaviour. A 15-minute data RPO may be unacceptable for a payment ledger unless additional transaction-level controls provide tighter integrity.

Twelve-step recovery runbook template

StepInstruction fieldEvidence / rollback
1Declare recovery mode and freeze conflicting changesIncident reference, approver, recovery target
2Confirm clean/trusted recovery pointBackup/replica timestamps, cyber assessment
3Establish privileged recovery accessBreak-glass or recovery-admin evidence
4Recover network/DNS/security prerequisitesConnectivity and policy tests
5Recover storage/database layerRestore/failover logs, consistency checks
6Recover middleware, messaging and cachesQueue state and replay rules
7Recover application servicesHealth endpoints, service logs
8Recover external integrations in controlled orderPartner acknowledgements, API checks
9Restore scheduled/batch processingJob dependencies and cutoff implications
10Technical validationAvailability, performance, monitoring, security
11Business validation and reconciliationKnown test cases, counts, balances, exceptions
12Controlled release and observeTraffic percentage, rollback threshold, owner approval

Cyber recovery: why failover can be the wrong first move

A conventional hardware outage often favours rapid failover. A cyber incident may require a different decision. If credentials, software images, replication streams or data integrity are compromised, failing over can reproduce the problem. The DRP should therefore define a branch for clean recovery: containment confirmation, trusted restore point, clean administrative access, malware/indicator checks, credential rotation where appropriate, controlled network reconnect, business data validation and staged release.

Thirty DR review questions

  1. Which business services depend on this application?
  2. Are RTO and RPO approved by business owners?
  3. Has actual recovery time been measured?
  4. Has actual recovered data age been measured?
  5. Are upstream and downstream dependencies mapped?
  6. Is identity recovery included?
  7. Are DNS/network/firewall/load-balancing dependencies included?
  8. Are certificates, secrets and licences recoverable?
  9. Is the recovery environment sized for required MBCO?
  10. Can infrastructure be rebuilt from controlled code/configuration?
  11. Are backups monitored?
  12. Has restore been tested from the same backup technology used in production?
  13. Is a protected/isolated recovery copy available where risk warrants it?
  14. Can encryption keys be recovered securely?
  15. How is the trusted recovery point selected after cyber compromise?
  16. Are queues and in-flight transactions reconciled?
  17. What prevents duplicate message or payment replay?
  18. Are third-party/SaaS dependencies in the runbook?
  19. Are vendor emergency contacts tested?
  20. Are external allow-lists/routes preconfigured?
  21. Are technical validation criteria explicit?
  22. Are business validation cases explicit?
  23. Who can approve controlled reopening?
  24. Are performance and capacity tested after recovery?
  25. Are monitoring/logging/security controls restored before full release?
  26. Is failback documented?
  27. Is rollback from failback possible?
  28. Are test defects tracked to closure?
  29. Does the test include people/process as well as technology?
  30. Does every material architecture change trigger a DR plan review?

Example DR test report

MeasureTargetObservedAssessment / action
Service RTO2h1h 47mMet; preserve automation and repeat under peak-load scenario
Database RPO15m7mMet for test; confirm during unplanned failover scenario
Business validation≤15m after technical recovery22mGap: automate validation data set and pre-assign validator
Critical API connectivityAll priority APIs4/5 at first checkGap: missing recovery-region allow-list; corrective action assigned
FailbackDocumented and rehearsedWalk-through onlySchedule controlled failback test before declaring full capability

Cloud, SaaS and on-premise variations

Cloud workloads

Include region/account/subscription failure, control-plane dependency, quotas, infrastructure-as-code, cloud identity, key management, private connectivity, DNS and shared services. “Multi-AZ” or “multi-region” labels do not prove business recovery unless the application and its dependencies have been exercised.

SaaS services

The provider may own platform restoration, while the customer still owns business continuity: identity configuration, data export, integration recovery, user communications, alternate process, contractual escalation and exit. Know the provider commitment, but also define what the business does while waiting.

On-premise / data-centre workloads

Include power, cooling, network carriers, storage, hardware replacement, virtualisation, facility access, media, spares and alternate-site dependencies. Test whether the DR site relies on the same staff, directory service, network path or supplier that could be affected at the primary site.

Official references and further reading

  • NIST SP 800-34 Rev. 1 — US NIST contingency-planning guidance for federal information systems; useful structured reference for BIA, recovery strategy, plans, testing and maintenance.
  • CERT-In government security guidance — Includes documented backup, separation and regular backup expectations for Indian government entities.
  • ISO 22301:2019 — BCMS requirements connecting recovery capability to business continuity needs.
  • ISO/TS 22331:2018 — Continuity strategy guidance.

Build an RTO budget instead of giving every component the same RTO

A business RTO is an end-to-end outcome. If the service must resume in four hours, technology cannot consume all four hours because detection, declaration, infrastructure, applications, interfaces, data validation and business checks each consume time. An RTO budget makes those assumptions visible and testable.

StageIllustrative budget for a 4-hour service RTOEvidence to capture
Detect / assess / declare0:00–0:20Monitoring alert, incident ticket, declaration timestamp
Platform / network readiness0:20–1:00Environment and connectivity health checks
Database / core data recovery1:00–1:50Restore/replication evidence and integrity checks
Application recovery1:50–2:40Runbook timestamps, service health
Interfaces / identity / batch2:40–3:15Integration and authentication tests
Business validation3:15–3:40Representative transactions and controls
Release / communications buffer3:40–4:00Go/no-go decision and user notification

Illustrative only: the allocation above is not an industry benchmark. Build the budget from your architecture and test evidence. The useful metric is whether the complete service repeatedly returns within the approved target.

Recovery dependency matrix

ComponentDepends onRecovery orderValidationTypical hidden failure
IdentityDirectory, MFA, network, time syncBefore user-facing appsEmergency and standard loginDR infrastructure works but authentication points to failed primary service
DatabaseStorage, keys, backup/replicationBefore application data accessIntegrity, transaction consistency, RPOBackup exists but encryption key or credential is unavailable
ApplicationDatabase, secrets, DNS, certificatesAfter platform/dataHealth + business transactionCertificate/DNS/secret differs in DR
API / integrationNetwork routes, external parties, credentialsBefore end-to-end validationRequest/response + queue reconciliationPartner firewall allows only primary IP range
Batch / schedulerApplication, file transfer, calendarAfter core serviceRun due jobs without duplicationJobs re-run and duplicate transactions
Monitoring / loggingAgents, collectors, storageDuring recoveryAlerts and audit records visibleDR works but operators are blind to degradation

Backup evidence register

Evidence fieldWhy it matters
Protected datasets and systemsProves scope follows business criticality, not convenience.
Backup frequencyAllows comparison with approved RPO.
Retention and version historySupports recovery from delayed discovery or corruption.
Logical/physical separationReduces common-mode loss and ransomware exposure.
Immutability / protected credentialsLimits destructive access from the production identity plane.
Restore test date and data setA successful backup job is not proof of recoverability.
Observed restore timeFeeds the RTO budget.
Integrity / reconciliation resultConfirms restored data is usable, complete and consistent.
Key / certificate / secret recoveryPrevents technically restored systems failing because dependencies are unavailable.
Owner and next testKeeps evidence actionable.

Cyber recovery decision points

  • Do not automatically fail over to a replicated environment if the incident may involve compromised identities, malware or corrupted data.
  • Define who determines a clean recovery point and what forensic/security evidence is required.
  • Protect backup administrative paths separately from normal production administration.
  • Document how privileged credentials, keys and secrets are rotated during recovery.
  • Validate that recovered systems are patched and trusted before reconnecting to production networks.
  • Decide how to handle data created during the outage or in isolated environments.
  • Include executive/business decisions when the safest recovery point creates data loss beyond normal RPO.
  • Practice communications when recovery estimates are uncertain rather than promising an unverified restoration time.

DR test types and what each proves

Test typeUseful forDoes not prove by itself
Runbook walkthroughFinding missing steps, owners, access and assumptionsThat systems can actually recover
Component restoreBackup usability, database/file restorationEnd-to-end service recovery
Technical failoverInfrastructure/application recoveryBusiness service, data and control correctness
Integrated service testInterfaces, identity, dependencies and business validationFull organisational response under crisis pressure
Cyber clean-room recoveryIsolated restore, trust validation, identity/key recoveryNormal-site/business continuity workarounds
Full business continuity exercisePeople, communications, decisions, workaround and technology togetherThat every future scenario will succeed

DR test report: evidence that should survive the exercise

FieldExample content
Scenario / scopePrimary region unavailable; recover Payment Service A and dependencies
Start / declaration09:00 detected; 09:12 DR declared
RTO / actualTarget 4h; business validated 3h 28m
RPO / actualTarget 15m; measured loss 7m; reconciliation completed
Recovery sequenceIdentity → network → database → app → interfaces → business validation
Failures / workaroundsPartner API firewall missing DR subnet; temporary approved rule added
Business validation20 representative transactions, reversals, reporting and audit logs passed
Residual riskManual settlement report delayed until 14:00
ActionsNetwork rule baseline, automated certificate check, retest due date
ApprovalTechnology recovery lead + business service owner
NIST connection

NIST SP 800-34 Rev. 1 is a useful U.S. federal information-system contingency reference. It links policy, BIA, preventive controls, recovery strategy, contingency plans, testing/training/exercises and maintenance. Use it where applicable; do not treat a technology contingency guide as a substitute for enterprise BCM.

Jurisdiction and regulated-sector DR overlay

Technical recovery targets may be affected by banking, market-infrastructure, public-sector, cyber-security or resilience rules. Add a short compliance matrix to the DRP with jurisdiction, regulated service/system, applicable source, required recovery/testing/notification control, owner and evidence location. This prevents the runbook from becoming a generic technology document disconnected from obligations.

Where a regulator specifies testing, recovery-site, transaction-integrity, cyber-recovery or third-party requirements, map them to concrete runbook steps and evidence rather than adding a citation only.

Choose the right DRP template variant without creating a separate plan for every keyword

A disaster recovery plan is most useful when one controlled document can be adapted to the technology and business context. An IT disaster recovery plan, a payroll recovery plan, a cloud recovery runbook and a supply-chain technology recovery plan can share the same control structure while using different inventories, dependencies and recovery procedures. The core questions remain the same: what must be restored, by when, from which recovery point, in what sequence, by whom, and what evidence proves the service works after recovery?

DRP variantWhat changesWhat should remain common
IT disaster recovery planApplications, infrastructure, identity, network, backups, cloud regions and technical runbooks.Activation, roles, dependency order, RTO/RPO, evidence, communications and failback controls.
Payroll disaster recovery planPayroll calendars, HR/finance data, banking interfaces, cutoff dates, statutory deadlines and manual fallback.Recovery targets, minimum service, data reconciliation, approvals and business validation.
Supply-chain recovery planSupplier systems, EDI/API connections, logistics platforms, alternate suppliers and order backlogs.Dependency mapping, priority transactions, workaround capacity and return-to-normal decisions.
Cloud/SaaS recovery planProvider responsibilities, tenant configuration, exports, identity dependencies, region options and support escalation.Business-owned RTO/RPO, recovery evidence, communications and third-party assurance.
Cyber recovery planContainment, clean-room rebuild, credential reset, forensic preservation, immutable backup and trust restoration.Authority, recovery sequence, validation gates and documented risk acceptance.

Disaster recovery policy, DR plan and DR project plan are different artefacts

A disaster recovery policy sets governance: scope, accountability, minimum testing expectations, recovery-objective ownership and evidence requirements. A disaster recovery plan is the operational document used during a disruption. A DR project plan is temporary delivery management used to implement or improve recovery capability, for example moving backups off-site, building a secondary environment or closing test findings. Mixing the three creates a document that is difficult to approve and even harder to use during an incident.

DR project plan milestones

  • Confirm critical services, systems and dependency owners.
  • Validate RTO/RPO requirements against current capability.
  • Select recovery architecture and approve the risk/cost decision.
  • Implement backups, replication, infrastructure, identity and network prerequisites.
  • Write system runbooks and business validation steps.
  • Run component tests, then integrated recovery exercises.
  • Record achieved recovery times, gaps, owners and due dates.
  • Move the capability into a scheduled maintenance and assurance cycle.

How ISO/IEC 27001 fits with a disaster recovery plan

ISO/IEC 27001 is an information security management standard rather than a prescriptive DRP-template standard. Organizations commonly use disaster recovery documentation, backup evidence, recovery testing and continuity arrangements as part of demonstrating that information and technology services can be protected and restored. The practical approach is to map your organization’s applicable security and continuity controls to evidence in the DR plan instead of claiming that one generic template is an “ISO 27001 certified DR plan.”

Simple DR plan format for a small environment

  1. Scope: systems and services covered.
  2. Activation: conditions, authority and contacts.
  3. Targets: RTO, RPO and minimum service expectations.
  4. Dependencies: people, identity, network, facilities, suppliers and data.
  5. Recovery sequence: numbered technical actions with owners and prerequisites.
  6. Validation: technical health checks and business transaction tests.
  7. Communications: status updates, escalation and stakeholder messages.
  8. Failback: criteria and steps for returning to the normal environment.
  9. Evidence: timestamps, screenshots/log references, results and corrective actions.

Use the full 16-section structure above when the environment is complex. For a small service, this shorter format is acceptable only if it still captures the decisions and evidence needed to recover safely.