Skip to main content
Category: Resilience and Concentration

Disaster Recovery

Also known as: DR, IT disaster recovery, DR planning
Simply put

Disaster recovery is the process of restoring access to IT systems, applications, and data after a disruptive event such as a fire, flood, cyberattack, or other natural or human-caused incident. It focuses on getting technology infrastructure back up and running so that an organization can resume normal operations. Disaster recovery is one part of a broader resilience effort and is narrower in scope than business continuity, which addresses the continuation of the organization's overall functions and processes.

Formal definition

Disaster Recovery (DR) refers to the IT technologies, processes, and practices designed to restore access and functionality to critical systems, applications, data, and infrastructure following an unexpected disruption, whether natural or human-caused. In many programs DR is scoped specifically to the recovery of technology assets and typically encompasses activities such as risk assessment, planning, and the restoration of critical systems and data. DR should be distinguished from business continuity: DR addresses the technical recovery of IT services, whereas business continuity addresses the broader continuation of business functions, processes, and personnel. As scoped in the available evidence, DR centers on IT infrastructure recovery and does not by itself cover non-IT operational, financial, or supplier-facing continuity concerns; those fall under adjacent disciplines. In third-party contexts, an organization's own DR capability does not extend to a vendor's environment, so the recovery posture of external providers typically requires separate assessment.

Why it matters

For third-party and supply chain risk professionals, disaster recovery matters because an organization's ability to restore its own IT systems after a disruption does not guarantee that the vendors, service providers, and business partners it depends on can do the same. When a critical supplier's applications, data, or infrastructure go offline following a fire, flood, cyberattack, or other disruptive event, the downstream impact can cascade into the buying organization's operations even if that organization's internal DR posture is sound. DR capability, in other words, is only as complete as the recovery readiness of the parties an organization relies upon.

DR is frequently conflated with business continuity, but the distinction has practical consequences for how risk is assessed. DR is scoped to the technical recovery of IT services, systems, applications, data, and infrastructure, whereas business continuity addresses the broader continuation of business functions, processes, and personnel. A vendor may be able to restore its servers while still being unable to deliver a service if non-IT dependencies remain disrupted. Treating a DR attestation as evidence of full operational resilience overstates what the term covers and can leave gaps in a third-party risk assessment.

Because an organization's own DR program does not extend into a vendor's environment, the recovery posture of external providers typically requires separate, direct evaluation. Relying on a supplier's self-description of its DR capabilities is a point-in-time indicator that may not reflect tested, current, or independently verified recovery performance. This is why DR readiness is a recurring theme in vendor onboarding and ongoing monitoring, rather than a matter that can be assumed from a single contractual clause.

Who it's relevant to

Third-Party Risk Managers
TPRM practitioners evaluate whether critical vendors can restore IT systems, applications, and data after a disruption, recognizing that the organization's own DR capability does not extend into a supplier's environment. This typically means assessing a provider's DR posture separately during onboarding and revisiting it through ongoing monitoring, since a point-in-time attestation may become stale and self-reported claims may lack independent verification.
Operational Resilience and Continuity Teams
These teams position DR as one component of a broader resilience effort, keeping it distinct from business continuity. DR addresses the technical recovery of IT services, while continuity planning addresses the continuation of business functions, processes, and personnel. Conflating the two can leave non-IT operational dependencies unaddressed.
IT and Infrastructure Owners
System and infrastructure owners are responsible for the technologies and practices that restore critical systems and data after natural or human-caused events. Their scope centers on IT assets, risk assessment, planning, and restoration, rather than the full range of business functions, and their work supports but does not by itself constitute overall organizational resilience.
Procurement and Vendor Management
Procurement teams incorporating DR expectations into contracts and vendor selection benefit from distinguishing a vendor's DR commitments from tested, verifiable recovery capability. Because internal DR controls do not cover external providers, contractual language should be paired with a direct assessment of the supplier's recovery posture.

Inside DR

Recovery Time Objective (RTO)
The targeted duration within which a system, application, or IT service should be restored following a disruption. RTO defines the acceptable downtime for a given asset and typically informs the design and prioritization of recovery procedures.
Recovery Point Objective (RPO)
The maximum acceptable amount of data loss measured in time, indicating how far back recovery must reach. RPO drives backup frequency and replication strategy; a shorter RPO generally requires more frequent or continuous data protection.
Backup and Replication
The technical mechanisms for copying data and system states to alternate locations so they can be restored after an incident. This may include periodic backups, snapshots, or continuous replication, depending on the RPO and criticality of the asset.
Recovery Site or Environment
The alternate infrastructure, physical, virtual, or cloud-based, used to restore IT services. Arrangements vary in readiness and typically differ in cost and activation speed depending on how much standby capacity is maintained.
DR Plan and Runbooks
Documented procedures specifying the technical steps, roles, and sequencing for restoring IT systems after disruption. Effective plans identify dependencies among systems and the order in which services should be recovered.
Testing and Validation
Exercises such as tabletop reviews, partial failovers, or full failover tests used to confirm that recovery procedures function as intended. Without periodic testing, documented recovery capabilities may not perform as expected during an actual event.

Common questions

Answers to the questions practitioners most commonly ask about DR.

Is disaster recovery the same as business continuity?
No. Disaster recovery is a subset of the broader business continuity discipline. DR typically focuses on restoring IT systems, applications, data, and technology infrastructure after a disruptive event, whereas business continuity addresses the continuation of essential business functions overall, including people, facilities, processes, and third-party dependencies. A program can have a technically sound DR capability while still lacking a complete business continuity plan, and vice versa. When assessing a third party, it is worth confirming which of the two the provider actually maintains, as the terms are often used interchangeably in vendor documentation.
If a vendor has a documented DR plan, does that mean recovery is assured?
Not necessarily. A documented DR plan describes intended recovery procedures, but documentation alone does not demonstrate that recovery objectives can be met under real conditions. In many programs, the more meaningful evidence is whether the plan has been tested, how recently, the scope of that testing, and the results. A self-reported attestation that a plan exists is distinct from independent verification that it works. Depending on the risk tier, requesting test results or evidence of exercises is generally more informative than confirming the plan's existence.
How should we distinguish RTO and RPO when evaluating a third party's DR capability?
Recovery Time Objective (RTO) typically refers to the targeted duration within which a system or service should be restored after a disruption, while Recovery Point Objective (RPO) typically refers to the maximum acceptable amount of data loss measured as a point in time before the disruption. The two address different dimensions: RTO concerns downtime tolerance, and RPO concerns data loss tolerance. When evaluating a provider, it helps to confirm both values, whether they are contractually committed or merely targets, and whether they align with your own tolerance for the service in question.
What evidence of DR testing is reasonable to request from a critical vendor?
Depending on the risk tier and the nature of the service, organizations often request the date and scope of the most recent DR test, the type of test conducted, a summary of results or issues identified, and evidence of remediation for any gaps. It is useful to understand whether testing covered a full failover or a more limited tabletop or component-level exercise, since these provide different levels of assurance. Test evidence is generally more probative than a plan document, though it remains point-in-time and may become stale between assessment cycles.
How does a vendor's reliance on subcontractors or cloud providers affect DR assessment?
A third party's DR capability may depend on fourth-party or Nth-party providers, such as hosting or cloud services, whose recovery arrangements the vendor may not fully control. This can create limited visibility beyond the direct contractual relationship. In many programs it is worth asking how the vendor's DR plan accounts for its own critical dependencies, whether recovery commitments flow down to those parties, and whether concentration on a shared underlying provider could affect recovery for multiple services simultaneously.
Should DR expectations be written into contracts, and what are the limits of doing so?
Contractual DR provisions, such as committed RTO and RPO values, testing frequency, and notification obligations, can help set expectations and provide a basis for accountability. However, a contractual commitment is a stated obligation rather than proof of capability, and it does not by itself guarantee that recovery will occur as specified. Contract terms are typically more effective when paired with ongoing monitoring and periodic evidence of testing, rather than relied upon as a one-time onboarding control. Enforcement expectations may also vary by jurisdiction and sector.

Common misconceptions

Disaster recovery and business continuity are the same thing.
DR is typically a subset focused on restoring IT systems, data, and technical infrastructure after a disruption, whereas business continuity addresses the broader continuation of critical business functions, including people, processes, facilities, and supplier dependencies. A robust DR capability does not by itself constitute a business continuity program.
Having backups means an organization has disaster recovery.
Backups are one component, but DR also depends on tested restoration procedures, defined RTO and RPO targets, alternate environments, and validated runbooks. Backups that have never been tested for restoration may not meet recovery objectives when needed.
A supplier's stated DR capability guarantees resilience for your organization.
A supplier attestation of DR capability is not the same as independent verification, and it may reflect a point-in-time claim that can become stale. Stated RTO and RPO figures may also not align with your organization's own recovery needs, and visibility often does not extend beyond the direct supplier to their own downstream dependencies.

Best practices

Define RTO and RPO for each critical system based on its business impact, and confirm that supplier DR commitments align with your organization's recovery requirements rather than accepting generic assurances.
Test recovery procedures periodically through failover or restoration exercises rather than relying on documented plans or untested backups, and record whether recovery objectives were actually met.
Request evidence of testing from critical suppliers where feasible, and treat self-reported DR attestations as claims that may warrant independent verification depending on the risk tier.
Map dependencies among systems and services so recovery sequencing accounts for interrelationships, and consider dependencies that may extend to a supplier's own subcontractors where visibility allows.
Distinguish DR scope from broader business continuity planning in contracts and assessments, ensuring that IT restoration commitments do not create a false impression of full operational resilience.
Review DR arrangements on a recurring basis, since point-in-time assessments can become stale as systems, suppliers, and infrastructure change over time.
Promotional banner for the Pentest Readiness checklist download