Design for failure. Prepare for recovery.

An outage is a business interruption before it is an infrastructure problem. Decide how long the operation can stop, how much data could be lost, and which failures the recovery plan must survive.

Resilience PlannerFive answers surface recovery priorities and conflicting requirements.5 questionsLive prioritiesNext steps Start planningClose section

Prepare a recovery brief.

Five questions about one business operation. Get priorities, unresolved constraints and what to test next. These are planning prompts, not guaranteed recovery times or a compliance decision.

  1. Question 1 of 5 · Time

    How long can this business operation be unavailable?

    Agree the target with its business owner. Include detection, decisions, restore and validation.

    Choose one No answer selected
  2. Question 2 of 5 · Data

    How much recent data could you lose?

    Think about the business changes you would have to reconstruct, not just the backup schedule.

    Choose one No answer selected
  3. Question 3 of 5 · Failure

    What is the widest infrastructure outage you need to plan for?

    Every plan also needs a way back from deletion or corruption. Select the widest infrastructure boundary in scope.

    Choose one No answer selected
  4. Question 4 of 5 · Location

    Can recovery data and operations move to another region?

    Check the permitted locations for production data, backups, keys and supporting services.

    Choose one No answer selected
  5. Question 5 of 5 · Readiness

    Have you restored the whole business operation under these conditions?

    A backup success message or a tabletop exercise alone does not measure an end-to-end restore.

    Choose one No answer selected
What it takes to recoverExplore time, lost changes, recovery copies and the failures your plan must survive.RTORPORecovery copiesFailure boundaries ExploreClose section

Four questions before the next outage.

Close a harbour, lose a record, test a backup. See what must recover—and what the business can afford to lose.

Figure 01 Recovery time
Recovery time includes recognising a harbour closure and restoring deliveries through an alternate harbour. N PRIMARY ALTERNATE RTO · TIME TO RESTORE THE OPERATION INTERRUPTION OPERATION RESTORED

Get the harbour working again.

RTO is the target time to restore a business operation. A closed harbour is recovered only when deliveries work again—not when an alternate route merely opens.

Ask about your workload

How long can this operation stop, including detection, decisions, dependencies and checking that users can work again?

Details & limits of the example

A closed harbour stops deliveries. Opening another route takes time: recognise the closure, make the decision, prepare the destination and confirm that cargo can actually arrive.

A provider SLA describes contractual service commitments. Your RTO includes your detection, decisions, dependencies and validation. This route is illustrative, not a measured recovery time.

Figure 02 Recovery point
A failure at 09:00 leaves a 15-minute data gap when the last usable recovery point is 08:45. RPO · MAXIMUM ACCEPTABLE DATA LOSS LIVE CARGO LOG USABLE COPY · 08:45 08:45 · 12 crates08:50 · +2 crates08:55 · +1 crate 08:45 · 12 crates LOG LOST AT 09:00Restore this point. EARLIER HISTORY08:4509:00 Example: 15-minute gap

How far back is the last safe record?

RPO is the acceptable data-loss window, measured in time. Like a saved cargo manifest, recovery can restore only the records in a usable recovery point.

Ask about your workload

How much recent work can the business afford to lose, and is the newest usable recovery point recent enough?

Details & limits of the example

The shore office has a saved cargo manifest. If the ship’s log is lost, you can recover only the entries that reached a usable copy. Newer cargo movements may need reconciliation.

Backup frequency alone does not prove RPO: check successful copies, replication lag, consistency and whether you can restore them.

Figure 03 Backups & replicas
A deletion propagates to the live replica. An earlier retained recovery point still contains the record. A LIVE MIRROR IS NOT A RECOVERY HISTORY PRIMARY Cargo #4212 crates LIVE REPLICA Cargo #4212 crates RETAINED COPY Cargo #4212 crates RESTORE A KNOWN-GOOD POINT · RECONCILE LATER WORK

Keep a history, not just a mirror.

A live replica can repeat a deletion or corruption, like two harbour offices mirroring a ledger. A protected earlier recovery point gives you a way back.

Ask about your workload

Where is the last known-good copy if every live record is damaged, and when did you last prove it restores?

Details & limits of the example

Two harbour offices mirror the same cargo ledger. A mistaken deletion reaches both. A protected earlier manifest lets the crew recover what a live mirror cannot.

Use retention, appropriate isolation and access controls. Where supported, immutable backups can reduce tampering risk. A backup job succeeding is not the same as a restore succeeding.

Figure 04 Failure boundaries
A common harbour closure disables two local piers. A prepared alternate harbour survives this particular failure. N SEPARATE COPIES · CHECK SHARED DEPENDENCIES PIER A PIER B OTHER HARBOUR ONE FAILURE BOUNDARYA PREPARED ALTERNATIVE

Two piers can share one storm.

Two piers share one harbour’s closure. Cloud copies can also fail together; separate instances, zones and regions cover different failure scopes.

Ask about your workload

What could disable both environments, and does the recovery location have independent access, data and enough capacity?

Details & limits of the example

Two piers in one harbour share its closure. A prepared harbour elsewhere may survive, but the crew still needs the records, permissions, route and capacity to operate there.

Zones are not a substitute for region-outage planning. A paired region does not automatically provide application failover. Verify services, data residency and recovery dependencies.

Recovery patterns and evidenceSelected articles and primary sources for the decisions that need a closer look.Practical articlesSource documents Browse notesClose section

Read the detail behind the decision.

Start with the target agreement, protection checklist and exercise runbook. The regional notes help verify placement and dependencies.