Design for failure. Prepare for recovery.
An outage is a business interruption before it is an infrastructure problem. Decide how long the operation can stop, how much data could be lost, and which failures the recovery plan must survive.
Resilience PlannerFive answers surface recovery priorities and conflicting requirements.5 questionsLive prioritiesNext steps Start planningClose section
Prepare a recovery brief.
Five questions about one business operation. Get priorities, unresolved constraints and what to test next. These are planning prompts, not guaranteed recovery times or a compliance decision.
What it takes to recoverExplore time, lost changes, recovery copies and the failures your plan must survive.RTORPORecovery copiesFailure boundaries ExploreClose section
Four questions before the next outage.
Close a harbour, lose a record, test a backup. See what must recover—and what the business can afford to lose.
Get the harbour working again.
RTO is the target time to restore a business operation. A closed harbour is recovered only when deliveries work again—not when an alternate route merely opens.
Deliveries use the main harbour; an alternate route is prepared.
Ask about your workload
How long can this operation stop, including detection, decisions, dependencies and checking that users can work again?
Details & limits of the example
A closed harbour stops deliveries. Opening another route takes time: recognise the closure, make the decision, prepare the destination and confirm that cargo can actually arrive.
A provider SLA describes contractual service commitments. Your RTO includes your detection, decisions, dependencies and validation. This route is illustrative, not a measured recovery time.
How far back is the last safe record?
RPO is the acceptable data-loss window, measured in time. Like a saved cargo manifest, recovery can restore only the records in a usable recovery point.
Saved manifest: 08:45. Live log: 09:00.
Ask about your workload
How much recent work can the business afford to lose, and is the newest usable recovery point recent enough?
Details & limits of the example
The shore office has a saved cargo manifest. If the ship’s log is lost, you can recover only the entries that reached a usable copy. Newer cargo movements may need reconciliation.
Backup frequency alone does not prove RPO: check successful copies, replication lag, consistency and whether you can restore them.
Keep a history, not just a mirror.
A live replica can repeat a deletion or corruption, like two harbour offices mirroring a ledger. A protected earlier recovery point gives you a way back.
Two live ledgers agree. A retained earlier copy is separate.
Ask about your workload
Where is the last known-good copy if every live record is damaged, and when did you last prove it restores?
Details & limits of the example
Two harbour offices mirror the same cargo ledger. A mistaken deletion reaches both. A protected earlier manifest lets the crew recover what a live mirror cannot.
Use retention, appropriate isolation and access controls. Where supported, immutable backups can reduce tampering risk. A backup job succeeding is not the same as a restore succeeding.
Two piers can share one storm.
Two piers share one harbour’s closure. Cloud copies can also fail together; separate instances, zones and regions cover different failure scopes.
Two local piers operate; another harbour has a prepared route.
Ask about your workload
What could disable both environments, and does the recovery location have independent access, data and enough capacity?
Details & limits of the example
Two piers in one harbour share its closure. A prepared harbour elsewhere may survive, but the crew still needs the records, permissions, route and capacity to operate there.
Zones are not a substitute for region-outage planning. A paired region does not automatically provide application failover. Verify services, data residency and recovery dependencies.
Recovery patterns and evidenceSelected articles and primary sources for the decisions that need a closer look.Practical articlesSource documents Browse notesClose section
Read the detail behind the decision.
Start with the target agreement, protection checklist and exercise runbook. The regional notes help verify placement and dependencies.
- RTO and RPO — From Business Impact to a Tested Recovery Target →
Agree what recovery means for one business operation, then measure time and data loss in a representative exercise.
- Backups, Replication and Point-in-Time Recovery — Three Different Jobs →
Separate live copies from recoverable history, choose a clean point and reconcile the valid work that came later.
- Run a Recovery Exercise — Failover, Validation and Safe Failback →
A practical exercise record for restoring the business operation, checking data and returning safely after a regional outage.
- Regions, Zones, Availability Domains — Where Your Data Actually Lives →
Region choice locks in data residency, resilience, and service availability for years. The portal calls it a dropdown. It is an architectural decision.
- Service Availability by Region — Why You Cannot Trust the Map →
Verify exact service features, quotas and recovery capacity in the regions your workload depends on.
Sources & method
Reviewed 14 September 2026. The illustrations simplify failure scenarios; the planner translates your stated targets into questions and priorities. It does not calculate achieved RTO, RPO or availability.