Intermediate Architecture

Run a Recovery Exercise — Failover, Validation and Safe Failback

Test a named failure and a named operation

A useful exercise has a narrow claim: under these conditions, this operation recovered in this time with this observed data loss. “Disaster recovery passed” leaves too much unspecified.

Start with an isolated environment and synthetic or appropriately protected test data. Identify what the test cannot reproduce. A successful isolated drill does not establish that public traffic, external providers or production identity will behave the same way during an incident.

Microsoft’s recovery guidance recommends documented plans, clear responsibilities and regular validation. Treat the runbook as an operational document with an owner.

Establish prerequisites and stopping conditions

Before starting, confirm:

  • The incident lead, recovery operator and business validator are available.
  • The alternate location is permitted for the data and dependencies.
  • Required services, quotas, capacity, keys and credentials are usable.
  • Test endpoints cannot send real payments, notifications or customer traffic.
  • Recovery points are available and the original evidence is preserved.
  • Abort conditions and the route back to the initial state are understood.

Define how you prevent two writable primaries. A network partition may leave the original service running even when the recovery team cannot reach it. Traffic switching alone does not fence an old writer. Use the database or application’s supported promotion and fencing procedures.

Worked example: a constrained alternate region

Imagine a test order service whose primary region is unavailable. The approved alternate region can initially serve only half the normal throughput. The business owner agrees that order submission has priority over reporting.

The exercise runbook should be explicit:

StepOwnerEvidence
Declare the simulated outage and start the clockIncident leadTimestamp and failure scope
Confirm the usable data point and promotion conditionsData ownerRecovery point, lag and fencing checks
Start dependencies and application capacityRecovery operatorDeployment results and dependency checks
Apply the agreed limited-service configurationApplication ownerReporting disabled; order path prioritised
Switch test trafficNetwork ownerRoute, name resolution and endpoint checks
Submit and reconcile test ordersBusiness validatorOrder IDs, duplicates, missing records and latency
Declare the agreed operation restoredIncident leadFinish time and accepted limitations

Measure whether that reduced capacity meets the agreed recovery acceptance criteria. Do not silently call a slow, incomplete operation “recovered” because its health endpoint returns success.

Failback is a separate change

After the primary location becomes usable, the recovered location may hold the newest writes. Reopening the old primary without synchronising and validating data can lose work or create conflicting histories.

Plan the direction of data synchronisation, the write cutover, traffic movement and rollback conditions. Validate the business operation again after returning. Keep recovery protection active while locations change roles.

Oracle Full Stack Disaster Recovery distinguishes planned switchover, unplanned failover and drills that create a replica application stack for validation. Their prerequisites and effects differ; consult the Oracle terminology and concepts for the workflow you actually use.

Publish the result with its limits

Record required RTO/RPO alongside measured interruption, recoverable point, missing transactions, achieved throughput and time to clear the backlog. List exclusions, manual decisions and any temporary access grants. Assign an owner and due date to every gap, then schedule the next representative exercise.

Use the recovery planner to prepare requirements and the cost guide to budget for standby resources, retained data and the exercise itself.

References