Command Palette

Search for a command to run...

Hectal
PHASE 13Advanced ~8 min· topic 3 of 5

Topic 13.3

Disaster Recovery: RPO, RTO, Failover and Drills

In one line

RPO is how much data you can afford to lose; RTO is how long you can be down. They drive the DR architecture: backups only (hours), a warm standby in another region (minutes, some data loss with async replication), or active-active (seconds, with conflict handling). Failover, failback and regular drills make the numbers real.

0/5 · 0%

Think of it like this

A hospital's backup generator. RTO is how many seconds the lights can be off before the generator kicks in; RPO is how much of the patient monitoring record can be missing. Both are tested monthly, not assumed.

Key ideas

  1. 01

    Tiers: backup and restore (RPO minutes via WAL, RTO hours); pilot light (minimal infrastructure in the DR region plus replicated data; RTO tens of minutes); warm standby (a scaled-down live copy; RTO minutes); active-active (both regions serve; RTO near zero).

  2. 02

    Cross-region replication is usually asynchronous (latency), so RPO = the replication lag at the moment of failure (typically seconds). Synchronous cross-region replication gives RPO 0 at tens of milliseconds per commit.

  3. 03

    Failover runbook: declare the incident; confirm the primary region is really down (avoid split brain); fence it; promote the DR database; switch DNS or global load balancer; scale up application tiers; communicate. Automate what you can, but keep a human decision for regional failover.

  4. 04

    Failback: rebuild the old region as a replica of the new primary, resync, then plan a controlled switchover. It's often harder than failover.

  5. 05

    Drills: game days that actually fail over production (or a production-like environment) at least twice a year, measuring RPO and RTO achieved and fixing what broke.

Code & diagrams

dr-tiers.txttext
Tier               RPO            RTO           Cost    Example
Backup/restore     ~1 min (WAL)   hours         $       pgBackRest to other-region S3
Pilot light        seconds-min    30-60 min     $$      cross-region replica, apps scaled to 0
Warm standby       seconds        5-15 min      $$$     replica + small app fleet, scale on failover
Active-active      ~0 / seconds   ~0            $$$$    geo-partitioned or multi-writer database
dr-failover.mermaiddiagram
Rendering diagram…

Interview problem

The problem

Set RPO/RTO and design DR for three systems

An online bank has: the core ledger, a customer-support ticketing system, and a marketing analytics warehouse. Propose RPO/RTO per system and a DR design for each.

When it breaks

DR region never tested

What you see

On the day of a regional outage, the DR database is fine but application configs, secrets, DNS TTLs and capacity limits in the DR region are missing; the real RTO is 9 hours.

Fix & prevent

Regular full failover drills, infrastructure as code for both regions, low DNS TTLs, and pre-approved capacity.

Explain it without notes

01

Define RPO and RTO and give the technology that most affects each.

Practice

01

What is the RPO of async cross-region replication with typical 2 s lag and spikes to 30 s?

Trade-offs

  • ↔

    Lower RPO and RTO cost more infrastructure, latency and complexity; match them to business impact per system.

Done when you can

  • I can set RPO/RTO per system and design, run and drill a failover and failback.