Command Palette

Search for a command to run...

Hectal
PHASE 13Advanced ~9 min· topic 5 of 5

Topic 13.5

Disaster Recovery and Multi-Region Kafka

In one line

Kafka clusters live in one region; for regional disasters you replicate topics to another cluster with MirrorMaker 2 (or Cluster Linking, MSK Replicator and similar). Plan RPO (replication lag), RTO (failover procedure), consumer offset translation, duplicate handling, and whether you run active-passive or active-active.

0/5 · 0%

Think of it like this

A company keeping a copy of its records at a second office in another city. Copies are sent continuously but arrive a little late, so after a disaster the second office may be missing the last few minutes, and staff need to know where they'd got to in each file.

Key ideas

  1. 01

    Why not stretch one cluster across regions: cross-region latency (70–250 ms) would slow acks=all writes and destabilise replication and controller quorum. Stretched clusters work only across nearby zones or metro sites.

  2. 02

    MirrorMaker 2 (MM2): a Kafka Connect-based replicator with source connectors copying topics (renamed source.topic by default to avoid loops), checkpoint connectors translating consumer group offsets, and heartbeats. Replication is asynchronous: RPO ≈ replication lag.

  3. 03

    Offset translation: offsets differ between clusters; MM2 checkpoints map source group offsets to target offsets so consumers can resume near where they were. Some duplicates are expected, so consumers must be idempotent.

  4. 04

    Active-passive: producers write to the primary region; the DR region receives mirrored data; on failover, producers and consumers switch (DNS or config), consumers start from translated offsets. Simpler, clearer ordering.

  5. 05

    Active-active: both regions accept writes to their own topics; each mirrors to the other; consumers read local and remote-prefixed topics. Lower latency for global users, but ordering across regions isn't defined and conflicts must be resolved in the application.

  6. 06

    Also plan: data residency (which topics may leave a region), cost (cross-region transfer), and regular failover drills.

Code & diagrams

mm2.propertiesproperties
clusters = mumbai, singapore
mumbai.bootstrap.servers = kafka-mum-1:9092,kafka-mum-2:9092
singapore.bootstrap.servers = kafka-sin-1:9092,kafka-sin-2:9092

mumbai->singapore.enabled = true
mumbai->singapore.topics = commerce\..*, payments\..*
mumbai->singapore.groups = .*
replication.factor = 3
sync.group.offsets.enabled = true          # write translated offsets into the target cluster
emit.checkpoints.interval.seconds = 10
refresh.topics.interval.seconds = 60
# replicated topics appear as mumbai.commerce.orders in singapore (DefaultReplicationPolicy)
dr.mermaiddiagram
Rendering diagram…

Interview problem

The problem

Region A is completely unavailable

Region A is down; consumers must continue in Region B. Design RPO, RTO, replication, consumer failover and duplicate handling.

You're given

  • Orders and payments topics
  • RPO under 1 minute
  • RTO under 15 minutes

The interviewer follows up

01

Why not have producers write to both regions directly?

When it breaks

DR cluster never tested

What you see

During a real outage, consumers in B start from the beginning (offsets not synced) or ACLs and quotas are missing; recovery takes hours and reprocesses days of events.

Fix & prevent

Mirror configs and ACLs too, sync offsets, and run scheduled failover drills with measured RTO.

Explain it without notes

01

Explain offset translation and why duplicates are expected after DR failover.

Practice

01

Run two lab clusters with MM2, mirror a topic and a consumer group's offsets, then fail the consumer over to the second cluster.

Trade-offs

  • ↔

    Active-passive is simpler and ordered; active-active gives local writes everywhere but no cross-region ordering and more complexity.

Done when you can

  • I can design Kafka DR with MM2, RPO/RTO, offset translation and idempotent consumers.