Topic 13.5
Disaster Recovery and Multi-Region Kafka
In one line
Kafka clusters live in one region; for regional disasters you replicate topics to another cluster with MirrorMaker 2 (or Cluster Linking, MSK Replicator and similar). Plan RPO (replication lag), RTO (failover procedure), consumer offset translation, duplicate handling, and whether you run active-passive or active-active.
Think of it like this
A company keeping a copy of its records at a second office in another city. Copies are sent continuously but arrive a little late, so after a disaster the second office may be missing the last few minutes, and staff need to know where they'd got to in each file.
Key ideas
- 01
Why not stretch one cluster across regions: cross-region latency (70–250 ms) would slow acks=all writes and destabilise replication and controller quorum. Stretched clusters work only across nearby zones or metro sites.
- 02
MirrorMaker 2 (MM2): a Kafka Connect-based replicator with source connectors copying topics (renamed
source.topicby default to avoid loops), checkpoint connectors translating consumer group offsets, and heartbeats. Replication is asynchronous: RPO ≈ replication lag. - 03
Offset translation: offsets differ between clusters; MM2 checkpoints map source group offsets to target offsets so consumers can resume near where they were. Some duplicates are expected, so consumers must be idempotent.
- 04
Active-passive: producers write to the primary region; the DR region receives mirrored data; on failover, producers and consumers switch (DNS or config), consumers start from translated offsets. Simpler, clearer ordering.
- 05
Active-active: both regions accept writes to their own topics; each mirrors to the other; consumers read local and remote-prefixed topics. Lower latency for global users, but ordering across regions isn't defined and conflicts must be resolved in the application.
- 06
Also plan: data residency (which topics may leave a region), cost (cross-region transfer), and regular failover drills.
Code & diagrams
clusters = mumbai, singapore
mumbai.bootstrap.servers = kafka-mum-1:9092,kafka-mum-2:9092
singapore.bootstrap.servers = kafka-sin-1:9092,kafka-sin-2:9092
mumbai->singapore.enabled = true
mumbai->singapore.topics = commerce\..*, payments\..*
mumbai->singapore.groups = .*
replication.factor = 3
sync.group.offsets.enabled = true # write translated offsets into the target cluster
emit.checkpoints.interval.seconds = 10
refresh.topics.interval.seconds = 60
# replicated topics appear as mumbai.commerce.orders in singapore (DefaultReplicationPolicy)Interview problem
The problem
Region A is completely unavailable
Region A is down; consumers must continue in Region B. Design RPO, RTO, replication, consumer failover and duplicate handling.
You're given
- Orders and payments topics
- RPO under 1 minute
- RTO under 15 minutes
The interviewer follows up
Why not have producers write to both regions directly?
When it breaks
DR cluster never tested
What you see
During a real outage, consumers in B start from the beginning (offsets not synced) or ACLs and quotas are missing; recovery takes hours and reprocesses days of events.
Fix & prevent
Mirror configs and ACLs too, sync offsets, and run scheduled failover drills with measured RTO.
Explain it without notes
Explain offset translation and why duplicates are expected after DR failover.
Practice
Run two lab clusters with MM2, mirror a topic and a consumer group's offsets, then fail the consumer over to the second cluster.
Trade-offs
- ↔
Active-passive is simpler and ordered; active-active gives local writes everywhere but no cross-region ordering and more complexity.
Done when you can
I can design Kafka DR with MM2, RPO/RTO, offset translation and idempotent consumers.