Topic 8.3
Multi-Region Filters and Failover
In one line
Each region keeps its own filters for low-latency checks, fed by replicated change streams (Kafka mirroring or CDC per region) and versioned snapshots. Regional filters are eventually consistent; recent-item fallback, residency rules and a failover plan (rebuild or load a peer region's snapshot) keep them correct.
Think of it like this
Branch offices each keeping a copy of the head-office customer list. Updates are mailed to every branch, so each branch is a little behind; if a branch loses its copy, it borrows another branch's latest copy and applies the recent updates.
Key ideas
- 01
Topology: region-local filters (in-process or regional Redis) for lookups under a millisecond; changes replicated between regions via Kafka MirrorMaker 2, per-region CDC from replicated databases, or a global event bus.
- 02
Eventual consistency: an item created in Region A appears in Region B's filter after replication lag. A user created in India and immediately served from Europe could be rejected if Europe's filter trusts "absent". Use watermark bypass (IDs newer than the regional filter's watermark go to the database) or route requests for recent items to the home region.
- 03
Data residency: if identifiers themselves are personal or regulated, a region's filter may only contain its own data; cross-region lookups must go to the owning region.
- 04
Failover: if a region's filter is lost, load the latest snapshot (from object storage, replicated across regions) and replay the change stream from the snapshot's offsets; until ready, fail open to the database with concurrency limits or fail closed for security use cases.
- 05
Global blocklists: same filter everywhere, built centrally, distributed as versioned snapshots plus incremental updates; nodes report their version and alerts fire on laggards.
Code & diagrams
Interview problem
The problem
Region A's filter is unavailable; Region B serves A's traffic
Region A's Bloom filter (and its Redis) becomes unavailable, and Region B starts serving traffic for Region A's data. Design fallback, replication, rebuild and the fail-open/closed decision.
Explain it without notes
How can replication lag between regions cause false negatives, and how do you prevent them?
Practice
Design the distribution of a 4 GB global blocklist filter to 5,000 edge nodes in four regions.
Trade-offs
- ↔
Region-local filters give low latency and resilience at the cost of eventual consistency and replication plumbing.
Done when you can
I can design multi-region filters with safe lag handling, residency and failover.