Topic 10.4
CAP, Network Partitions, and Split Brain
In one line
Don't label Redis "AP" or "CP". Analyse the specific deployment: during a partition, an isolated primary may keep accepting writes that are discarded when it rejoins (split brain), and settings like min-replicas-to-write and client behaviour decide how much. Redis replication is not consensus.
Think of it like this
Two branches of a shop lose their phone line to head office. If both keep selling from the same stock list, they may sell the last item twice (split brain). If one stops selling until the line is back, customers there are turned away (unavailable). CAP says you can't have both during the outage.
Key ideas
- 01
CAP in one line: during a network partition, a system must choose between consistency (refuse some requests) and availability (answer, possibly inconsistently). Partitions aren't optional, so the real question is what your Redis deployment does when one happens.
- 02
Scenario: the primary is on the minority side with some clients; Sentinels and replicas are on the majority side. The majority promotes a replica. Now two primaries accept writes. When the network heals, the old primary becomes a replica and discards everything it accepted during the partition.
- 03
Limiting the damage:
min-replicas-to-write 1+min-replicas-max-lag 10make the isolated primary stop accepting writes after ~10 s without replicas, bounding the loss window. Redis Cluster has a similar rule: a primary that can't reach the majority of primaries stops accepting writes aftercluster-node-timeout. - 04
Why Redis isn't a consensus system: replication is asynchronous and failover isn't coordinated with client writes, so even with these settings, some acknowledged writes can be lost. Systems like etcd, ZooKeeper or Consul use Raft/ZAB, where a write is acknowledged only after a majority stores it.
- 05
How to describe Redis honestly: a single primary is consistent (all reads and writes go to one place) but unavailable if it fails; with replicas and Sentinel/Cluster, it favours availability with bounded, not zero, data loss; with replica reads, it's eventually consistent. Your client routing and settings shift it along that spectrum.
- 06
Design implication: use Redis where brief inconsistency or small loss is tolerable, and put correctness-critical coordination (leader election for money movement, strict locks) on consensus systems or the database, or add fencing (Topic 12.2).
Code & diagrams
Interview problem
The problem
Two app groups, one Redis, and a partition
Application group A and group B both use one Redis primary with replicas and Sentinel. A network partition separates group A plus the primary from group B, the replicas and two of the three Sentinels. Discuss availability, consistency, failover, split brain and stale data.
The interviewer follows up
Why isn't a Redis lock equivalent to one from a consensus system?
When it breaks
Split brain with no write limits on the old primary
What you see
Minutes of writes from part of the application are accepted and then silently discarded when the partition heals.
Fix & prevent
Configure min-replicas-to-write/min-replicas-max-lag (or rely on Cluster's node-timeout write stop), keep Sentinels in three failure domains, and design clients to handle write errors.
Explain it without notes
Why is "Redis is AP" a weak answer?
Practice
Simulate a partition in Docker (disconnect the primary's container from the network) with and without min-replicas-to-write, writing from a client on the primary's side.
Trade-offs
- ↔
Stopping writes on isolated primaries reduces loss but turns partitions into errors for the minority side.
Done when you can
I can analyse a Redis deployment's behaviour during a partition.
I can explain split brain and how to bound it.
I can explain why Redis replication isn't consensus.