Topic 14.2
Failure Injection: Breaking Redis on Purpose
In one line
For each Redis role, rehearse the failures: Redis down, slow, primary or replica lost, partition, maxmemory, hot key, huge key, stampede, slow Lua, consumer crash, duplicate message, restart, persistence recovery and cluster node failure. For every scenario, know what breaks, what's lost, and how you recover.
Think of it like this
Fire drills. Nobody wants a fire, but a building that has rehearsed evacuation handles a real one calmly. Game days do the same for Redis.
Key ideas
- 01
The eight questions for every scenario: (1) what breaks, (2) what keeps working, (3) what data can be lost, (4) what consistency guarantees remain, (5) what happens to latency, (6) what clients see, (7) how recovery works, (8) how to prevent recurrence.
- 02
Tools: stop or pause containers (
docker pause,kill -STOP),DEBUG SLEEPin a lab, toxiproxy ortc netemfor latency and packet loss, network policies or iptables for partitions,CLIENT PAUSEfor write pauses, Chaos Mesh or AWS FIS in real environments, and filling memory to hitmaxmemory. - 03
Run in staging first, with a hypothesis ("checkout still succeeds with Redis down, with p99 under 800 ms"), a blast-radius limit, and an abort condition. Then run carefully in production-like settings.
- 04
Cover the roles, not just the server: cache, sessions, locks, rate limiter, queues and streams each behave differently when Redis fails, so each needs its own expected outcome.
Code & diagrams
Scenario What breaks / client sees Data at risk Recovery / prevention
A Redis down timeouts -> fallbacks; DB load spikes cache: none; state: all fallbacks, circuit breakers, replicas
B Redis latency 500ms thread pools fill if no timeouts none timeouts, breakers, find cause
C Replica down less read capacity; failover target gone none replace replica, alert
D Primary down writes fail ~5-20 s until failover unreplicated writes Sentinel/Cluster, WAIT for critical
E Partition split brain; minority writes discarded minority-side writes min-replicas-to-write, 3 zones
F maxmemory reached OOM errors (noeviction) or evictions evicted keys alerts at 85%, right policy
G Hot key one node saturated, cluster-wide latency none L1 cache, key replication
H Huge key seconds-long blocking on access/delete none UNLINK, split keys
I Cache stampede DB overload after hot key expiry none locks, SWR, coalescing
J Slow Lua script BUSY errors for everyone none bounded scripts, timeouts
K Consumer crash pending entries stuck none (if acked properly) XAUTOCLAIM, idempotency
L Duplicate message double side effects correctness idempotent handlers
M Redis restart empty or reloading; LOADING errors since last fsync/snapshot AOF/RDB, replicas
N Persistence recovery corrupt/truncated AOF, long load tail of AOF redis-check-aof, backups
O Cluster node failure slots unavailable until failover unreplicated writes replicas per shard, zones# B: add 500 ms latency between the app and Redis (Linux, lab)
tc qdisc add dev eth0 root netem delay 500ms
# ... run the load test, observe fallbacks and latency ...
tc qdisc del dev eth0 root
# D: kill the primary and time failover
docker kill redis-primary
redis-cli -p 26379 SENTINEL get-master-addr-by-name mymaster
# F: hit maxmemory
redis-cli CONFIG SET maxmemory 50mb
redis-benchmark -t set -n 2000000 -d 1024 -r 10000000
# J: slow script (lab only)
redis-cli EVAL "local i = 0 while i < 1e9 do i = i + 1 end return i" 0 &
redis-cli PING # (error) BUSY Redis is busy running a script...
redis-cli SCRIPT KILLInterview problem
The problem
Game day for a checkout system
Checkout uses Redis for carts, rate limiting, idempotency keys and a product cache. Plan a game day: which Redis failures you inject, what you expect for each role, and what you'll measure.
The interviewer follows up
What does a client see during a primary failover?
Explain it without notes
List the eight questions to answer for each failure scenario.
Practice
Pick scenarios A, D and I and run them in your lab against a small app, writing down the answers to the eight questions for each.
Trade-offs
- ↔
Game days take time and carry some risk, but they're far cheaper than discovering these behaviours during a real outage.
Done when you can
I can answer the eight questions for every failure scenario A–O.
I've injected at least three Redis failures in a lab and verified the app's behaviour.