Command Palette

Search for a command to run...

Hectal
PHASE 14Advanced ~10 min· topic 2 of 5

Topic 14.2

Failure Injection: Breaking Redis on Purpose

In one line

For each Redis role, rehearse the failures: Redis down, slow, primary or replica lost, partition, maxmemory, hot key, huge key, stampede, slow Lua, consumer crash, duplicate message, restart, persistence recovery and cluster node failure. For every scenario, know what breaks, what's lost, and how you recover.

0/5 · 0%

Think of it like this

Fire drills. Nobody wants a fire, but a building that has rehearsed evacuation handles a real one calmly. Game days do the same for Redis.

Key ideas

  1. 01

    The eight questions for every scenario: (1) what breaks, (2) what keeps working, (3) what data can be lost, (4) what consistency guarantees remain, (5) what happens to latency, (6) what clients see, (7) how recovery works, (8) how to prevent recurrence.

  2. 02

    Tools: stop or pause containers (docker pause, kill -STOP), DEBUG SLEEP in a lab, toxiproxy or tc netem for latency and packet loss, network policies or iptables for partitions, CLIENT PAUSE for write pauses, Chaos Mesh or AWS FIS in real environments, and filling memory to hit maxmemory.

  3. 03

    Run in staging first, with a hypothesis ("checkout still succeeds with Redis down, with p99 under 800 ms"), a blast-radius limit, and an abort condition. Then run carefully in production-like settings.

  4. 04

    Cover the roles, not just the server: cache, sessions, locks, rate limiter, queues and streams each behave differently when Redis fails, so each needs its own expected outcome.

Code & diagrams

scenarios.txttext
Scenario                 What breaks / client sees                      Data at risk                 Recovery / prevention
A Redis down             timeouts -> fallbacks; DB load spikes          cache: none; state: all      fallbacks, circuit breakers, replicas
B Redis latency 500ms    thread pools fill if no timeouts               none                         timeouts, breakers, find cause
C Replica down           less read capacity; failover target gone       none                         replace replica, alert
D Primary down           writes fail ~5-20 s until failover             unreplicated writes          Sentinel/Cluster, WAIT for critical
E Partition              split brain; minority writes discarded         minority-side writes         min-replicas-to-write, 3 zones
F maxmemory reached      OOM errors (noeviction) or evictions           evicted keys                 alerts at 85%, right policy
G Hot key                one node saturated, cluster-wide latency       none                         L1 cache, key replication
H Huge key               seconds-long blocking on access/delete         none                         UNLINK, split keys
I Cache stampede         DB overload after hot key expiry               none                         locks, SWR, coalescing
J Slow Lua script        BUSY errors for everyone                       none                         bounded scripts, timeouts
K Consumer crash         pending entries stuck                          none (if acked properly)     XAUTOCLAIM, idempotency
L Duplicate message      double side effects                            correctness                  idempotent handlers
M Redis restart          empty or reloading; LOADING errors             since last fsync/snapshot    AOF/RDB, replicas
N Persistence recovery   corrupt/truncated AOF, long load               tail of AOF                  redis-check-aof, backups
O Cluster node failure   slots unavailable until failover               unreplicated writes          replicas per shard, zones
chaos.shbash
# B: add 500 ms latency between the app and Redis (Linux, lab)
tc qdisc add dev eth0 root netem delay 500ms
# ... run the load test, observe fallbacks and latency ...
tc qdisc del dev eth0 root

# D: kill the primary and time failover
docker kill redis-primary
redis-cli -p 26379 SENTINEL get-master-addr-by-name mymaster

# F: hit maxmemory
redis-cli CONFIG SET maxmemory 50mb
redis-benchmark -t set -n 2000000 -d 1024 -r 10000000

# J: slow script (lab only)
redis-cli EVAL "local i = 0 while i < 1e9 do i = i + 1 end return i" 0 &
redis-cli PING        # (error) BUSY Redis is busy running a script...
redis-cli SCRIPT KILL

Interview problem

The problem

Game day for a checkout system

Checkout uses Redis for carts, rate limiting, idempotency keys and a product cache. Plan a game day: which Redis failures you inject, what you expect for each role, and what you'll measure.

The interviewer follows up

01

What does a client see during a primary failover?

Explain it without notes

01

List the eight questions to answer for each failure scenario.

Practice

01

Pick scenarios A, D and I and run them in your lab against a small app, writing down the answers to the eight questions for each.

Trade-offs

  • ↔

    Game days take time and carry some risk, but they're far cheaper than discovering these behaviours during a real outage.

Done when you can

  • I can answer the eight questions for every failure scenario A–O.

  • I've injected at least three Redis failures in a lab and verified the app's behaviour.