Topic 6.6
Cache Avalanche and Cache Warming
In one line
A cache avalanche is many keys expiring or disappearing at the same time (synchronized TTLs, a Redis restart, a cold deployment), sending a flood of misses to the database. Prevent it with TTL jitter, staggered expiry, warming, multi-level caches and database protection.
Think of it like this
A supermarket that restocks every shelf at exactly 9:00 and has everything go off at exactly 9:00 the next day. For a moment every shelf is empty and the storeroom is swamped. Staggering restock times keeps shelves full.
Key ideas
- 01
Causes: (1) bulk loads that set the same TTL for millions of keys at once, so they all expire together; (2) a Redis restart without persistence, a failover to an empty replica, or a
FLUSHALL; (3) a deployment that changes the key version (all new keys, all misses); (4) mass eviction from memory pressure. - 02
TTL jitter:
ttl = base × (1 + random(−0.1, 0.1))or base + random 0–20%. Keys written together expire over a window instead of an instant. - 03
Cache warming: before sending traffic to a cold cache (new cluster, new key version, after a flush), pre-load the hot set. Sources: yesterday's top keys (from Top-K or access logs), a list of popular products, or replaying recent reads against the new cache. Warm gradually so the database isn't hit all at once.
- 04
Deploy-safe key versioning: roll out a new key version gradually (canary instances first) or dual-read (new key, fall back to old key) during the transition so the whole fleet doesn't go cold together.
- 05
Survive restarts: for large caches whose cold start would hurt, enable persistence (RDB snapshots are enough for a cache) or keep replicas so a failover gets a warm copy.
- 06
Multi-level cache and database protection: an L1 in-process cache absorbs part of the flood; circuit breakers, concurrency limits and load shedding keep the database alive while the cache refills.
Code & diagrams
import random
def ttl_with_jitter(base: int, spread: float = 0.15) -> int:
return int(base * (1 + random.uniform(-spread, spread)))
def warm(r, db, product_ids, batch=500, pause=0.05):
"""Warm a cold cache from a list of hot IDs without flooding the database."""
import time
for i in range(0, len(product_ids), batch):
rows = db.fetch_products(product_ids[i:i + batch]) # one query per batch
pipe = r.pipeline(transaction=False)
for row in rows:
pipe.set(f"product:v2:{row.id}", serialize(row), ex=ttl_with_jitter(3600))
pipe.execute()
time.sleep(pause) # pace the databaseInterview problem
The problem
Morning avalanche after a nightly reload
A nightly job loads 5M product entries into Redis at 02:00 with EX 86400. Every day at 02:00 the database and API latency spike badly. Also, a planned Redis migration to a new cluster is coming up. Diagnose and fix both.
You're given
- 5M keys
- Database handles ~5K queries/sec
- Traffic at 02:00 is low but nonzero
- Migration must not cause an outage
The interviewer follows up
How is an avalanche different from a stampede?
Explain it without notes
List four causes of a cache avalanche and one prevention for each.
Practice
Write a warming script that reads the 10,000 most accessed keys from yesterday's Top-K structure and pre-loads them into a new cluster.
Trade-offs
- ↔
Jitter spreads load but makes expiry times less predictable.
- ↔
Warming protects the database but takes time and must be kept in sync with what's actually hot.
Done when you can
I can prevent avalanches with jitter, warming, persistence and gradual rollouts.
I can explain the difference between avalanche and stampede.