Command Palette

Search for a command to run...

Hectal
PHASE 6Intermediate ~9 min· topic 6 of 7

Topic 6.6

Cache Avalanche and Cache Warming

In one line

A cache avalanche is many keys expiring or disappearing at the same time (synchronized TTLs, a Redis restart, a cold deployment), sending a flood of misses to the database. Prevent it with TTL jitter, staggered expiry, warming, multi-level caches and database protection.

0/7 · 0%

Think of it like this

A supermarket that restocks every shelf at exactly 9:00 and has everything go off at exactly 9:00 the next day. For a moment every shelf is empty and the storeroom is swamped. Staggering restock times keeps shelves full.

Key ideas

  1. 01

    Causes: (1) bulk loads that set the same TTL for millions of keys at once, so they all expire together; (2) a Redis restart without persistence, a failover to an empty replica, or a FLUSHALL; (3) a deployment that changes the key version (all new keys, all misses); (4) mass eviction from memory pressure.

  2. 02

    TTL jitter: ttl = base × (1 + random(−0.1, 0.1)) or base + random 0–20%. Keys written together expire over a window instead of an instant.

  3. 03

    Cache warming: before sending traffic to a cold cache (new cluster, new key version, after a flush), pre-load the hot set. Sources: yesterday's top keys (from Top-K or access logs), a list of popular products, or replaying recent reads against the new cache. Warm gradually so the database isn't hit all at once.

  4. 04

    Deploy-safe key versioning: roll out a new key version gradually (canary instances first) or dual-read (new key, fall back to old key) during the transition so the whole fleet doesn't go cold together.

  5. 05

    Survive restarts: for large caches whose cold start would hurt, enable persistence (RDB snapshots are enough for a cache) or keep replicas so a failover gets a warm copy.

  6. 06

    Multi-level cache and database protection: an L1 in-process cache absorbs part of the flood; circuit breakers, concurrency limits and load shedding keep the database alive while the cache refills.

Code & diagrams

jitter-and-warm.pypython
import random

def ttl_with_jitter(base: int, spread: float = 0.15) -> int:
    return int(base * (1 + random.uniform(-spread, spread)))

def warm(r, db, product_ids, batch=500, pause=0.05):
    """Warm a cold cache from a list of hot IDs without flooding the database."""
    import time
    for i in range(0, len(product_ids), batch):
        rows = db.fetch_products(product_ids[i:i + batch])      # one query per batch
        pipe = r.pipeline(transaction=False)
        for row in rows:
            pipe.set(f"product:v2:{row.id}", serialize(row), ex=ttl_with_jitter(3600))
        pipe.execute()
        time.sleep(pause)                                       # pace the database
avalanche.mermaiddiagram
Rendering diagram…

Interview problem

The problem

Morning avalanche after a nightly reload

A nightly job loads 5M product entries into Redis at 02:00 with EX 86400. Every day at 02:00 the database and API latency spike badly. Also, a planned Redis migration to a new cluster is coming up. Diagnose and fix both.

You're given

  • 5M keys
  • Database handles ~5K queries/sec
  • Traffic at 02:00 is low but nonzero
  • Migration must not cause an outage

The interviewer follows up

01

How is an avalanche different from a stampede?

Explain it without notes

01

List four causes of a cache avalanche and one prevention for each.

Practice

01

Write a warming script that reads the 10,000 most accessed keys from yesterday's Top-K structure and pre-loads them into a new cluster.

Trade-offs

  • ↔

    Jitter spreads load but makes expiry times less predictable.

  • ↔

    Warming protects the database but takes time and must be kept in sync with what's actually hot.

Done when you can

  • I can prevent avalanches with jitter, warming, persistence and gradual rollouts.

  • I can explain the difference between avalanche and stampede.