Command Palette

Search for a command to run...

Hectal
PHASE 13Advanced ~10 min· topic 1 of 5

Topic 13.1

Hot Keys: One Key, Millions of Requests

In one line

A hot key concentrates traffic on one shard's single thread and network link, so adding shards doesn't help. Find hot keys with LFU-based tools, then spread or absorb the load: local caching, read replicas, key replication with random suffixes, sharded counters, and request coalescing.

0/5 · 0%

Think of it like this

A stadium with 40 entrances where everyone insists on using gate 7 because it's printed on the poster. The other 39 gates are idle; gate 7 has a two-hour queue. Adding gates doesn't help until you redirect people.

Key ideas

  1. 01

    Why it hurts: each key lives in one slot on one primary, and commands run on one thread. A single key at 500K reads/sec saturates that node's CPU or network while the rest of the cluster idles. Typical sources: a celebrity profile, a viral product, global config, a global counter, a trending list.

  2. 02

    Finding them: redis-cli --hotkeys (needs an LFU maxmemory-policy, since it reads OBJECT FREQ), per-node ops skew in monitoring, client-side sampling of keys per request, a Top-K structure fed by the app (Topic 4.4), and short MONITOR samples in emergencies (costly).

  3. 03

    Read-heavy fixes: an L1 in-process cache with a short TTL (the biggest win: 500K reads/sec becomes a few hundred), client-side caching with invalidation, read replicas for that shard, or key replication: store N copies (product:iphone:{0..7} in different slots) and read a random one, writing all copies on update.

  4. 04

    Write-heavy fixes: sharded counters (write random sub-keys, sum on read), local aggregation with periodic INCRBY, or moving the hot aggregate out of Redis into a stream-processing job.

  5. 05

    Request coalescing: many concurrent requests for the same key in one instance share one Redis call, which cuts load and also protects against stampedes.

  6. 06

    Prevention: design keys so no single key is a global bottleneck; beware hash tags that pull many hot keys onto one slot; load-test with realistic skew (Zipf distributions), not uniform keys.

Code & diagrams

find-hot-keys.shbash
redis-cli CONFIG SET maxmemory-policy allkeys-lfu
redis-cli --hotkeys
# Scanning the entire keyspace to find hot keys ...
[00.00%] Hot key '"product:iphone"' found so far with counter 255
[00.00%] Hot key '"config:flags"' found so far with counter 221
-------- summary -------
Sampled 1203455 keys in the keyspace!
hot key found with counter: 255   keyname: "product:iphone"

# Per-node skew in a cluster: one node far above the rest is a hot slot
for n in 10.0.0.1 10.0.0.2 10.0.0.3; do
  echo -n "$n "; redis-cli -h $n INFO stats | grep instantaneous_ops_per_sec
done
10.0.0.1 instantaneous_ops_per_sec:31022
10.0.0.2 instantaneous_ops_per_sec:412877      # <-- hot
10.0.0.3 instantaneous_ops_per_sec:29870
hot-key-fixes.mermaiddiagram
Rendering diagram…

Interview problem

The problem

Viral product: 5M requests per second

product:iphone receives 5M requests/sec during a launch. How do you stop one Redis shard from being overloaded?

You're given

  • 2,000 app instances
  • Product data changes every few minutes (price, stock badge)
  • Redis Cluster, 12 shards

The interviewer follows up

01

Why doesn't adding more shards fix a hot key?

When it breaks

A hash tag like {catalog} on all product keys

What you see

Every product lives in one slot; the whole catalogue's traffic hits one node and the cluster gives no scaling.

Fix & prevent

Remove broad tags; tag only small groups that need multi-key atomicity.

Explain it without notes

01

List three read-side and two write-side fixes for hot keys.

Practice

01

Generate Zipf-distributed traffic against a lab cluster and find the hot key with --hotkeys; then add a 1 s L1 cache and measure the drop in node ops.

Trade-offs

  • ↔

    L1 caching and key replication remove hot spots at the cost of staleness or more complex writes.

Done when you can

  • I can detect hot keys and explain why sharding doesn't fix them.

  • I can apply the right read or write mitigation.