Command Palette

Search for a command to run...

Hectal
PHASE 13Advanced ~10 min· topic 5 of 5

Topic 13.5

Debugging a Latency Incident: 2 ms to 2 s

In one line

When Redis latency jumps, work through a structured tree: is it the client or the server; network; CPU; slow commands or big keys; hot keys; persistence forks and fsync; memory pressure, swapping or eviction; replication; connection storms; cluster migrations. Each branch has a specific piece of evidence.

0/5 · 0%

Think of it like this

A doctor with a patient's fever doesn't guess; they check temperature, pulse, blood tests in order, ruling out causes one by one. Latency debugging is the same: a checklist of measurements, each confirming or eliminating a cause.

Key ideas

  1. 01

    Client or server? Compare client-side latency with redis-cli --latency from the app host and from the Redis host. If Redis-host latency is fine but app latency is bad, suspect the network, the client (pool exhaustion, GC pauses, thread starvation) or DNS/TLS.

  2. 02

    Server-side intrinsic latency: redis-cli --intrinsic-latency 30 on the Redis host shows latency caused by the machine or VM (noisy neighbours, CPU steal).

  3. 03

    Slow commands: SLOWLOG GET, LATENCY LATEST "command" events, INFO commandstats. Big keys: --bigkeys, --memkeys.

  4. 04

    Persistence: LATENCY events fork, aof-fsync-always, aof-write-pending-fsync; latest_fork_usec; disk latency with iostat. THP enabled makes forks worse.

  5. 05

    Memory: used_memory near maxmemory with evictions (eviction work on the write path); RSS above physical memory or fragmentation ratio under 1 (swapping: check si/so in vmstat); expire cycles of millions of keys.

  6. 06

    Traffic: ops/sec spike, hot keys (per-node skew), connection storms (total_connections_received rate), blocked clients, large replies (omem in CLIENT LIST).

  7. 07

    Topology: replication full syncs (fork plus network), cluster resharding (MIGRATE of big keys), failovers in progress.

Code & diagrams

latency-tree.mermaiddiagram
Rendering diagram…
triage.shbash
# 1. From an app host, then from the Redis host
redis-cli -h redis-prod --latency-history -i 5
redis-cli --intrinsic-latency 30          # on the Redis host
# 2. Server evidence
redis-cli SLOWLOG GET 20
redis-cli LATENCY LATEST
redis-cli INFO stats | egrep "instantaneous_ops|total_connections_received|evicted_keys|expired_keys"
redis-cli INFO persistence | egrep "fork|bgsave_in_progress|aof_rewrite_in_progress"
redis-cli INFO memory | egrep "used_memory_human|used_memory_rss_human|mem_fragmentation_ratio"
redis-cli INFO replication | egrep "role|master_link_status|lag"
# 3. Host evidence
vmstat 1 5                                 # si/so > 0 means swapping
cat /sys/kernel/mm/transparent_hugepage/enabled
iostat -x 1 3                              # disk latency during AOF/RDB

Interview problem

The problem

Production incident: Redis latency 2 ms → 2 s

You're paged: Redis latency went from 2 ms to 2 seconds. Build your debugging tree and walk through it, naming the evidence for each branch: network, CPU, memory, big key, slow command, hot key, fork, persistence, replication, connection pool, client issue, cluster migration, server overload.

The interviewer follows up

01

Redis memory keeps increasing. How do you find the cause?

Explain it without notes

01

Explain how you'd tell whether a latency problem is on the Redis server or in the client.

Practice

01

In the lab, cause three kinds of latency (a big-key HGETALL, BGSAVE on a large dataset, and a DEBUG SLEEP) and find each using only the triage commands.

Trade-offs

  • ↔

    Deep diagnostics (MONITOR, big-key scans) can add load during an incident; prefer cheap evidence first.

Done when you can

  • I can walk a structured latency debugging tree with evidence for each branch.

  • I can diagnose CPU saturation and unbounded memory growth.