Topic 13.5
Debugging a Latency Incident: 2 ms to 2 s
In one line
When Redis latency jumps, work through a structured tree: is it the client or the server; network; CPU; slow commands or big keys; hot keys; persistence forks and fsync; memory pressure, swapping or eviction; replication; connection storms; cluster migrations. Each branch has a specific piece of evidence.
Think of it like this
A doctor with a patient's fever doesn't guess; they check temperature, pulse, blood tests in order, ruling out causes one by one. Latency debugging is the same: a checklist of measurements, each confirming or eliminating a cause.
Key ideas
- 01
Client or server? Compare client-side latency with
redis-cli --latencyfrom the app host and from the Redis host. If Redis-host latency is fine but app latency is bad, suspect the network, the client (pool exhaustion, GC pauses, thread starvation) or DNS/TLS. - 02
Server-side intrinsic latency:
redis-cli --intrinsic-latency 30on the Redis host shows latency caused by the machine or VM (noisy neighbours, CPU steal). - 03
Slow commands:
SLOWLOG GET,LATENCY LATEST"command" events,INFO commandstats. Big keys:--bigkeys,--memkeys. - 04
Persistence:
LATENCYeventsfork,aof-fsync-always,aof-write-pending-fsync;latest_fork_usec; disk latency withiostat. THP enabled makes forks worse. - 05
Memory:
used_memorynearmaxmemorywith evictions (eviction work on the write path); RSS above physical memory or fragmentation ratio under 1 (swapping: checksi/soinvmstat); expire cycles of millions of keys. - 06
Traffic: ops/sec spike, hot keys (per-node skew), connection storms (
total_connections_receivedrate), blocked clients, large replies (omeminCLIENT LIST). - 07
Topology: replication full syncs (fork plus network), cluster resharding (
MIGRATEof big keys), failovers in progress.
Code & diagrams
# 1. From an app host, then from the Redis host
redis-cli -h redis-prod --latency-history -i 5
redis-cli --intrinsic-latency 30 # on the Redis host
# 2. Server evidence
redis-cli SLOWLOG GET 20
redis-cli LATENCY LATEST
redis-cli INFO stats | egrep "instantaneous_ops|total_connections_received|evicted_keys|expired_keys"
redis-cli INFO persistence | egrep "fork|bgsave_in_progress|aof_rewrite_in_progress"
redis-cli INFO memory | egrep "used_memory_human|used_memory_rss_human|mem_fragmentation_ratio"
redis-cli INFO replication | egrep "role|master_link_status|lag"
# 3. Host evidence
vmstat 1 5 # si/so > 0 means swapping
cat /sys/kernel/mm/transparent_hugepage/enabled
iostat -x 1 3 # disk latency during AOF/RDBInterview problem
The problem
Production incident: Redis latency 2 ms → 2 s
You're paged: Redis latency went from 2 ms to 2 seconds. Build your debugging tree and walk through it, naming the evidence for each branch: network, CPU, memory, big key, slow command, hot key, fork, persistence, replication, connection pool, client issue, cluster migration, server overload.
The interviewer follows up
Redis memory keeps increasing. How do you find the cause?
Explain it without notes
Explain how you'd tell whether a latency problem is on the Redis server or in the client.
Practice
In the lab, cause three kinds of latency (a big-key HGETALL, BGSAVE on a large dataset, and a DEBUG SLEEP) and find each using only the triage commands.
Trade-offs
- ↔
Deep diagnostics (
MONITOR, big-key scans) can add load during an incident; prefer cheap evidence first.
Done when you can
I can walk a structured latency debugging tree with evidence for each branch.
I can diagnose CPU saturation and unbounded memory growth.