Topic 0.4
How a Command Executes: RESP, the Event Loop, and Threads
In one line
A Redis command is a RESP message on a TCP socket. It's read and parsed, queued, executed on the main thread, and its reply is buffered back. Knowing this path tells you where latency comes from and why some operations block everyone.
Think of it like this
A post office with one clerk who stamps letters. Many assistants can open the envelopes and seal the replies (I/O threads), but only the clerk stamps, one letter at a time. If someone hands the clerk a 5,000-page parcel to stamp page by page, everyone waits.
Key ideas
- 01
RESP (REdis Serialization Protocol) is a simple text-prefixed protocol.
SET k vtravels as an array of bulk strings:*3\r\n$3\r\nSET\r\n$1\r\nk\r\n$1\r\nv\r\n. Replies start with a type byte:+simple string,-error,:integer,$bulk string,*array. RESP3 (Redis 6+, enabled withHELLO 3) adds maps, sets, doubles, booleans and push messages, which client-side caching uses for invalidations. - 02
The path of one command: the socket becomes readable → Redis reads bytes into the client's query buffer → parses a complete command → looks it up in the command table → checks ACL permissions, cluster slot and memory limits → executes on the main thread → appends the reply to the client's output buffer → writes it back when the socket is writable. Replication and AOF propagation happen right after execution.
- 03
Where latency hides: large requests (big query buffers), large replies (
HGETALLon a 1M-field hash builds a huge output buffer), slow commands (O(N)), fork for persistence (a pause proportional to memory size),fsyncwhen usingappendfsync always, transparent huge pages, swapping, and CPU contention with noisy neighbours. - 04
Background threads (
bio) do work that would block the main thread: closing big files,fsyncof the AOF, and lazy freeing of big values (UNLINK,FLUSHALL ASYNC, andlazyfree-*options). Separate from these, the persistence child is a forked process, not a thread. - 05
Output buffer limits protect memory:
client-output-buffer-limit normal 0 0 0,replica 256mb 64mb 60,pubsub 32mb 8mb 60. A slow subscriber or replica that can't keep up is disconnected when it crosses the hard limit, or the soft limit for the given seconds. This is a common hidden reason for "my subscriber keeps disconnecting".
Code & diagrams
Talk to Redis without a client library to see the protocol itself.
printf '*3\r\n$3\r\nSET\r\n$4\r\nuser\r\n$4\r\nasha\r\n*2\r\n$3\r\nGET\r\n$4\r\nuser\r\n' | nc -q1 localhost 6379
+OK
$4
asha
# An error reply:
printf 'INCR user\r\n' | nc -q1 localhost 6379
-ERR value is not an integer or out of range`qbuf` and `omem` show per-client query and output buffers. Huge `omem` means a big reply or a slow reader.
127.0.0.1:6379> CLIENT LIST
id=7 addr=172.18.0.1:51844 laddr=172.18.0.2:6379 fd=8 name=api-7 age=312 idle=0 flags=N db=0 sub=0 psub=0 ssub=0 multi=-1 qbuf=26 qbuf-free=20448 argv-mem=10 obl=0 oll=0 omem=0 tot-mem=22426 events=r cmd=get user=default resp=2
id=9 addr=172.18.0.5:40110 ... flags=P ... sub=3 ... omem=33554432 ... cmd=subscribe
127.0.0.1:6379> CONFIG GET client-output-buffer-limit
1) "client-output-buffer-limit"
2) "normal 0 0 0 slave 268435456 67108864 60 pubsub 33554432 8388608 60"Interview problem
The problem
One client makes everyone slow
p99 latency for all services jumps from 1 ms to 400 ms for a few seconds, several times an hour. Redis CPU averages 15%. How do you find the cause, using what you know about how commands execute?
You're given
- Single Redis primary
- ~30K ops/sec
- Spikes last 1–5 seconds
The interviewer follows up
Why doesn't SLOWLOG show the network time?
When it breaks
A slow Pub/Sub subscriber keeps getting disconnected
What you see
The subscriber's output buffer exceeds the pubsub limit (32 MB hard, or 8 MB for 60 s), Redis closes the connection, and messages sent while it reconnects are lost.
Fix & prevent
Make consumers faster, shard channels, or switch to Streams if you need durability. Raise limits only with a memory budget in mind.
Replicas enter a full-resync loop
What you see
During a full sync the primary buffers new writes for the replica; if they exceed the replica output buffer limit the sync is aborted and restarted, forever.
Fix & prevent
Raise client-output-buffer-limit replica and repl-backlog-size to fit the write rate during sync time; use diskless sync; check network bandwidth.
Explain it without notes
Walk through the path of a single SET command from client socket to reply.
What does RESP3 add over RESP2, and which Redis feature depends on it?
Practice
Use nc or telnet to send a raw RESP INCR counter and read the integer reply.
Set latency-monitor-threshold 50, run DEBUG SLEEP 0.2 in the lab, then read LATENCY LATEST.
Trade-offs
- ↔
Bigger output buffer limits reduce disconnects for slow consumers but let one client consume a lot of memory; limits are a memory-safety valve.
- ↔
RESP3 gives richer replies and push messages, but older client libraries may not support it; check your client before relying on it.
Done when you can
I can read a RESP message and name the reply types.
I can list the steps a command goes through and where latency can hide.
I know what I/O threads, background threads and the fork child each do.