Topic 3.5
Pipelining: Paying the Round Trip Once
In one line
Pipelining sends many commands without waiting for each reply, turning N round trips into roughly one. It's the single biggest throughput win for bulk work, but it's not atomic, and very large pipelines cost memory on both sides.
Think of it like this
Mailing 100 letters. You can walk to the post office 100 times, one letter per trip, or take all 100 in one trip. The post office processes each letter separately either way (no atomicity), but you save 99 walks.
Key ideas
- 01
Cost model: without pipelining, N commands cost N × RTT + N × server time. With pipelining, roughly 1 × RTT + N × server time + transfer. With RTT = 0.5 ms and 10,000 GETs: ~5 seconds sequentially vs ~10–30 ms pipelined.
- 02
It also cuts syscalls: Redis reads many commands per
read()and writes many replies perwrite(), so server CPU per command drops too. That's whyredis-benchmark -P 16shows up to ~10× more ops/sec. - 03
Pipelines aren't atomic: other clients' commands can interleave, and each command succeeds or fails on its own. Wrap commands in
MULTI/EXECinside the pipeline if you need atomicity with one round trip. - 04
Batch size: replies are buffered in Redis's output buffer and in the client until read. Pipelining 1M commands at once can use hundreds of MB. Use batches of ~100–1,000 commands and read replies per batch.
- 05
Alternatives: multi-key commands (
MGET,MSET,HMGET,DEL k1 k2 ...) are even cheaper than pipelined single-key commands. In Redis Cluster, multi-key commands need one slot, while pipelines work across slots: cluster-aware clients split the pipeline per node and run the parts in parallel. - 06
redis-cli --pipedoes mass insertion from a file of raw commands, the fastest way to bulk-load data.
Code & diagrams
import time, redis
r = redis.Redis()
keys = [f"product:{i}" for i in range(10_000)]
t = time.perf_counter()
for k in keys:
r.get(k) # 10,000 round trips
print(f"sequential: {time.perf_counter() - t:.2f}s")
t = time.perf_counter()
pipe = r.pipeline(transaction=False) # plain pipeline, not MULTI
for i, k in enumerate(keys, 1):
pipe.get(k)
if i % 1000 == 0:
pipe.execute() # send 1,000 at a time
pipe.execute()
print(f"pipelined: {time.perf_counter() - t:.2f}s")
t = time.perf_counter()
for i in range(0, len(keys), 1000):
r.mget(keys[i:i + 1000]) # multi-key command
print(f"mget: {time.perf_counter() - t:.2f}s")
# sequential: 4.91s (RTT ~0.5 ms)
# pipelined: 0.06s
# mget: 0.03s# Generate RESP-free inline commands and stream them in one go
for i in $(seq 1 1000000); do echo "SET key:$i value:$i"; done > data.txt
cat data.txt | redis-cli --pipe
All data transferred. Waiting for the last reply...
Last reply received from server.
errors: 0, replies: 1000000Interview problem
The problem
10,000 GETs
A report endpoint needs 10,000 values from Redis. Compare 10,000 sequential round trips with pipelining and with MGET. Calculate the latency, and explain the differences in a cluster.
You're given
- RTT = 0.5 ms
- Server time ~1 µs per GET
- Values ~200 bytes
- Redis Cluster with 6 primaries
The interviewer follows up
Does pipelining make Redis execute commands faster?
When it breaks
A job pipelines 5 million commands in one batch
What you see
Redis builds a huge output buffer for the replies, memory spikes, and the client runs out of memory buffering replies. Other clients may see latency spikes.
Fix & prevent
Batch 100–1,000 commands per execute(); for loads, use redis-cli --pipe or chunked multi-key commands.
Assuming a pipeline is atomic
What you see
A pipeline that "moves" data with GET then SET then DEL races with other clients; partial failures leave half-moved data.
Fix & prevent
Use MULTI inside the pipeline, a single command (LMOVE, COPY, RENAME), or Lua.
Explain it without notes
Compute the latency for 1,000 sequential commands at 1 ms RTT vs one pipeline.
Why do cluster clients split a pipeline by node?
Practice
Run pipeline.py in your lab (or its equivalent in your language) and record the three timings. Then add artificial latency with tc or a remote Redis and run again.
Trade-offs
- ↔
Bigger pipeline batches mean fewer round trips but more memory and longer head-of-line waits; 100–1,000 is a practical range.
- ↔
Pipelines work across cluster slots but give no atomicity; multi-key commands and transactions are atomic but need one slot.
Done when you can
I can calculate the latency of sequential vs pipelined commands.
I can explain pipeline vs transaction without hesitation.
I batch bulk operations and know how cluster clients handle pipelines.