Command Palette

Search for a command to run...

Hectal
PHASE 3Intermediate ~10 min· topic 5 of 5

Topic 3.5

Pipelining: Paying the Round Trip Once

In one line

Pipelining sends many commands without waiting for each reply, turning N round trips into roughly one. It's the single biggest throughput win for bulk work, but it's not atomic, and very large pipelines cost memory on both sides.

0/5 · 0%

Think of it like this

Mailing 100 letters. You can walk to the post office 100 times, one letter per trip, or take all 100 in one trip. The post office processes each letter separately either way (no atomicity), but you save 99 walks.

Key ideas

  1. 01

    Cost model: without pipelining, N commands cost N × RTT + N × server time. With pipelining, roughly 1 × RTT + N × server time + transfer. With RTT = 0.5 ms and 10,000 GETs: ~5 seconds sequentially vs ~10–30 ms pipelined.

  2. 02

    It also cuts syscalls: Redis reads many commands per read() and writes many replies per write(), so server CPU per command drops too. That's why redis-benchmark -P 16 shows up to ~10× more ops/sec.

  3. 03

    Pipelines aren't atomic: other clients' commands can interleave, and each command succeeds or fails on its own. Wrap commands in MULTI/EXEC inside the pipeline if you need atomicity with one round trip.

  4. 04

    Batch size: replies are buffered in Redis's output buffer and in the client until read. Pipelining 1M commands at once can use hundreds of MB. Use batches of ~100–1,000 commands and read replies per batch.

  5. 05

    Alternatives: multi-key commands (MGET, MSET, HMGET, DEL k1 k2 ...) are even cheaper than pipelined single-key commands. In Redis Cluster, multi-key commands need one slot, while pipelines work across slots: cluster-aware clients split the pipeline per node and run the parts in parallel.

  6. 06

    redis-cli --pipe does mass insertion from a file of raw commands, the fastest way to bulk-load data.

Code & diagrams

pipeline-vs-rtt.mermaiddiagram
Rendering diagram…
pipeline.pypython
import time, redis

r = redis.Redis()
keys = [f"product:{i}" for i in range(10_000)]

t = time.perf_counter()
for k in keys:
    r.get(k)                                   # 10,000 round trips
print(f"sequential: {time.perf_counter() - t:.2f}s")

t = time.perf_counter()
pipe = r.pipeline(transaction=False)           # plain pipeline, not MULTI
for i, k in enumerate(keys, 1):
    pipe.get(k)
    if i % 1000 == 0:
        pipe.execute()                         # send 1,000 at a time
pipe.execute()
print(f"pipelined:  {time.perf_counter() - t:.2f}s")

t = time.perf_counter()
for i in range(0, len(keys), 1000):
    r.mget(keys[i:i + 1000])                   # multi-key command
print(f"mget:       {time.perf_counter() - t:.2f}s")

# sequential: 4.91s   (RTT ~0.5 ms)
# pipelined:  0.06s
# mget:       0.03s
mass-insert.shbash
# Generate RESP-free inline commands and stream them in one go
for i in $(seq 1 1000000); do echo "SET key:$i value:$i"; done > data.txt
cat data.txt | redis-cli --pipe
All data transferred. Waiting for the last reply...
Last reply received from server.
errors: 0, replies: 1000000

Interview problem

The problem

10,000 GETs

A report endpoint needs 10,000 values from Redis. Compare 10,000 sequential round trips with pipelining and with MGET. Calculate the latency, and explain the differences in a cluster.

You're given

  • RTT = 0.5 ms
  • Server time ~1 µs per GET
  • Values ~200 bytes
  • Redis Cluster with 6 primaries

The interviewer follows up

01

Does pipelining make Redis execute commands faster?

When it breaks

A job pipelines 5 million commands in one batch

What you see

Redis builds a huge output buffer for the replies, memory spikes, and the client runs out of memory buffering replies. Other clients may see latency spikes.

Fix & prevent

Batch 100–1,000 commands per execute(); for loads, use redis-cli --pipe or chunked multi-key commands.

Assuming a pipeline is atomic

What you see

A pipeline that "moves" data with GET then SET then DEL races with other clients; partial failures leave half-moved data.

Fix & prevent

Use MULTI inside the pipeline, a single command (LMOVE, COPY, RENAME), or Lua.

Explain it without notes

01

Compute the latency for 1,000 sequential commands at 1 ms RTT vs one pipeline.

02

Why do cluster clients split a pipeline by node?

Practice

01

Run pipeline.py in your lab (or its equivalent in your language) and record the three timings. Then add artificial latency with tc or a remote Redis and run again.

Trade-offs

  • ↔

    Bigger pipeline batches mean fewer round trips but more memory and longer head-of-line waits; 100–1,000 is a practical range.

  • ↔

    Pipelines work across cluster slots but give no atomicity; multi-key commands and transactions are atomic but need one slot.

Done when you can

  • I can calculate the latency of sequential vs pipelined commands.

  • I can explain pipeline vs transaction without hesitation.

  • I batch bulk operations and know how cluster clients handle pipelines.