Command Palette

Search for a command to run...

Hectal
PHASE 5Intermediate ~10 min· topic 2 of 5

Topic 5.2

Streams: An Append-Only Log in Redis

In one line

A stream is an append-only log of entries, each with a time-based ID and field-value pairs. You append with XADD, read ranges with XRANGE, tail it with XREAD BLOCK, and cap it with MAXLEN or MINID. Unlike Pub/Sub, entries stay until you trim them, so readers can catch up.

0/5 · 0%

Think of it like this

A ship's logbook. Every event is written on the next line with a timestamp, lines are never rewritten, and anyone can read from any point: "show me everything since 14:00". When the book gets too thick, old pages are archived (trimmed).

Key ideas

  1. 01

    XADD key [NOMKSTREAM] [MAXLEN|MINID [=|~] threshold] *|id field value ... appends and returns the ID. With *, the ID is <ms-timestamp>-<sequence>, for example 1727520000123-0, always increasing (if the clock goes backwards, Redis keeps using the last timestamp and bumps the sequence).

  2. 02

    Reading: XRANGE key - + COUNT 10 (oldest first), XREVRANGE key + - COUNT 10 (newest first), XRANGE key 1727520000000 1727523600000 (a time range), XLEN, XINFO STREAM. Tailing: XREAD COUNT 100 BLOCK 5000 STREAMS key $ waits for entries newer than now; replace $ with the last ID you saw to resume.

  3. 03

    Trimming: XADD ... MAXLEN ~ 100000 * keeps roughly the last 100K entries; MINID ~ <id> drops entries older than an ID (time-based retention). The ~ makes trimming approximate and much cheaper (whole internal nodes are removed). XTRIM trims explicitly; XDEL deletes single entries (rarely what you want).

  4. 04

    Internals: a radix tree of listpacks, so entries with the same fields are compressed well and IDs are stored as deltas. Appends are O(1); ranges are O(log N + M).

  5. 05

    Streams are part of the dataset: persisted in RDB/AOF and replicated. Durability is still Redis durability (async replication, fsync policy), which is weaker than Kafka's replicated log with acks.

  6. 06

    XREAD without groups gives every reader all entries (fan-out, like Pub/Sub with history). For work distribution, where each entry goes to one worker with acknowledgements, use consumer groups (next topic).

Code & diagrams

streams.redisredis
127.0.0.1:6379> XADD orders MAXLEN ~ 100000 * orderId 9 status CREATED amount 1299
"1727520000123-0"
127.0.0.1:6379> XADD orders * orderId 10 status CREATED amount 450
"1727520000456-0"
127.0.0.1:6379> XLEN orders
(integer) 2
127.0.0.1:6379> XRANGE orders - + COUNT 1
1) 1) "1727520000123-0"
   2) 1) "orderId" 2) "9" 3) "status" 4) "CREATED" 5) "amount" 6) "1299"
127.0.0.1:6379> XREAD COUNT 10 STREAMS orders 1727520000123-0     # everything after this ID
1) 1) "orders"
   2) 1) 1) "1727520000456-0"
         2) 1) "orderId" 2) "10" ...
127.0.0.1:6379> XREAD BLOCK 5000 STREAMS orders $                # wait for new entries
(nil)                                                            # none in 5 s
127.0.0.1:6379> XTRIM orders MINID ~ 1727433600000               # keep ~24 h
(integer) 0
stream-log.mermaiddiagram
Rendering diagram…

Interview problem

The problem

Activity feed with catch-up

Build an activity log for a collaborative document: every edit event is appended, clients show a live feed, and a client that reconnects after being offline for a minute must receive what it missed. Keep only the last 24 hours.

You're given

  • ~2,000 events/sec across 50K documents
  • Clients reconnect often (mobile)
  • 24-hour retention

The interviewer follows up

01

Why not Pub/Sub for this?

When it breaks

A stream without any trimming

What you see

It grows forever; memory creeps up until eviction or OOM. Because streams are one key, it also becomes a big key that's slow to replicate and delete.

Fix & prevent

Use MAXLEN ~ or MINID ~ on every XADD, or a scheduled XTRIM; alert on XLEN and memory per key.

Explain it without notes

01

What does a stream ID encode, and why is that useful?

02

Why is approximate trimming (~) preferred?

Practice

01

Append 5 entries, read the 3 newest with XREVRANGE, then read everything after the second entry's ID.

Trade-offs

  • ↔

    Streams give ordered, replayable history in Redis, but durability and retention are bounded by memory and Redis persistence.

  • ↔

    Per-entity streams scale and isolate; one global stream is simpler but becomes a hot, big key.

Done when you can

  • I can append, range-read, tail and trim a stream.

  • I can implement catch-up after reconnect using stream IDs.