Command Palette

Search for a command to run...

Hectal
PHASE 10Advanced ~10 min· topic 2 of 4

Topic 10.2

Replica Lag, Stale Reads, and Lost Acknowledged Writes

In one line

Because replication is asynchronous, a replica may not yet have a write the primary acknowledged, so replica reads can be stale and a failover can lose acknowledged writes. You manage this with read routing, WAIT, min-replicas-to-write, and designs that tolerate the gap.

0/4 · 0%

Think of it like this

Posting a letter and immediately phoning the recipient to ask "did you get it?". The post office accepted it (acknowledged), but the recipient hasn't received it yet (replica lag). If the post office burns down tonight, your letter may be gone even though they accepted it.

Key ideas

  1. 01

    Stale reads: a client writes to the primary and immediately reads from a replica that's 50 ms behind: it sees the old value. Users perceive this as "my change didn't save" or "I got logged out".

  2. 02

    Consistency models: strong (every read sees the latest write, which Redis replicas don't give), eventual (replicas converge), read-your-writes (a client sees its own writes), monotonic reads (a client never sees data go backwards, which breaks if it bounces between replicas with different lag), and session consistency.

  3. 03

    Getting read-your-writes: read from the primary after writing (for that user, for a few seconds: "sticky to primary"), or read from a replica only if its offset has passed the write's offset, or simply route all reads that must be fresh to the primary.

  4. 04

    WAIT numreplicas timeout blocks until the client's previous writes reach N replicas (or the timeout passes) and returns how many acknowledged. It reduces the chance of losing a write on failover but isn't a consensus protocol: a timeout returns fewer, and a failover can still pick a replica that didn't get it.

  5. 05

    min-replicas-to-write 1 and min-replicas-max-lag 10: the primary refuses writes if fewer than 1 replica has pinged within 10 s. This limits how much a primary isolated by a partition can accept (and later lose), at the cost of availability when replicas are down.

  6. 06

    When replica reads are fine: analytics dashboards, product catalogue browsing, leaderboards, anything where a second of staleness is acceptable. When not: sessions right after login, balances, inventory decisions, anything read-modify-write.

Code & diagrams

stale-read.mermaiddiagram
Rendering diagram…
wait.redisredis
127.0.0.1:6379> SET order:9:status PAID
OK
127.0.0.1:6379> WAIT 1 100          # wait up to 100 ms for 1 replica to acknowledge
(integer) 1
127.0.0.1:6379> CONFIG SET min-replicas-to-write 1
127.0.0.1:6379> CONFIG SET min-replicas-max-lag 10
# With the replica down for >10 s:
127.0.0.1:6379> SET x 1
(error) NOREPLICAS Not enough good replicas to write.

Interview problem

The problem

Username change and a 500 ms replica

A user changes their username. The primary updates immediately; the replica is 500 ms behind and serves profile reads. What can the user see, how would you design around it, should this read go to the replica, and is stale data acceptable here?

The interviewer follows up

01

Replica is 2 seconds behind. Should the application read from it?

When it breaks

Login writes the session to the primary; the next request reads it from a replica

What you see

Random "you've been logged out" right after login, especially under load when lag grows.

Fix & prevent

Read sessions from the primary, or retry on the primary when a replica returns nothing.

Explain it without notes

01

Why can a Redis failover lose writes that clients received OK for?

02

What do WAIT and min-replicas-to-write each protect against, and what don't they?

Practice

01

Measure lag: write a timestamp to the primary every 10 ms and read it from the replica; plot the difference under load.

Trade-offs

  • ↔

    Replica reads scale throughput and reduce primary load at the cost of staleness.

  • ↔

    WAIT and min-replicas-to-write improve durability at the cost of latency and availability.

Done when you can

  • I can explain stale reads, non-monotonic reads and lost acknowledged writes.

  • I can design read-your-writes and decide which reads may go to replicas.