Command Palette

Search for a command to run...

Hectal
PHASE 9Advanced ~9 min· topic 1 of 5

Topic 9.1

Replication: Replicas, Lag, Consistency and Failover

In one line

A primary streams its WAL to replicas that replay it. Replicas scale reads and provide failover targets. Asynchronous replication is fast but can lose recent commits on failover and serves stale reads; synchronous replication waits for replica confirmation. Failover needs fencing to avoid split brain, and applications need read-your-writes routing.

0/5 · 0%

Think of it like this

A teacher (primary) writing on the board while assistants (replicas) copy it into their notebooks for students at the back. The assistants are always a little behind, and if the teacher leaves, one assistant takes over, possibly missing the last line written.

Key ideas

  1. 01

    Physical streaming replication: replicas receive WAL over a replication connection and replay it; they're read-only and byte-identical. Logical replication publishes row changes per table (selective, cross-version, used for upgrades and CDC).

  2. 02

    Replication lag = WAL written on the primary but not yet replayed on the replica. Caused by heavy write bursts, long queries on the replica conflicting with replay, network or disk limits. Measure with pg_stat_replication (write/flush/replay lag).

  3. 03

    Async vs sync: async (default) acknowledges commits before replicas have them, so failover can lose the last seconds (RPO > 0). synchronous_standby_names = 'ANY 1 (r1, r2)' makes commits wait for one replica's flush: RPO ≈ 0 at the cost of commit latency, and writes stall if no sync replica is available.

  4. 04

    Read-your-writes: after a user updates their profile, reading from a lagging replica shows old data. Options: read from the primary for N seconds after a write (sticky), route by session LSN (wait until the replica's replay LSN ≥ the write's LSN), or read critical pages from the primary.

  5. 05

    Failover: detect primary failure (Patroni with etcd, or a managed service), promote the most up-to-date replica, repoint clients (DNS, VIP, or proxy), and fence the old primary (stop it, revoke its access) so two primaries never accept writes: split brain.

Code & diagrams

replication.sqlsql
-- on the primary
SELECT client_addr, state, sync_state,
       write_lag, flush_lag, replay_lag,
       pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn)) AS behind
FROM pg_stat_replication;
--  client_addr | state     | sync_state | write_lag | flush_lag | replay_lag | behind
--  10.0.2.11   | streaming | quorum     | 00:00:00.0004 | 00:00:00.0011 | 00:00:00.0019 | 24 kB
--  10.0.3.12   | streaming | async      | 00:00:00.0021 | 00:00:00.0040 | 00:00:04.2031 | 61 MB

-- on a replica
SELECT pg_is_in_recovery(), now() - pg_last_xact_replay_timestamp() AS replay_delay;

-- read-your-writes by LSN: after a write, remember pg_current_wal_lsn() in the session,
-- then on a replica read only if pg_last_wal_replay_lsn() >= that LSN (else use primary)
failover.mermaiddiagram
Rendering diagram…

Interview problem

The problem

Stale reads after profile updates

Users update their display name and immediately see the old one; support tickets pile up. Reads go to 3 async replicas with ~200 ms typical lag, spiking to 30 s during batch jobs. Design a fix that keeps most reads on replicas.

When it breaks

Split brain after a network partition

What you see

The old primary keeps accepting writes from some clients while a replica is promoted; two diverging histories must be reconciled by hand, and some writes are lost.

Fix & prevent

Leader leases in a consensus store (Patroni/etcd), fencing (the old primary demotes itself when it loses the lease; STONITH), and clients that connect only via the leader endpoint.

Replica query cancelled by recovery conflicts

What you see

Long reports on replicas fail with "canceling statement due to conflict with recovery" because VACUUM on the primary removed rows they need.

Fix & prevent

A dedicated analytics replica with max_standby_streaming_delay raised, or hot_standby_feedback = on (which causes bloat on the primary); better, a warehouse.

Explain it without notes

01

What does synchronous replication guarantee and cost?

Practice

01

Set up a primary and replica with Docker, run a write burst, and observe replay_lag.

Trade-offs

  • ↔

    Replicas scale reads and availability but introduce staleness and failover complexity; sync replicas trade latency for zero data loss.

Done when you can

  • I can explain replication modes, measure lag, design read-your-writes, and describe safe failover.