Command Palette

Search for a command to run...

Hectal
PHASE 11Advanced ~9 min· topic 4 of 5

Topic 11.4

Cluster Failover and Failure Modes

In one line

When a primary stops responding for cluster-node-timeout, other primaries mark it PFAIL then FAIL, its replicas hold an election that needs a majority of primaries, and the winner takes over the slots with a higher config epoch. If a shard loses its primary and all replicas, the cluster stops serving by default.

0/5 · 0%

Think of it like this

Branch managers of a chain voting to replace a manager who stopped answering the phone. Only the missing manager's own deputies can stand, and they need votes from most of the other branch managers. If the branch has no deputies left, headquarters closes that branch's services instead of guessing.

Key ideas

  1. 01

    Detection: nodes ping each other; a node not answering for cluster-node-timeout (default 15 s) is marked PFAIL locally. When a majority of primaries report it via gossip, it becomes FAIL, broadcast to all.

  2. 02

    Election: the failed primary's replicas wait a short, rank-based delay (the replica with the most data goes first), then request votes. A replica wins with votes from a majority of primaries for that epoch, promotes itself, and claims the slots with a new, higher config epoch, which everyone accepts.

  3. 03

    Manual failover: CLUSTER FAILOVER on a replica does a safe switchover (waits until the replica has all of the primary's data, then swaps), used for maintenance. FORCE and TAKEOVER variants skip safety checks for emergencies.

  4. 04

    Full coverage: by default (cluster-require-full-coverage yes), if any slot has no working owner, the whole cluster stops accepting queries. Setting it to no lets healthy shards keep serving while failed slots error. cluster-allow-reads-when-down yes allows reads during a down state.

  5. 05

    Minority protection: a primary cut off from the majority of primaries stops accepting writes after cluster-node-timeout, which bounds split-brain writes. Acknowledged writes not yet replicated can still be lost at failover, as with Sentinel.

  6. 06

    Replica migration: with cluster-migration-barrier, spare replicas move automatically to primaries left without replicas, so the cluster keeps some redundancy after failures.

Code & diagrams

cluster-failover.mermaiddiagram
Rendering diagram…
cluster-nodes.redisredis
127.0.0.1:6379> CLUSTER NODES
a1b2... 10.0.0.1:6379@16379 master,fail - 1727520000 1727519990 3 disconnected 0-5460
c3d4... 10.0.0.4:6379@16379 master - 0 1727520010 7 connected 0-5460
e5f6... 10.0.0.2:6379@16379 master - 0 1727520011 2 connected 5461-10922
...
# On a replica, for planned maintenance of its primary:
127.0.0.1:6379> CLUSTER FAILOVER
OK

Interview problem

The problem

Primary A dies in a 3-shard cluster

A cluster has primaries A, B, C with one replica each (A1, B1, C1). A's host crashes. Explain detection, election, promotion and the topology update. Then: what if A1's host also died?

You're given

  • cluster-node-timeout 15000
  • cluster-require-full-coverage yes

The interviewer follows up

01

Why must a majority of primaries vote, not just the replicas?

When it breaks

Primaries and their replicas placed in the same availability zone

What you see

A zone outage takes out a whole shard (primary and replica), so with full coverage the entire cluster goes down.

Fix & prevent

Spread each shard across zones; verify placement after every failover, since roles move.

Explain it without notes

01

Explain PFAIL vs FAIL and why failover needs votes from primaries.

Practice

01

Stop a primary in your lab cluster and time how long until its replica is promoted; then stop both a primary and its replica and observe CLUSTER INFO.

Trade-offs

  • ↔

    Lower node timeouts mean faster failover and more false positives; full coverage protects correctness at the cost of total availability.

Done when you can

  • I can walk through a cluster failover step by step.

  • I know what full coverage means and how to place replicas.