Command Palette

Search for a command to run...

Hectal
PHASE 12Advanced ~8 min· topic 3 of 4

Topic 12.3

Controller Failure and What Keeps Working

In one line

If the active controller fails, the remaining voters elect a new leader via Raft within seconds; the new controller already has the metadata log and continues. During the gap, data traffic to existing partition leaders continues; only metadata operations (leader elections, topic creation, reassignments) wait. Losing quorum majority stops metadata changes but not existing data flow.

0/4 · 0%

Think of it like this

A school principal falls ill. Classes continue with their current teachers (data traffic keeps flowing), but decisions like hiring or changing timetables wait until the deputies elect an acting principal, which they do quickly because they already have all the records.

Key ideas

  1. 01

    Detection: followers stop receiving fetch responses from the leader within controller.quorum.fetch.timeout.ms and start an election with a higher epoch; the voter with an up-to-date log that gathers a majority wins.

  2. 02

    During the election: produce and fetch requests to current partition leaders keep working, since brokers don't need the controller for normal data traffic. What pauses: electing new partition leaders for failed brokers, creating or deleting topics, config changes, partition reassignments, and fencing or registering brokers.

  3. 03

    Brokers fence themselves if they can't reach the active controller for too long (broker.session.timeout.ms on the controller side), so a broker isolated from the quorum eventually stops serving as a leader.

  4. 04

    Losing the majority (2 of 3 controllers down): no active controller; metadata is frozen. Existing leaders keep serving, but any additional broker failure can't be handled (no new leaders elected), so restoring quorum is urgent.

  5. 05

    Double failure scenario to rehearse: a broker failure while the controller quorum is down means its partitions go offline until the quorum returns.

Code & diagrams

controller-failover.mermaiddiagram
Rendering diagram…

Interview problem

The problem

The active controller crashes

The KRaft controller leader crashes. Explain the failure → election → quorum → new controller sequence and what happens to partition leaders, producers, consumers and metadata operations. Then: what if two of three controllers fail?

Explain it without notes

01

Why doesn't a controller failure stop producers and consumers?

Practice

01

In the lab, stop the active controller while producing, then stop a second controller, and try creating a topic.

Trade-offs

  • ↔

    Five controllers tolerate two failures but add a little metadata latency and cost.

Done when you can

  • I can explain KRaft controller failover and the impact of losing quorum.