Topic 4.2
Leader Election and Unclean Elections
In one line
When a partition leader fails, the controller elects a new leader from the ISR, so no committed record is lost. If no ISR member is alive, Kafka waits (unavailable) unless unclean leader election is enabled, which elects an out-of-sync replica and loses data to restore availability.
Think of it like this
A relay team choosing a replacement runner. They pick someone who's been training with the team (ISR), not a spectator. If no trained runner is available, they either stop the race (unavailability) or send a spectator who doesn't know the route (unclean election: the race continues, some distance is lost).
Key ideas
- 01
Detection: in KRaft, brokers send heartbeats to the active controller; missed heartbeats (
broker.session.timeout.ms) or a controlled shutdown tell the controller a broker is gone. - 02
Election: for each partition the failed broker led, the controller picks the first live replica in the ISR (preferring the replica order), updates metadata, and brokers and clients learn the new leader. Producers and consumers refresh metadata on errors like
NOT_LEADER_OR_FOLLOWERand continue. - 03
Controlled shutdown (a normal restart) moves leadership away before the broker stops, so clients see almost no interruption. Preferred leader election later moves leadership back to restore balance (
auto.leader.rebalance.enable). - 04
Unclean leader election (
unclean.leader.election.enable=true, default false): if all ISR members are dead, an out-of-sync replica becomes leader; records it didn't have are permanently lost, and consumers may see offsets reused with different data. Enable only where availability matters more than completeness (some metrics or logs). - 05
Newer Kafka versions add Eligible Leader Replicas (KIP-966) to make clean elections possible in more failure cases; check your version's release notes.
Code & diagrams
Interview problem
The problem
Broker B1 crashes while leading
Three brokers B1, B2, B3. A partition has leader B1 and ISR {B1, B2, B3}. B1 crashes. Explain exactly what happens to leadership, producers, consumers and data.
You're given
- RF 3, min.insync.replicas 2
- Producers acks=all, idempotent
- Consumers read committed data
The interviewer follows up
What if B2 and B3 were not in the ISR when B1 died?
When it breaks
unclean.leader.election.enable=true on a payments topic
What you see
After a double failure, an out-of-sync replica becomes leader; acknowledged payments vanish and offsets are reused with different records, confusing consumers.
Fix & prevent
Keep unclean election off for important data; use RF=3, min.insync=2 and rack-aware placement so double failures are rare.
Explain it without notes
What does unclean leader election trade?
Practice
In the lab, kill the leader of a partition (docker kill) while producing and verify the consumer count equals the producer's acknowledged count.
Trade-offs
- ↔
Clean elections protect data but can leave partitions offline in rare multi-failure cases.
Done when you can
I can walk through a leader failure and explain why acknowledged data survives.
I can explain unclean leader election and when (not) to enable it.