Topic 12.2
KRaft: The Controller Quorum and the Metadata Log
In one line
In KRaft mode, a quorum of controllers (usually 3 or 5) replicates cluster metadata as an event log using the Raft consensus protocol. One active controller handles metadata changes; brokers fetch the metadata log and keep a local cache. Kafka 4.0 removed ZooKeeper entirely, and KRaft enabled faster failover and far more partitions.
Think of it like this
A committee of three secretaries keeping the official minutes. One chairs and writes each decision; the others copy every line, and a decision is official once a majority has it. If the chair falls ill, the others elect a new chair from whoever has the most complete minutes.
Key ideas
- 01
Why ZooKeeper existed: early Kafka used ZooKeeper to store metadata (brokers, topics, partition assignments, ACLs) and elect a single controller broker. It worked, but meant running a second distributed system, slow controller failover (the new controller had to reload all metadata from ZooKeeper), and practical limits around a few hundred thousand partitions.
- 02
KRaft: metadata is an append-only log (
__cluster_metadata) replicated among controller nodes with Raft. The active controller is the Raft leader. Brokers are observers that fetch the log and apply it to their metadata cache, so every broker has an up-to-date, versioned view. - 03
Roles:
process.roles=controller,broker, or both (combined mode, fine for small clusters and labs). Production usually runs dedicated controllers on separate nodes withcontroller.quorum.voters(static) or bootstrap servers with dynamic quorum membership (newer versions, KIP-853). - 04
Benefits: one system to operate, controller failover in seconds (the standby controllers already have the log in memory), and support for millions of partitions per cluster.
- 05
Quorum sizing: 3 controllers tolerate 1 failure, 5 tolerate 2. Place them in separate failure domains (three AZs).
- 06
Tools:
kafka-storage.sh format(initialise storage with a cluster ID),kafka-metadata-quorum.sh describe --status/--replication,kafka-metadata-shell.shto inspect the metadata log.
Code & diagrams
process.roles=controller
node.id=1001
controller.quorum.voters=1001@ctrl-a:9093,1002@ctrl-b:9093,1003@ctrl-c:9093
listeners=CONTROLLER://:9093
controller.listener.names=CONTROLLER
log.dirs=/var/lib/kafka/metadata
# one-time storage format with a shared cluster id
# kafka-storage.sh random-uuid
# kafka-storage.sh format -t <cluster-id> -c controller.propertieskafka-metadata-quorum.sh --bootstrap-controller ctrl-a:9093 describe --replication
NodeId DirectoryId LogEndOffset Lag LastFetchTimestamp LastCaughtUpTimestamp Status
1001 ... 913402 0 1727520011000 1727520011000 Leader
1002 ... 913402 0 1727520010950 1727520010950 Follower
1003 ... 913401 1 1727520010900 1727520010900 Follower
1 ... 913400 2 1727520010880 1727520010880 Observer
2 ... 913402 0 1727520010990 1727520010990 ObserverExplain it without notes
Why did Kafka replace ZooKeeper with KRaft?
Practice
Run a cluster with 3 dedicated controllers and 3 brokers in Docker and describe the quorum.
Trade-offs
- ↔
Combined broker+controller nodes are simpler but couple metadata stability to broker load; dedicated controllers isolate them.
Done when you can
I can explain KRaft roles, the metadata log, quorum sizing and why ZooKeeper was removed.