Command Palette

Search for a command to run...

Hectal
Phase 12Advanced13 of 18 in Apache Kafka

Cluster Architecture & KRaft

What brokers do with each request, why Kafka replaced ZooKeeper with KRaft, how the controller quorum and metadata log work, controller failover, migrating from ZooKeeper, and day-2 cluster operations.

Kafka 4.0 (2025) completed a multi-year change: ZooKeeper is gone, and cluster metadata lives in Kafka's own Raft-replicated log managed by a controller quorum. Anything you learned about Kafka's control plane before then needs updating.

This phase covers the broker's request path, KRaft in depth, what happens when controllers fail, how existing clusters migrated, and the routine operations every Kafka team performs.

0/4 · 0%
4 topics ~32 min 8 code blocks & diagrams
Start with the first topic
1
12.1

Brokers: Request Handling, Leadership, and Balance

A broker accepts client and inter-broker requests on network threads, processes them on I/O threads (appends, fetches, metadata, group and transaction coordination), stores partitions, leads some and follows others, and hosts coordinators for groups and transactions. Balanced leadership and partition placement keep load even.

7 min 1 diagram 1 code practice

2
12.2

KRaft: The Controller Quorum and the Metadata Log

In KRaft mode, a quorum of controllers (usually 3 or 5) replicates cluster metadata as an event log using the Raft consensus protocol. One active controller handles metadata changes; brokers fetch the metadata log and keep a local cache. Kafka 4.0 removed ZooKeeper entirely, and KRaft enabled faster failover and far more partitions.

8 min 1 diagram 2 code practice

3
12.3

Controller Failure and What Keeps Working

If the active controller fails, the remaining voters elect a new leader via Raft within seconds; the new controller already has the metadata log and continues. During the gap, data traffic to existing partition leaders continues; only metadata operations (leader elections, topic creation, reassignments) wait. Losing quorum majority stops metadata changes but not existing data flow.

8 min 1 diagram practice

4
12.4

ZooKeeper-to-KRaft Migration and Day-2 Operations

ZooKeeper-based clusters migrate to KRaft on Kafka 3.x (3.9 is the last ZooKeeper-capable release) using a bridge phase where KRaft controllers take over metadata while brokers are rolled, before upgrading to 4.x. Routine operations include rolling upgrades, partition reassignment and rebalancing (often with Cruise Control), adding and removing brokers, and config management.

9 min 2 code practice