Topic 12.4
ZooKeeper-to-KRaft Migration and Day-2 Operations
In one line
ZooKeeper-based clusters migrate to KRaft on Kafka 3.x (3.9 is the last ZooKeeper-capable release) using a bridge phase where KRaft controllers take over metadata while brokers are rolled, before upgrading to 4.x. Routine operations include rolling upgrades, partition reassignment and rebalancing (often with Cruise Control), adding and removing brokers, and config management.
Think of it like this
Moving a company's records from an old filing system to a new one without closing the office. For a while both systems are kept in sync, staff switch over desk by desk, and only when everyone has moved is the old cabinet removed.
Key ideas
- 01
Migration path (high level): upgrade to a 3.x version that supports migration (ideally 3.9), deploy KRaft controllers configured with
zookeeper.metadata.migration.enable=true, which copy metadata from ZooKeeper; roll brokers into migration mode (still writing to ZooKeeper for rollback); roll brokers again as KRaft brokers; finalise by removing ZooKeeper config from controllers. After finalising, rollback isn't possible. Then upgrade to 4.x. - 02
Rolling upgrades: one broker at a time with controlled shutdown (leadership moves away first), waiting for under-replicated partitions to return to zero before the next. Upgrade controllers and brokers per the release notes;
metadata.version(feature level) is bumped after all nodes run the new version. - 03
Partition reassignment:
kafka-reassign-partitions.shwith a JSON plan moves replicas between brokers (adding brokers, decommissioning, rebalancing). Throttle replication (--throttle) so reassignment traffic doesn't starve production traffic. - 04
Cruise Control (LinkedIn's open-source tool) continuously models broker load (CPU, disk, network) and proposes or executes balanced reassignments; Strimzi integrates it with
KafkaRebalanceresources. - 05
Adding brokers doesn't move existing partitions automatically; new brokers get only new partitions until you rebalance. Removing a broker requires moving its replicas off first.
Code & diagrams
cat > move.json <<'EOF'
{"version":1,"partitions":[
{"topic":"commerce.orders","partition":0,"replicas":[4,2,3]},
{"topic":"commerce.orders","partition":1,"replicas":[2,4,1]}
]}
EOF
kafka-reassign-partitions.sh --bootstrap-server $B --reassignment-json-file move.json \
--execute --throttle 50000000 # 50 MB/s replication throttle
kafka-reassign-partitions.sh --bootstrap-server $B --reassignment-json-file move.json --verify
Status of partition reassignment:
Reassignment of partition commerce.orders-0 is completed.
Reassignment of partition commerce.orders-1 is still in progress.
# --verify also removes the throttle once everything is completefor each broker (one at a time):
1. check: under-replicated partitions == 0, offline partitions == 0
2. stop broker (controlled shutdown moves leadership away)
3. upgrade binaries/config, start broker
4. wait: broker rejoins ISR for all its partitions (URP back to 0)
5. optional: preferred leader election to restore leadership balance
after all nodes upgraded and stable:
kafka-features.sh --bootstrap-server $B upgrade --metadata <new version>Interview problem
The problem
Migrate a ZooKeeper-era cluster to KRaft
Your production cluster runs Kafka 3.4 with ZooKeeper, 30 brokers, 40K partitions. Management wants Kafka 4.x. Describe the migration plan, risks and rollback options.
When it breaks
Reassignment without a throttle on a busy cluster
What you see
Replication traffic saturates broker network and disks; produce latency spikes; followers drop out of ISR; under-min-ISR errors appear for producers.
Fix & prevent
Always throttle reassignments, move in batches, and monitor latency and URP.
Explain it without notes
Why doesn't adding brokers rebalance existing partitions?
Practice
Add a fourth broker to your lab and move half of a topic's replicas to it with a throttled reassignment.
Trade-offs
- ↔
Automated rebalancing (Cruise Control) keeps clusters balanced but moves data continually; manual reassignment is controlled but labour-intensive.
Done when you can
I can outline a ZooKeeper-to-KRaft migration and perform rolling upgrades and throttled reassignments.