Topic 3.4
Rebalancing: Protocols, Timeouts, and Storms
In one line
A rebalance reassigns partitions when members join, leave, crash or exceed the poll interval. Eager rebalancing stops the whole group; cooperative rebalancing moves only affected partitions; static membership avoids rebalances on restarts; and Kafka 4.0's new consumer protocol (KIP-848) moves assignment to the broker with incremental, per-member reconciliation.
Think of it like this
Reassigning desks in an office. The old way: everyone stands up and leaves their desk every time one person joins or leaves (eager). The better way: only the people whose desks change move (cooperative). Giving everyone a named desk that's held for them while they step out means a coffee break doesn't reshuffle anyone (static membership).
Key ideas
- 01
Triggers: a member joins or leaves, a member misses heartbeats for
session.timeout.ms(45 s default since 3.0; heartbeats everyheartbeat.interval.ms, 3 s), a member doesn't poll withinmax.poll.interval.ms(5 min), subscription or partition count changes. - 02
Classic protocol, eager: all members revoke all partitions, rejoin, the group leader computes a new assignment, everyone resumes: a stop-the-world pause proportional to group size.
- 03
Classic protocol, cooperative (
CooperativeStickyAssignor): members keep partitions that don't move and only revoke those being reassigned, over two quick rounds. Most consumers keep processing throughout. - 04
Static membership (
group.instance.id, one stable ID per instance): a restarting member rejoins with the same ID within the session timeout and gets its old partitions back with no rebalance, ideal for rolling deploys on Kubernetes with StatefulSets. - 05
KIP-848 consumer protocol (GA in Kafka 4.0, opt in with
group.protocol=consumeron clients): the group coordinator on the broker computes assignments and reconciles each member incrementally via heartbeats, removing the global synchronisation barrier and making rebalances much less disruptive. - 06
Hooks:
ConsumerRebalanceListener.onPartitionsRevoked(commit offsets and flush state before losing partitions) andonPartitionsAssigned(load state, seek if needed).
Code & diagrams
# Classic protocol, modern settings
partition.assignment.strategy=org.apache.kafka.clients.consumer.CooperativeStickyAssignor
group.instance.id=payments-7 # static membership: stable per pod (e.g. StatefulSet ordinal)
session.timeout.ms=45000
heartbeat.interval.ms=3000
max.poll.interval.ms=300000
max.poll.records=200
# Kafka 4.0+ brokers and clients: the new consumer group protocol (KIP-848)
# group.protocol=consumer
# group.remote.assignor=uniformInterview problem
The problem
Rebalance storm with 100 consumers
100 consumers in a group keep joining and leaving (autoscaling, crashes, slow processing). Throughput collapses. Explain what's happening and fix it, covering session timeout, max.poll.interval.ms and consumer health.
The interviewer follows up
What's the difference between session.timeout.ms and max.poll.interval.ms?
When it breaks
Rolling deploy of 50 consumer pods with eager rebalancing
What you see
Each pod restart triggers two full rebalances (leave, then join); the deploy causes minutes of near-zero throughput and a lag spike.
Fix & prevent
Static membership plus cooperative rebalancing (or KIP-848); deploy with maxUnavailable tuned so pods return within the session timeout.
Explain it without notes
Compare eager, cooperative, static membership and the KIP-848 protocol.
Practice
Run 4 consumers with eager and then cooperative assignment; add a 5th and compare the processing pause in logs.
Trade-offs
- ↔
Shorter session timeouts detect crashes faster but cause false rebalances during GC pauses or network blips.
Done when you can
I can name every rebalance trigger and the settings that control it.
I use cooperative or KIP-848 rebalancing and static membership in production.