Command Palette

Search for a command to run...

Hectal
PHASE 3Intermediate ~9 min· topic 4 of 5

Topic 3.4

Rebalancing: Protocols, Timeouts, and Storms

In one line

A rebalance reassigns partitions when members join, leave, crash or exceed the poll interval. Eager rebalancing stops the whole group; cooperative rebalancing moves only affected partitions; static membership avoids rebalances on restarts; and Kafka 4.0's new consumer protocol (KIP-848) moves assignment to the broker with incremental, per-member reconciliation.

0/5 · 0%

Think of it like this

Reassigning desks in an office. The old way: everyone stands up and leaves their desk every time one person joins or leaves (eager). The better way: only the people whose desks change move (cooperative). Giving everyone a named desk that's held for them while they step out means a coffee break doesn't reshuffle anyone (static membership).

Key ideas

  1. 01

    Triggers: a member joins or leaves, a member misses heartbeats for session.timeout.ms (45 s default since 3.0; heartbeats every heartbeat.interval.ms, 3 s), a member doesn't poll within max.poll.interval.ms (5 min), subscription or partition count changes.

  2. 02

    Classic protocol, eager: all members revoke all partitions, rejoin, the group leader computes a new assignment, everyone resumes: a stop-the-world pause proportional to group size.

  3. 03

    Classic protocol, cooperative (CooperativeStickyAssignor): members keep partitions that don't move and only revoke those being reassigned, over two quick rounds. Most consumers keep processing throughout.

  4. 04

    Static membership (group.instance.id, one stable ID per instance): a restarting member rejoins with the same ID within the session timeout and gets its old partitions back with no rebalance, ideal for rolling deploys on Kubernetes with StatefulSets.

  5. 05

    KIP-848 consumer protocol (GA in Kafka 4.0, opt in with group.protocol=consumer on clients): the group coordinator on the broker computes assignments and reconciles each member incrementally via heartbeats, removing the global synchronisation barrier and making rebalances much less disruptive.

  6. 06

    Hooks: ConsumerRebalanceListener.onPartitionsRevoked (commit offsets and flush state before losing partitions) and onPartitionsAssigned (load state, seek if needed).

Code & diagrams

rebalance.mermaiddiagram
Rendering diagram…
consumer-rebalance.propertiesproperties
# Classic protocol, modern settings
partition.assignment.strategy=org.apache.kafka.clients.consumer.CooperativeStickyAssignor
group.instance.id=payments-7          # static membership: stable per pod (e.g. StatefulSet ordinal)
session.timeout.ms=45000
heartbeat.interval.ms=3000
max.poll.interval.ms=300000
max.poll.records=200

# Kafka 4.0+ brokers and clients: the new consumer group protocol (KIP-848)
# group.protocol=consumer
# group.remote.assignor=uniform

Interview problem

The problem

Rebalance storm with 100 consumers

100 consumers in a group keep joining and leaving (autoscaling, crashes, slow processing). Throughput collapses. Explain what's happening and fix it, covering session timeout, max.poll.interval.ms and consumer health.

The interviewer follows up

01

What's the difference between session.timeout.ms and max.poll.interval.ms?

When it breaks

Rolling deploy of 50 consumer pods with eager rebalancing

What you see

Each pod restart triggers two full rebalances (leave, then join); the deploy causes minutes of near-zero throughput and a lag spike.

Fix & prevent

Static membership plus cooperative rebalancing (or KIP-848); deploy with maxUnavailable tuned so pods return within the session timeout.

Explain it without notes

01

Compare eager, cooperative, static membership and the KIP-848 protocol.

Practice

01

Run 4 consumers with eager and then cooperative assignment; add a 5th and compare the processing pause in logs.

Trade-offs

  • ↔

    Shorter session timeouts detect crashes faster but cause false rebalances during GC pauses or network blips.

Done when you can

  • I can name every rebalance trigger and the settings that control it.

  • I use cooperative or KIP-848 rebalancing and static membership in production.