Command Palette

Search for a command to run...

Hectal
PHASE 3Intermediate ~8 min· topic 3 of 5

Topic 3.3

Consumer Groups and Scaling

In one line

Within a consumer group, each partition is assigned to exactly one consumer, so a group's parallelism is capped by the partition count; extra consumers sit idle. Different groups read the same topic independently, each with its own offsets.

0/5 · 0%

Think of it like this

A team of cashiers sharing checkout lanes. Each lane has exactly one cashier at a time. With 4 lanes and 10 cashiers, 6 cashiers wait for a lane to free up. A second store (another consumer group) has its own cashiers serving the same customers' receipts independently.

Key ideas

  1. 01

    Assignment: one partition → at most one consumer in the group; one consumer → zero or more partitions. With P partitions and C consumers: if C < P, some consumers get several partitions; if C > P, C − P consumers are idle (useful only as hot standbys).

  2. 02

    Multiple groups: payment-group, inventory-group, analytics-group each receive every record on the topic and track their own offsets. That's how one event feeds many services.

  3. 03

    Scaling limits: to scale consumers, you need enough partitions, which is why partition count is planned for peak consumer parallelism (Topic 1.4). Within one consumer, you can also parallelise with threads per partition (Topic 3.5).

  4. 04

    Assignment strategies: RangeAssignor (co-locates the same partition numbers of multiple topics on one consumer, useful for joins), RoundRobinAssignor, StickyAssignor, and CooperativeStickyAssignor (sticky and incremental, the usual choice with the classic protocol). With the new KIP-848 protocol, the broker computes assignments.

Code & diagrams

groups.mermaiddiagram
Rendering diagram…
describe-groups.shbash
kafka-consumer-groups.sh --bootstrap-server $B --describe --group payment-group --members --verbose
GROUP          CONSUMER-ID        HOST        #PARTITIONS  ASSIGNMENT
payment-group  consumer-1-9f2...  /10.0.1.11  2            commerce.orders(0,1)
payment-group  consumer-2-a71...  /10.0.1.12  2            commerce.orders(2,3)
payment-group  consumer-3-c05...  /10.0.1.13  0            -

Interview problem

The problem

10 partitions and 20 consumers, then the reverse

You have 10 partitions and 20 consumers in one group. What happens? Then 20 partitions and 10 consumers: what changes?

Explain it without notes

01

Why does adding consumers beyond the partition count not increase throughput?

Practice

01

Start 1, then 3, then 8 console consumers in one group on a 6-partition topic and watch assignments with --describe --members.

Trade-offs

  • ↔

    Spare idle consumers speed up recovery but add rebalance participants; usually autoscale consumers up to the partition count instead.

Done when you can

  • I can predict assignments for any partition and consumer count and explain groups vs consumers.