Topic 6.3
Consumer Lag and Backpressure
In one line
Lag is how far a group's committed offsets trail the end of the log. It grows whenever producers outpace consumers. Kafka handles this gracefully (records wait on disk), but lag means delay, and if it outgrows retention it means data loss. Diagnose why consumers are slow before adding consumers.
Think of it like this
An inbox at work. A backlog isn't a disaster by itself, since letters wait patiently. But if they arrive faster than you handle them for long enough, old ones get shredded by the archive policy (retention) before you read them.
Key ideas
- 01
Lag per partition = log end offset − committed offset. Group lag = sum. Time lag (how old the oldest unprocessed record is) is often more meaningful than record counts.
- 02
Absolute lag alone is not enough: watch its trend. Lag of 1M that's shrinking by 10K/s is recovering; lag of 10K growing every minute is a problem. Alert on growth rate and time lag.
- 03
Backpressure in Kafka is natural: producers aren't slowed by slow consumers (Kafka isn't a bounded queue), records accumulate on disk up to retention. The risk is consumers falling so far behind that retention deletes unread data.
- 04
Fixes, in order: make processing faster (batch database writes, async I/O, remove per-record API calls), add consumers up to the partition count, increase partitions if the ceiling is hit, split slow workflows to their own topic or group, and shed or sample load for non-critical consumers.
- 05
Tools:
kafka-consumer-groups --describe, consumer metricrecords-lag-max, Burrow or exporters for Prometheus, and managed-service lag metrics.
Code & diagrams
kafka-consumer-groups.sh --bootstrap-server $B --describe --group payment-service
GROUP TOPIC PARTITION CURRENT-OFFSET LOG-END-OFFSET LAG CONSUMER-ID
payment-service commerce.orders 0 10421 10433 12 consumer-1-...
payment-service commerce.orders 1 9811 9820 9 consumer-1-...
payment-service commerce.orders 2 2104410 12891022 10786612 consumer-2-... <-- one partition
payment-service commerce.orders 3 10210 10222 12 consumer-3-...Interview problem
The problem
Lag going 0 → 100 → 10K → 1M → 10M
Consumer lag climbs steadily from zero to 10 million. What could cause it? Build a diagnostic tree covering consumer CPU, database, external APIs, GC, network, partition assignment, consumer and partition counts, slow processing, rebalances and broker fetch latency.
When it breaks
Lag exceeds retention
What you see
Segments containing unread records are deleted; consumers jump to the earliest available offset (with auto.offset.reset=earliest) or error out. Data is permanently skipped.
Fix & prevent
Alert on time lag vs retention, temporarily increase retention during incidents, and size consumer capacity for peak.
Explain it without notes
Why is lag growth rate more important than absolute lag?
Practice
Slow a consumer artificially (sleep 50 ms per record), produce at 100 records/s, and chart lag; then fix it with batching.
Trade-offs
- ↔
Scaling consumers is easy up to the partition count; beyond it you need more partitions or smarter processing.
Done when you can
I can diagnose growing lag systematically and protect against retention loss.