Command Palette

Search for a command to run...

Hectal
PHASE 6Intermediate ~9 min· topic 3 of 5

Topic 6.3

Consumer Lag and Backpressure

In one line

Lag is how far a group's committed offsets trail the end of the log. It grows whenever producers outpace consumers. Kafka handles this gracefully (records wait on disk), but lag means delay, and if it outgrows retention it means data loss. Diagnose why consumers are slow before adding consumers.

0/5 · 0%

Think of it like this

An inbox at work. A backlog isn't a disaster by itself, since letters wait patiently. But if they arrive faster than you handle them for long enough, old ones get shredded by the archive policy (retention) before you read them.

Key ideas

  1. 01

    Lag per partition = log end offset − committed offset. Group lag = sum. Time lag (how old the oldest unprocessed record is) is often more meaningful than record counts.

  2. 02

    Absolute lag alone is not enough: watch its trend. Lag of 1M that's shrinking by 10K/s is recovering; lag of 10K growing every minute is a problem. Alert on growth rate and time lag.

  3. 03

    Backpressure in Kafka is natural: producers aren't slowed by slow consumers (Kafka isn't a bounded queue), records accumulate on disk up to retention. The risk is consumers falling so far behind that retention deletes unread data.

  4. 04

    Fixes, in order: make processing faster (batch database writes, async I/O, remove per-record API calls), add consumers up to the partition count, increase partitions if the ceiling is hit, split slow workflows to their own topic or group, and shed or sample load for non-critical consumers.

  5. 05

    Tools: kafka-consumer-groups --describe, consumer metric records-lag-max, Burrow or exporters for Prometheus, and managed-service lag metrics.

Code & diagrams

lag.shbash
kafka-consumer-groups.sh --bootstrap-server $B --describe --group payment-service
GROUP            TOPIC            PARTITION  CURRENT-OFFSET  LOG-END-OFFSET  LAG       CONSUMER-ID
payment-service  commerce.orders  0          10421           10433           12        consumer-1-...
payment-service  commerce.orders  1          9811            9820            9         consumer-1-...
payment-service  commerce.orders  2          2104410         12891022        10786612  consumer-2-...   <-- one partition
payment-service  commerce.orders  3          10210           10222           12        consumer-3-...
lag-tree.mermaiddiagram
Rendering diagram…

Interview problem

The problem

Lag going 0 → 100 → 10K → 1M → 10M

Consumer lag climbs steadily from zero to 10 million. What could cause it? Build a diagnostic tree covering consumer CPU, database, external APIs, GC, network, partition assignment, consumer and partition counts, slow processing, rebalances and broker fetch latency.

When it breaks

Lag exceeds retention

What you see

Segments containing unread records are deleted; consumers jump to the earliest available offset (with auto.offset.reset=earliest) or error out. Data is permanently skipped.

Fix & prevent

Alert on time lag vs retention, temporarily increase retention during incidents, and size consumer capacity for peak.

Explain it without notes

01

Why is lag growth rate more important than absolute lag?

Practice

01

Slow a consumer artificially (sleep 50 ms per record), produce at 100 records/s, and chart lag; then fix it with batching.

Trade-offs

  • ↔

    Scaling consumers is easy up to the partition count; beyond it you need more partitions or smarter processing.

Done when you can

  • I can diagnose growing lag systematically and protect against retention loss.