Topic 14.4
Throughput Incident: A Broker-Side Debugging Tree
In one line
When throughput drops or latency spikes, check in order: client-side batching and compression changes, broker CPU and handler saturation, disk latency and fullness, network saturation, page cache pressure from lagging consumers, GC pauses, partition and leader skew, replication and ISR churn, reassignments, and controller problems.
Think of it like this
A factory line slowing down. You check whether raw materials arrive in smaller batches (producer changes), whether machines are overloaded (CPU), whether the warehouse is full (disk), whether trucks are stuck (network), and whether one station is doing all the work (skew).
Key ideas
- 01
Producer side: a deploy that set
linger.ms=0or disabled compression can multiply request counts; checkbatch-size-avg,records-per-request-avgand request rate per broker. - 02
CPU: TLS encryption, compression conversions (if topic compression differs from producer compression, the broker recompresses), too many small requests, and down-conversion for old clients all burn CPU.
- 03
Disk: saturation from replaying consumers reading old segments, reassignments, or disk full (brokers stop accepting writes on full log dirs). Check iostat latency and utilisation.
- 04
Network: replication plus consumer fan-out plus reassignments exceeding NIC capacity; cross-AZ throttling in some clouds.
- 05
Page cache: lagging consumers reading cold data evict hot pages, turning tail reads into disk reads for everyone.
- 06
GC and JVM: long pauses cause ISR shrinks and request timeouts. Skew: one broker leading hot partitions. Controller: metadata churn or quorum problems stall elections.
Code & diagrams
Interview problem
The problem
Kafka throughput suddenly drops by 70%
Producers report throughput down 70% and latency up. Build your debugging tree across CPU, disk, network, page cache, GC, broker load, partition skew, producer batching, compression, replication and consumer lag.
Explain it without notes
How can a single lagging consumer group slow down producers?
Practice
In the lab, run a steady producer, then start a consumer reading from the beginning of a large topic, and watch produce latency and disk reads.
Trade-offs
- ↔
Quotas and throttles protect production traffic at the cost of slower replays and reassignments.
Done when you can
I can walk a structured tree from a throughput drop to its cause.