Command Palette

Search for a command to run...

Hectal
Phase 14Advanced15 of 18 in Apache Kafka

Observability & Debugging

The broker, producer and consumer metrics that matter, lag observability that tells you whether consumers are catching up, a debugging playbook for client and broker errors, a throughput-incident tree, and failure injection for every scenario.

A Kafka platform is only as good as your ability to see what's wrong. This phase gives you a short list of signals that map to real failures, then turns them into diagnosis trees for the incidents you'll actually get paged for.

It ends with fifteen failure scenarios to inject on purpose, each answered with what breaks, what's lost or duplicated, and what prevents it.

0/5 · 0%
5 topics ~39 min 6 code blocks & diagrams
Start with the first topic
1
14.1

Metrics That Matter: Brokers, Producers, Consumers

Watch a small set of metrics: under-replicated, under-min-ISR and offline partitions, active controller count, request latency and handler idle percentages, bytes in/out, disk usage; producer error, retry and latency rates; consumer lag, rebalance rate and processing time. Alert on symptoms first.

8 min 1 code practice

2
14.2

Lag Observability: Records, Time, and Trend

Measure lag per partition and group in records and in time, and track its trend: is the consumer catching up, holding steady or falling behind? Alert on time lag against SLOs and on sustained growth, not on absolute record counts.

7 min 1 code practice

3
14.3

Debugging Producers and Consumers

Most client errors map to a short list of causes: timeouts and NotEnoughReplicas (broker or ISR health), RecordTooLarge (size limits), serialization and schema errors, authentication and authorization failures, unknown topics, rebalance loops, commit failures, deserialization failures, processing timeouts and slow downstreams.

7 min 1 code practice

4
14.4

Throughput Incident: A Broker-Side Debugging Tree

When throughput drops or latency spikes, check in order: client-side batching and compression changes, broker CPU and handler saturation, disk latency and fullness, network saturation, page cache pressure from lagging consumers, GC pauses, partition and leader skew, replication and ISR churn, reassignments, and controller problems.

8 min 1 diagram practice

5
14.5

Failure Injection: Breaking Kafka on Purpose

Rehearse fifteen failures (producer network loss, broker crash, leader crash, lagging follower, consumer crash before commit, slow processing, rebalance storms, hot partition, full disk, network partition, incompatible schema, poison message, slow database, region failure, massive backlog) and answer ten questions for each.

9 min 2 code practice