Topic 14.1
Metrics That Matter: Brokers, Producers, Consumers
In one line
Watch a small set of metrics: under-replicated, under-min-ISR and offline partitions, active controller count, request latency and handler idle percentages, bytes in/out, disk usage; producer error, retry and latency rates; consumer lag, rebalance rate and processing time. Alert on symptoms first.
Think of it like this
A pilot's instrument panel. Dozens of gauges exist, but a few red lights (engine fire, low fuel, stall) demand action now; the rest help diagnose.
Key ideas
- 01
Broker health:
UnderReplicatedPartitions(> 0 means replicas lagging or brokers down),UnderMinIsrPartitionCount(> 0 means acks=all writes are failing for those partitions),OfflinePartitionsCount(> 0 means no leader: data unavailable),ActiveControllerCount(exactly 1 across controllers),IsrShrinksPerSec/IsrExpandsPerSec(flapping means unstable followers). - 02
Broker load:
BytesInPerSec,BytesOutPerSec,MessagesInPerSecper broker and topic;RequestHandlerAvgIdlePercentandNetworkProcessorAvgIdlePercent(< 30% = saturated); requestTotalTimeMsp99 for Produce and Fetch; disk usage and latency; leader count per broker (imbalance). - 03
Producer:
record-error-rate,record-retry-rate,request-latency-avg/max,batch-size-avg,compression-rate-avg,buffer-available-bytes(near zero means the producer is blocking), throttle times. - 04
Consumer:
records-lag-maxand per-partition lag,records-consumed-rate,fetch-latency-avg, commit rate and failures, rebalance rate (rebalance-rate-per-hour, join/sync time), plus your own processing-time histogram and error counts. - 05
Collection: JMX exporter or the broker's metrics reporter into Prometheus; Burrow or lag exporters for lag; managed-service metrics in CloudWatch or equivalents. Dashboards per layer (cluster, topic, consumer group).
Code & diagrams
groups:
- name: kafka
rules:
- alert: KafkaOfflinePartitions
expr: sum(kafka_controller_kafkacontroller_offlinepartitionscount) > 0
for: 1m
- alert: KafkaUnderMinIsr
expr: sum(kafka_server_replicamanager_underminisrpartitioncount) > 0
for: 2m
- alert: KafkaUnderReplicated
expr: sum(kafka_server_replicamanager_underreplicatedpartitions) > 0
for: 10m
- alert: KafkaNoActiveController
expr: sum(kafka_controller_kafkacontroller_activecontrollercount) != 1
for: 1m
- alert: KafkaRequestHandlersSaturated
expr: avg by (instance) (kafka_server_kafkarequesthandlerpool_requesthandleravgidlepercent_oneminuterate) < 0.3
for: 10m
- alert: ConsumerLagGrowing
expr: deriv(sum by (consumergroup) (kafka_consumergroup_lag)[15m:]) > 0 and sum by (consumergroup) (kafka_consumergroup_lag) > 10000
for: 15mInterview problem
The problem
Under-replicated partitions > 0
An alert fires: under-replicated partitions > 0. What do you investigate? Build the path through broker health, disk, network, replica lag, CPU, ISR, broker failure and partition imbalance.
Explain it without notes
Distinguish URP, under-min-ISR and offline partitions by impact.
Practice
Build a Grafana dashboard with the broker health metrics and trigger each alert in the lab by stopping brokers.
Trade-offs
- ↔
Too many alerts cause fatigue; alert on user-impacting symptoms and use the rest for diagnosis.
Done when you can
I know the key broker, producer and consumer metrics and their alert thresholds.