Command Palette

Search for a command to run...

Hectal
PHASE 14Advanced ~8 min· topic 1 of 5

Topic 14.1

Metrics That Matter: Brokers, Producers, Consumers

In one line

Watch a small set of metrics: under-replicated, under-min-ISR and offline partitions, active controller count, request latency and handler idle percentages, bytes in/out, disk usage; producer error, retry and latency rates; consumer lag, rebalance rate and processing time. Alert on symptoms first.

0/5 · 0%

Think of it like this

A pilot's instrument panel. Dozens of gauges exist, but a few red lights (engine fire, low fuel, stall) demand action now; the rest help diagnose.

Key ideas

  1. 01

    Broker health: UnderReplicatedPartitions (> 0 means replicas lagging or brokers down), UnderMinIsrPartitionCount (> 0 means acks=all writes are failing for those partitions), OfflinePartitionsCount (> 0 means no leader: data unavailable), ActiveControllerCount (exactly 1 across controllers), IsrShrinksPerSec/IsrExpandsPerSec (flapping means unstable followers).

  2. 02

    Broker load: BytesInPerSec, BytesOutPerSec, MessagesInPerSec per broker and topic; RequestHandlerAvgIdlePercent and NetworkProcessorAvgIdlePercent (< 30% = saturated); request TotalTimeMs p99 for Produce and Fetch; disk usage and latency; leader count per broker (imbalance).

  3. 03

    Producer: record-error-rate, record-retry-rate, request-latency-avg/max, batch-size-avg, compression-rate-avg, buffer-available-bytes (near zero means the producer is blocking), throttle times.

  4. 04

    Consumer: records-lag-max and per-partition lag, records-consumed-rate, fetch-latency-avg, commit rate and failures, rebalance rate (rebalance-rate-per-hour, join/sync time), plus your own processing-time histogram and error counts.

  5. 05

    Collection: JMX exporter or the broker's metrics reporter into Prometheus; Burrow or lag exporters for lag; managed-service metrics in CloudWatch or equivalents. Dashboards per layer (cluster, topic, consumer group).

Code & diagrams

kafka-alerts.ymlyaml
groups:
- name: kafka
  rules:
  - alert: KafkaOfflinePartitions
    expr: sum(kafka_controller_kafkacontroller_offlinepartitionscount) > 0
    for: 1m
  - alert: KafkaUnderMinIsr
    expr: sum(kafka_server_replicamanager_underminisrpartitioncount) > 0
    for: 2m
  - alert: KafkaUnderReplicated
    expr: sum(kafka_server_replicamanager_underreplicatedpartitions) > 0
    for: 10m
  - alert: KafkaNoActiveController
    expr: sum(kafka_controller_kafkacontroller_activecontrollercount) != 1
    for: 1m
  - alert: KafkaRequestHandlersSaturated
    expr: avg by (instance) (kafka_server_kafkarequesthandlerpool_requesthandleravgidlepercent_oneminuterate) < 0.3
    for: 10m
  - alert: ConsumerLagGrowing
    expr: deriv(sum by (consumergroup) (kafka_consumergroup_lag)[15m:]) > 0 and sum by (consumergroup) (kafka_consumergroup_lag) > 10000
    for: 15m

Interview problem

The problem

Under-replicated partitions > 0

An alert fires: under-replicated partitions > 0. What do you investigate? Build the path through broker health, disk, network, replica lag, CPU, ISR, broker failure and partition imbalance.

Explain it without notes

01

Distinguish URP, under-min-ISR and offline partitions by impact.

Practice

01

Build a Grafana dashboard with the broker health metrics and trigger each alert in the lab by stopping brokers.

Trade-offs

  • ↔

    Too many alerts cause fatigue; alert on user-impacting symptoms and use the rest for diagnosis.

Done when you can

  • I know the key broker, producer and consumer metrics and their alert thresholds.