Command Palette

Search for a command to run...

Hectal
PHASE 14Advanced ~7 min· topic 3 of 5

Topic 14.3

Debugging Producers and Consumers

In one line

Most client errors map to a short list of causes: timeouts and NotEnoughReplicas (broker or ISR health), RecordTooLarge (size limits), serialization and schema errors, authentication and authorization failures, unknown topics, rebalance loops, commit failures, deserialization failures, processing timeouts and slow downstreams.

0/5 · 0%

Think of it like this

A mechanic's fault-code manual. The dashboard shows a code; the manual tells you which part to check first. Kafka's client exceptions are those codes.

Key ideas

  1. 01

    Producer: TimeoutException: Expiring N record(s) (couldn't deliver within delivery.timeout.ms: broker down, network, NotEnoughReplicas, or producer overloaded), NotEnoughReplicasException (ISR below min.insync), RecordTooLargeException, SerializationException (schema or serializer mismatch), SaslAuthenticationException, TopicAuthorizationException, UnknownTopicOrPartitionException (typo or auto-create disabled), buffer exhaustion (max.block.ms exceeded), duplicates (idempotence disabled).

  2. 02

    Consumer: lag growth (processing too slow), rebalance loops (poll interval exceeded, crashes, flapping pods), CommitFailedException (partitions reassigned before commit), RecordDeserializationException (poison record or schema mismatch), processing timeouts, DLT growth, slow downstream database, consumer crashes (OOM with huge batches).

  3. 03

    First steps for any client problem: client logs with the exception chain; client metrics (error and retry rates, lag, rebalances); broker-side view of the same partitions (leader, ISR); recent changes (deploys, configs, schema registrations).

  4. 04

    Make clients debuggable: set client.id, log topic/partition/offset/key on errors, export client metrics, and propagate trace IDs in headers.

Code & diagrams

error-map.txttext
Symptom / exception                     Likely cause                              First check
TimeoutException (Expiring records)     broker down, ISR < min, network, overload  URP/under-min-ISR, broker logs, request latency
NotEnoughReplicasException              ISR shrank below min.insync.replicas       which brokers are out of ISR, why
RecordTooLargeException                 payload > max.request.size / topic limit   payload size; claim-check pattern
SerializationException                  schema/serializer mismatch                 registry subject, compatibility, client config
TopicAuthorizationException             missing ACL                                principal name, kafka-acls --list
CommitFailedException                   rebalanced away before commit              max.poll.interval.ms, processing time
RecordDeserializationException          poison record, wrong deserializer          partition/offset in the error, DLT routing
Rebalance every few minutes             slow polls, crashes, autoscaling flaps     "leaving group" reasons in logs
Lag growing on all partitions           slow processing, downstream latency        processing-time histogram, DB/API p99

Explain it without notes

01

A consumer logs CommitFailedException repeatedly. What's happening and how do you fix it?

Practice

01

Trigger each of five producer errors in the lab (wrong topic, oversize record, missing ACL, two brokers down, bad serializer) and record the exact exception.

Trade-offs

  • ↔

    Verbose client logging helps debugging but can be expensive at high throughput; log errors with context, sample the rest.

Done when you can

  • I can map common producer and consumer errors to causes and first checks.