Topic 14.3
Debugging Producers and Consumers
In one line
Most client errors map to a short list of causes: timeouts and NotEnoughReplicas (broker or ISR health), RecordTooLarge (size limits), serialization and schema errors, authentication and authorization failures, unknown topics, rebalance loops, commit failures, deserialization failures, processing timeouts and slow downstreams.
Think of it like this
A mechanic's fault-code manual. The dashboard shows a code; the manual tells you which part to check first. Kafka's client exceptions are those codes.
Key ideas
- 01
Producer:
TimeoutException: Expiring N record(s)(couldn't deliver within delivery.timeout.ms: broker down, network, NotEnoughReplicas, or producer overloaded),NotEnoughReplicasException(ISR below min.insync),RecordTooLargeException,SerializationException(schema or serializer mismatch),SaslAuthenticationException,TopicAuthorizationException,UnknownTopicOrPartitionException(typo or auto-create disabled), buffer exhaustion (max.block.msexceeded), duplicates (idempotence disabled). - 02
Consumer: lag growth (processing too slow), rebalance loops (poll interval exceeded, crashes, flapping pods),
CommitFailedException(partitions reassigned before commit),RecordDeserializationException(poison record or schema mismatch), processing timeouts, DLT growth, slow downstream database, consumer crashes (OOM with huge batches). - 03
First steps for any client problem: client logs with the exception chain; client metrics (error and retry rates, lag, rebalances); broker-side view of the same partitions (leader, ISR); recent changes (deploys, configs, schema registrations).
- 04
Make clients debuggable: set
client.id, log topic/partition/offset/key on errors, export client metrics, and propagate trace IDs in headers.
Code & diagrams
Symptom / exception Likely cause First check
TimeoutException (Expiring records) broker down, ISR < min, network, overload URP/under-min-ISR, broker logs, request latency
NotEnoughReplicasException ISR shrank below min.insync.replicas which brokers are out of ISR, why
RecordTooLargeException payload > max.request.size / topic limit payload size; claim-check pattern
SerializationException schema/serializer mismatch registry subject, compatibility, client config
TopicAuthorizationException missing ACL principal name, kafka-acls --list
CommitFailedException rebalanced away before commit max.poll.interval.ms, processing time
RecordDeserializationException poison record, wrong deserializer partition/offset in the error, DLT routing
Rebalance every few minutes slow polls, crashes, autoscaling flaps "leaving group" reasons in logs
Lag growing on all partitions slow processing, downstream latency processing-time histogram, DB/API p99Explain it without notes
A consumer logs CommitFailedException repeatedly. What's happening and how do you fix it?
Practice
Trigger each of five producer errors in the lab (wrong topic, oversize record, missing ACL, two brokers down, bad serializer) and record the exact exception.
Trade-offs
- ↔
Verbose client logging helps debugging but can be expensive at high throughput; log errors with context, sample the rest.
Done when you can
I can map common producer and consumer errors to causes and first checks.