Topic 13.4
Large Messages and the Claim-Check Pattern
In one line
Kafka's default max message size is about 1 MB. Very large records (20 MB JSON) cause memory pressure, slow replication, consumer instability and head-of-line blocking. Store large payloads in object storage and send a small event with a reference (claim-check), or split and compress.
Think of it like this
A mailroom that handles letters well but chokes on pianos. For a piano, you send it via a freight company and mail a letter saying "your piano is at warehouse 7, shelf 3".
Key ideas
- 01
Limits: broker/topic
message.max.bytes/max.message.bytes(~1 MB default), producermax.request.size(1 MB), consumermax.partition.fetch.bytes(1 MB) andfetch.max.bytes, replica fetch sizes. Raising one without the others causesRecordTooLargeExceptionor stuck consumers. - 02
Why large messages hurt: producer and broker memory spikes, replication delays (followers fall out of ISR), consumers need big buffers, one huge record delays all records behind it in the partition, and page cache efficiency drops.
- 03
Claim-check: upload the payload to S3/GCS (with a content hash), publish an event with the object key, size, hash and metadata. Consumers fetch the object if they need it. Lifecycle rules delete objects after retention.
- 04
Alternatives: compress (a 20 MB JSON may compress to 2 MB), trim fields consumers don't need, or split into chunks with sequence numbers (complex; reassembly must handle missing chunks).
- 05
Guidance: keep typical records under ~100 KB, alarm on anything above 1 MB.
Code & diagrams
Interview problem
The problem
A product team wants 20 MB JSON events
A team wants to publish 20 MB JSON events to Kafka. Explain why you'd challenge the design and propose an alternative.
Explain it without notes
Which configs must all agree to send a 5 MB record, and what else suffers?
Practice
Send a 2 MB record with defaults and read the error, then implement a claim-check version.
Trade-offs
- ↔
Claim-check adds a storage dependency and an extra fetch, but keeps Kafka fast and predictable.
Done when you can
I can explain the costs of large records and implement the claim-check pattern.