Topic 10.4
Retention: Time, Size, and Cleanup Policies
In one line
With cleanup.policy=delete, whole closed segments are deleted once they're older than retention.ms or when the partition exceeds retention.bytes. With compact, Kafka keeps the latest record per key. compact,delete does both. Retention is approximate at segment granularity and drives storage cost.
Think of it like this
A CCTV system that keeps 7 days of footage in hour-long files. Each hour a whole file older than 7 days is deleted, not individual frames, so footage may occasionally be kept a little over 7 days.
Key ideas
- 01
Time-based:
retention.ms(default 7 days) deletes a closed segment when its newest record's timestamp is older than the retention window. - 02
Size-based:
retention.bytes(per partition, default -1 = unlimited) deletes the oldest segments while the partition is over the limit. Total topic size ≈ retention.bytes × partitions × replication factor. - 03
Granularity: the active segment is never deleted, so a low-traffic partition with a 1 GB segment can hold data far beyond retention until the segment rolls (
segment.msbounds this). - 04
Cleanup policies:
deletefor event streams,compactfor latest-state-per-key topics (Topic 10.5),compact,deleteto keep the latest per key but also drop very old keys. - 05
Retention is a business decision: replay window for consumers (how long can a consumer be down and still recover?), audit requirements, and cost. Longer retention on local disks gets expensive; tiered storage (Topic 10.6) changes the economics.
Code & diagrams
# 3 days or 50 GB per partition, whichever comes first; roll segments at least hourly
kafka-configs.sh --bootstrap-server $B --alter --entity-type topics --entity-name clickstream \
--add-config retention.ms=259200000,retention.bytes=53687091200,segment.ms=3600000
# Broker log shows deletions:
# INFO [UnifiedLog partition=clickstream-7] Deleting segment LogSegment(baseOffset=9912003, size=1073740911,
# lastModifiedTime=...) due to log retention time 259200000ms breachExplain it without notes
Why might data remain longer than retention.ms?
Practice
Set retention.ms=60000 and segment.ms=10000 on a lab topic, produce continuously, and watch segments disappear.
Trade-offs
- ↔
Longer retention enables recovery and replay; shorter retention cuts cost and reduces the window for slow consumers.
Done when you can
I can configure time and size retention and explain segment granularity.