Topic 10.5
Log Compaction and Tombstones
In one line
Compacted topics keep at least the latest record for each key, removing older values in the background. A record with a null value (a tombstone) marks a key as deleted and is itself removed after delete.retention.ms. Compaction turns a topic into a replayable, latest-state table: ideal for user profiles, configuration, CDC and Kafka Streams changelogs.
Think of it like this
An address book where you add a new entry each time someone moves. A tidy-up job periodically crosses out old addresses and keeps only the latest one per person; a "deceased" note removes the person entirely after a while.
Key ideas
- 01
Mechanics: the log cleaner copies closed segments, keeping only the latest offset for each key, and swaps in the cleaned segments. The "head" (recent, uncleaned) part may still contain duplicates; the "tail" is compacted. Offsets never change; compacted topics just have gaps.
- 02
Controls:
min.cleanable.dirty.ratio(0.5: clean when half the log is uncleaned),min.compaction.lag.ms(keep records at least this long before compacting, so consumers see every update for a while),max.compaction.lag.ms(force cleaning within a bound),segment.ms(compaction only touches closed segments). - 03
Tombstones: produce
key → nullto delete. Consumers see the tombstone and delete their copy; afterdelete.retention.ms(1 day), the tombstone itself is removed, so consumers reading from scratch after that never see the key at all. - 04
Keys are required on compacted topics; records without keys are rejected.
- 05
Uses: latest user profile per ID, feature flags, CDC topics, Kafka Streams KTable changelogs,
__consumer_offsetsitself. New consumers bootstrap full state by reading the compacted topic from the beginning.
Code & diagrams
Before cleaning (offset: key=value)
0: u1=v1 1: u2=v1 2: u1=v2 3: u3=v1 4: u2=v2 5: u1=v3 6: u3=null (tombstone)
After cleaning
1: (gone) 2: (gone) 4: u2=v2 5: u1=v3 6: u3=null <- tombstone kept for delete.retention.ms
After delete.retention.ms
4: u2=v2 5: u1=v3kafka-topics.sh --bootstrap-server $B --create --topic users.state --partitions 12 --replication-factor 3 \
--config cleanup.policy=compact \
--config min.compaction.lag.ms=3600000 \
--config delete.retention.ms=86400000 \
--config segment.ms=3600000
# delete user u3: produce a tombstone (null value)
kcat -b $B -t users.state -P -K: -Z <<'EOF'
u3:
EOFInterview problem
The problem
A topic holding the latest user profile state
Design a topic that maintains the latest user profile for 50M users so services can bootstrap and stay in sync. Why is compaction appropriate? What happens when a user is deleted (including GDPR erasure)?
The interviewer follows up
Why must values be full state rather than deltas?
When it breaks
Compacted topic with a huge segment.ms and low traffic
What you see
The active segment never rolls, so nothing is compacted; storage grows with every update and deleted users' data remains.
Fix & prevent
Set segment.ms (for example 1 hour) and max.compaction.lag.ms on compacted topics.
Explain it without notes
How do tombstones work and why are they eventually removed?
Practice
Produce 3 versions for 3 keys and a tombstone for one, force frequent compaction in the lab, and read from the beginning.
Trade-offs
- ↔
Compaction gives compact latest-state topics but loses history; keep a separate delete-policy topic if you need both.
Done when you can
I can design compacted topics, use tombstones, and handle deletion and erasure correctly.