Topic 4.3
Durability as a Combination: RF, acks, min.insync, and Friends
In one line
No single setting makes Kafka durable. Replication factor sets how many copies exist, acks=all makes producers wait for the ISR, min.insync.replicas sets the minimum copies for a write, unclean election decides what happens when all in-sync copies die, and producer idempotence, retries and the outbox stop loss and duplicates at the edges.
Think of it like this
Protecting a legal document. You need several copies (replication factor), a rule that the office won't say "filed" until at least two copies are made (min.insync with acks=all), a rule against using an unverified draft if the originals burn (no unclean election), and the author keeping their own copy until they get the receipt (outbox).
Key ideas
- 01
RF 1: no redundancy; any broker loss loses data. RF 2: survives one loss, but with min.insync=2 any single outage blocks writes, and with min.insync=1 writes continue with one copy. RF 3: the standard; tolerates one failure with min.insync=2. RF 5: tolerates two failures with min.insync=3, at 5× storage and replication traffic, used for the most critical data.
- 02
Fsync: Kafka doesn't fsync each write by default; durability comes from replication across machines (and ideally zones). Flushing per message (
flush.messages=1) kills throughput and is rarely worth it. - 03
The durable recipe: RF=3, min.insync.replicas=2, acks=all, enable.idempotence=true, unclean.leader.election.enable=false, rack-aware replica placement, and producers that never drop a record on send failure (outbox or persistent retry).
- 04
Consumers are part of durability too: commit after processing, and use
isolation.level=read_committedwhen producers use transactions so aborted records are never processed.
Code & diagrams
kafka-topics.sh --bootstrap-server $B --create --topic payments \
--partitions 24 --replication-factor 3 \
--config min.insync.replicas=2 \
--config unclean.leader.election.enable=false
# producer: acks=all enable.idempotence=true delivery.timeout.ms=120000
# broker: broker.rack=<az>, default.replication.factor=3, min.insync.replicas=2Interview problem
The problem
Acknowledged messages must survive one broker failure
Design producer and topic configuration so that any acknowledged message survives the loss of one broker, and explain why the combination matters rather than any single setting.
Explain it without notes
Explain why min.insync.replicas has no effect with acks=1.
Practice
Tabulate what survives for RF 2 and 3 with min.insync 1 and 2 under one and two broker failures.
Trade-offs
- ↔
Higher RF and min.insync increase durability and cost, and reduce write availability during failures.
Done when you can
I can design and justify the durable configuration combination.