Topic 2.3
acks and Producer Durability
In one line
acks decides when a write counts as successful: 0 (don't wait), 1 (leader wrote it), all (all in-sync replicas have it). Only acks=all combined with replication factor ≥ 3 and min.insync.replicas ≥ 2 makes an acknowledged record survive a broker failure.
Think of it like this
Sending a signed contract. acks=0 is dropping it in a mailbox and walking away; acks=1 is the receptionist confirming they got it (but the only copy is on their desk); acks=all is waiting until copies are filed in two separate offices before you consider it done.
Key ideas
- 01
acks=0: the producer doesn't wait for a response. Highest throughput, lowest latency; records can be lost silently, and retries can't happen because the producer never learns of failures. Only for data you can lose (some metrics). - 02
acks=1: the leader appends and responds. If the leader crashes before followers copy the record, and a follower becomes leader, the acknowledged record is gone. - 03
acks=all(-1): the leader waits until every replica in the current ISR has the record. If the ISR has shrunk to just the leader,acks=alldegenerates toacks=1, unlessmin.insync.replicasforbids it. - 04
min.insync.replicas(topic or broker config): withacks=all, a write is rejected withNotEnoughReplicasif the ISR is smaller than this. RF=3 + min.insync=2 tolerates one replica down while still accepting durable writes; two down means writes fail rather than becoming unsafe. - 05
Latency cost of
acks=all: an extra follower fetch round trip within the cluster, usually a few milliseconds; batching hides most of it.
Code & diagrams
# RF=3, min.insync.replicas=2, two brokers stopped, producer acks=all:
WARN [Producer clientId=payments] Got error produce response with correlation id 812 on topic-partition
payments-4, retrying (2147483646 attempts left). Error: NOT_ENOUGH_REPLICAS
...
ERROR send failed: org.apache.kafka.common.errors.TimeoutException:
Expiring 12 record(s) for payments-4:120000 ms has passed since batch creationInterview problem
The problem
Payment events must not be silently lost
Compare acks=0, acks=1 and acks=all for payment events, then design the full durability configuration including min.insync.replicas, replication factor and producer retries.
You're given
- Payment events
- 3 brokers across 3 AZs
- Must survive one broker loss
The interviewer follows up
Why not set min.insync.replicas=3 with RF=3?
When it breaks
acks=all with min.insync.replicas=1 (the broker default)
What you see
When followers lag and drop out of the ISR, the leader acknowledges alone; a leader crash then loses acknowledged records even though the producer used acks=all.
Fix & prevent
Set min.insync.replicas=2 on important topics (or as a broker default) with RF=3.
Explain it without notes
Why does acks=all need min.insync.replicas to be meaningful?
Practice
In the 3-node lab, produce with acks=all to a topic with min.insync.replicas=2, stop one broker, then a second, and record what happens.
Trade-offs
- ↔
Stronger acks cost a few milliseconds of latency and reject writes during multi-broker failures in exchange for no silent loss.
Done when you can
I can explain acks 0, 1 and all and why RF=3 + min.insync=2 + acks=all is the standard durable setup.