Command Palette

Search for a command to run...

Hectal
PHASE 2Intermediate ~9 min· topic 3 of 4

Topic 2.3

acks and Producer Durability

In one line

acks decides when a write counts as successful: 0 (don't wait), 1 (leader wrote it), all (all in-sync replicas have it). Only acks=all combined with replication factor ≥ 3 and min.insync.replicas ≥ 2 makes an acknowledged record survive a broker failure.

0/4 · 0%

Think of it like this

Sending a signed contract. acks=0 is dropping it in a mailbox and walking away; acks=1 is the receptionist confirming they got it (but the only copy is on their desk); acks=all is waiting until copies are filed in two separate offices before you consider it done.

Key ideas

  1. 01

    acks=0: the producer doesn't wait for a response. Highest throughput, lowest latency; records can be lost silently, and retries can't happen because the producer never learns of failures. Only for data you can lose (some metrics).

  2. 02

    acks=1: the leader appends and responds. If the leader crashes before followers copy the record, and a follower becomes leader, the acknowledged record is gone.

  3. 03

    acks=all (-1): the leader waits until every replica in the current ISR has the record. If the ISR has shrunk to just the leader, acks=all degenerates to acks=1, unless min.insync.replicas forbids it.

  4. 04

    min.insync.replicas (topic or broker config): with acks=all, a write is rejected with NotEnoughReplicas if the ISR is smaller than this. RF=3 + min.insync=2 tolerates one replica down while still accepting durable writes; two down means writes fail rather than becoming unsafe.

  5. 05

    Latency cost of acks=all: an extra follower fetch round trip within the cluster, usually a few milliseconds; batching hides most of it.

Code & diagrams

acks.mermaiddiagram
Rendering diagram…
not-enough-replicas.txttext
# RF=3, min.insync.replicas=2, two brokers stopped, producer acks=all:
WARN [Producer clientId=payments] Got error produce response with correlation id 812 on topic-partition
     payments-4, retrying (2147483646 attempts left). Error: NOT_ENOUGH_REPLICAS
...
ERROR send failed: org.apache.kafka.common.errors.TimeoutException:
      Expiring 12 record(s) for payments-4:120000 ms has passed since batch creation

Interview problem

The problem

Payment events must not be silently lost

Compare acks=0, acks=1 and acks=all for payment events, then design the full durability configuration including min.insync.replicas, replication factor and producer retries.

You're given

  • Payment events
  • 3 brokers across 3 AZs
  • Must survive one broker loss

The interviewer follows up

01

Why not set min.insync.replicas=3 with RF=3?

When it breaks

acks=all with min.insync.replicas=1 (the broker default)

What you see

When followers lag and drop out of the ISR, the leader acknowledges alone; a leader crash then loses acknowledged records even though the producer used acks=all.

Fix & prevent

Set min.insync.replicas=2 on important topics (or as a broker default) with RF=3.

Explain it without notes

01

Why does acks=all need min.insync.replicas to be meaningful?

Practice

01

In the 3-node lab, produce with acks=all to a topic with min.insync.replicas=2, stop one broker, then a second, and record what happens.

Trade-offs

  • ↔

    Stronger acks cost a few milliseconds of latency and reject writes during multi-broker failures in exchange for no silent loss.

Done when you can

  • I can explain acks 0, 1 and all and why RF=3 + min.insync=2 + acks=all is the standard durable setup.