Topic 5.1
At-Most-Once, At-Least-Once, Exactly-Once
In one line
The order of "process" and "commit offset" decides delivery semantics. Commit then process: a crash loses records (at-most-once). Process then commit: a crash reprocesses records (at-least-once, the usual default). Exactly-once needs either transactions (Kafka to Kafka) or idempotent processing that makes duplicates harmless.
Think of it like this
Paying bills from a pile. If you tick a bill as paid before paying it and then get interrupted, it never gets paid (at-most-once). If you pay first and get interrupted before ticking, you might pay it twice (at-least-once). Exactly-once means either paying and ticking in one indivisible step, or having the biller reject a second payment of the same invoice.
Key ideas
- 01
At-most-once: commit offsets before (or independent of) processing, as with auto-commit plus async processing. No duplicates, possible loss. Acceptable for metrics or logs where occasional loss is fine.
- 02
At-least-once: process, then commit. A crash between them means the next owner reprocesses. No loss, possible duplicates. The standard for business events, paired with idempotent consumers.
- 03
Exactly-once (effectively-once): each record's effect happens once. In Kafka-to-Kafka pipelines, transactions make output records and input offsets commit atomically. For external systems, you get effectively-once via idempotency: dedupe by event ID or use natural idempotent operations (upsert, set status to X).
- 04
Producer side matters too: without idempotence, producer retries add duplicates before consumers even see them; with acks=0 or 1, records can be lost before consumers see them.
- 05
Say it precisely in interviews: "Kafka delivers at-least-once by default; I make processing idempotent so duplicates have no additional effect; where the pipeline is Kafka-to-Kafka, I use transactions for exactly-once."
Code & diagrams
Explain it without notes
State the three semantics in terms of where a crash happens.
Practice
For each consumer in a system you know, decide which semantics it needs and why.
Trade-offs
- ↔
Stronger semantics cost latency and complexity (transactions, dedup storage); pick the weakest that's correct for the use.
Done when you can
I can explain all three semantics and implement at-least-once with idempotent processing.