Command Palette

Search for a command to run...

Hectal
PHASE 2Intermediate ~9 min· topic 4 of 4

Topic 2.4

Retries, Timeouts, the Idempotent Producer, and In-Flight Requests

In one line

When a produce request times out, the producer can't know whether the broker wrote the batch, so retrying can duplicate it. The idempotent producer (on by default since Kafka 3.0) attaches a producer ID and per-partition sequence numbers so brokers drop duplicates and keep order even with up to 5 in-flight requests.

0/4 · 0%

Think of it like this

Sending a bank transfer and the page times out. Did it go through? If you just click again, you might pay twice. A transfer reference number lets the bank say "already processed this one" and ignore the repeat.

Key ideas

  1. 01

    Timeouts: request.timeout.ms (30 s) for one request; delivery.timeout.ms (120 s) as the total budget for a record including retries; retry.backoff.ms between attempts; retries defaults to effectively infinite, bounded by the delivery timeout.

  2. 02

    The duplicate problem: the broker appends the batch, the response is lost (network blip), the producer times out and resends. Without idempotence, the batch is written twice.

  3. 03

    Idempotent producer: each producer gets a producer ID (PID) and epoch; each batch carries a sequence number per partition. The leader tracks the last sequence numbers per PID and partition, rejects duplicates (acknowledging them as success) and rejects out-of-order sequences. Works within a producer session; a restarted producer gets a new PID unless it's transactional.

  4. 04

    max.in.flight.requests.per.connection: how many unacknowledged requests per broker connection. Without idempotence, values > 1 can reorder records on retry (batch 1 fails, batch 2 succeeds, batch 1 retried after 2). With idempotence, ordering is preserved for up to 5 in flight.

  5. 05

    Idempotence requires acks=all, retries > 0 and max in-flight ≤ 5. Kafka 3.0+ enables it by default unless you set conflicting options, and silently disables it if you do, so check your config.

  6. 06

    Limits: idempotence dedupes producer retries, not application-level duplicates (your service calling send twice for the same business event after a crash). For those, use event IDs and idempotent consumers, or transactions (Phase 5).

Code & diagrams

duplicate-on-retry.mermaiddiagram
Rendering diagram…
reliable-producer.propertiesproperties
acks=all
enable.idempotence=true
max.in.flight.requests.per.connection=5
retries=2147483647
delivery.timeout.ms=120000
request.timeout.ms=30000
retry.backoff.ms=100
# If you set acks=1 or max.in.flight > 5, idempotence is disabled (or the client errors
# when enable.idempotence=true is explicit). Check the producer's startup log.

Interview problem

The problem

Duplicate PaymentCompleted events

A producer sends PaymentCompleted. The broker processes it, but the producer times out before receiving the response and retries. What can happen, and how does the idempotent producer help? What does it not protect against?

The interviewer follows up

01

Can max.in.flight.requests=1 replace idempotence for ordering?

When it breaks

A team sets acks=1 for speed on a Kafka 3.x client

What you see

Idempotence is disabled along with it; retries after timeouts now produce duplicates and can reorder records, which nobody notices until a reconciliation mismatch.

Fix & prevent

Keep acks=all with idempotence; tune batching for throughput instead.

Explain it without notes

01

How do producer IDs and sequence numbers prevent duplicates and reordering?

Practice

01

Use a toxiproxy between producer and broker to drop responses, with idempotence on and off, and count duplicates in the topic.

Trade-offs

  • ↔

    Idempotence costs almost nothing today; disabling it for speed trades correctness for negligible gains.

Done when you can

  • I can explain timeout-induced duplicates, the idempotent producer, and in-flight ordering.

  • I know what idempotence doesn't cover and how to cover it.