Command Palette

Search for a command to run...

Hectal
PHASE 6Intermediate ~7 min· topic 1 of 5

Topic 6.1

Consumer Retry Strategies: Backoff, Jitter, and Retry Storms

In one line

Transient failures (timeouts, 503s, deadlocks) deserve retries; permanent ones (bad data, validation errors) don't. Retry in place with exponential backoff and jitter for brief blips, move to delayed retry topics for longer outages, and never hammer a struggling downstream with immediate retries.

0/5 · 0%

Think of it like this

Calling a busy customer-service line. Redialling instantly a hundred times jams the line for everyone. Waiting a bit longer between each try, with a random extra delay, gives the line a chance to clear.

Key ideas

  1. 01

    Classify errors: transient (network timeout, HTTP 429/503, lock timeout) → retry; permanent (deserialization error, schema mismatch, business rule violation, 400) → don't retry, send to the DLT immediately.

  2. 02

    In-place (blocking) retry: retry the same record a few times with backoff inside the consumer. Preserves partition order, but blocks the whole partition while retrying, so keep total retry time well under max.poll.interval.ms.

  3. 03

    Exponential backoff with jitter: delay = min(cap, base × 2^attempt) × random(0.5, 1.5). Jitter spreads retries from many consumers so they don't all hit the recovering service at the same instant.

  4. 04

    Retry storms: when a downstream fails, every consumer retrying immediately multiplies its load by the retry count, keeping it down. Combine backoff with a circuit breaker: after repeated failures, stop calling and pause the partition for a while.

  5. 05

    Non-blocking retries: for outages longer than seconds, move the record to a delayed retry topic (next topic) so other records keep flowing, accepting that ordering for that key may change.

Code & diagrams

BackoffRetry.javajava
<T> T withRetry(Supplier<T> call, int maxAttempts) {
    for (int attempt = 0; ; attempt++) {
        try {
            return call.get();
        } catch (TransientException e) {
            if (attempt + 1 >= maxAttempts) throw e;
            long base = Math.min(10_000, 200L * (1L << attempt));          // 200ms, 400, 800 ... cap 10s
            long jittered = (long) (base * (0.5 + ThreadLocalRandom.current().nextDouble()));
            sleep(jittered);
        }
        // PermanentException is not caught here: it goes straight to the DLT handler.
    }
}

When it breaks

Every consumer retries 503s immediately in a tight loop

What you see

The payment service, already overloaded, receives many times its normal traffic; it never recovers while retries continue.

Fix & prevent

Exponential backoff with jitter, a circuit breaker, and pausing partitions during outages.

Explain it without notes

01

Why add jitter to exponential backoff?

Practice

01

Classify ten errors your consumers have seen as transient or permanent and decide the retry policy for each.

Trade-offs

  • ↔

    Blocking retries preserve order but stall the partition; non-blocking retries keep throughput but can reorder a key's events.

Done when you can

  • I classify errors and retry transient ones with backoff, jitter and circuit breaking.