Command Palette

Search for a command to run...

Hectal
PHASE 12Advanced ~10 min· topic 6 of 6

Topic 12.6

Delayed Jobs, Retries, and Distributed Scheduling

In one line

A sorted set scored by run-at time is a delayed queue: workers atomically claim jobs whose score is ≤ now. Add exponential backoff with jitter for retries, a dead-letter set for poison jobs, idempotent handlers, and leader election or atomic claiming so schedules run once.

0/6 · 0%

Think of it like this

A hospital ward's medication schedule on a board sorted by time. Nurses take the next due item and write their initials on it so no one else gives the same dose. If a dose can't be given, it's rescheduled later; after too many failures, the doctor is called.

Key ideas

  1. 01

    Scheduling: ZADD jobs:delayed <runAtMs> <jobId> with job data in a hash job:{id}. Workers poll ZRANGE jobs:delayed -inf <now> BYSCORE LIMIT 0 10 and must claim atomically: a Lua script that removes due jobs from the sorted set and pushes them to a ready list or stream, so two workers can't take the same job.

  2. 02

    Retries: on failure, attempts += 1, compute delay = base × 2^attempts capped (5 s, 30 s, 5 min, 1 h) plus random jitter, and ZADD back with the new run-at. After max attempts, move to jobs:dead for inspection.

  3. 03

    Why jitter: if 10,000 jobs fail together (a downstream outage), identical backoff makes them all retry at the same instant, a retry storm that knocks the downstream over again. Jitter spreads them out.

  4. 04

    Worker crash: once claimed jobs move to a ready stream with a consumer group, pending entries and XAUTOCLAIM recover them (Topic 5.3). Handlers must be idempotent because a job can run twice.

  5. 05

    Clocks: run-at times compared against worker clocks can drift. Use Redis TIME inside the claim script as the single clock.

  6. 06

    Recurring schedules (cron-like) across many instances: either one elected leader enqueues occurrences (leader lease SET leader:scheduler <id> NX PX 10000, renewed every few seconds), or every instance tries SET run:{job}:{occurrence} 1 NX EX ... and only the winner enqueues. Record the last run so missed schedules can be caught up after downtime.

Code & diagrams

claim-due.lualua
-- KEYS[1] = delayed zset, KEYS[2] = ready stream (same hash tag); ARGV[1] = batch size
local t = redis.call('TIME')
local now = tonumber(t[1]) * 1000 + math.floor(tonumber(t[2]) / 1000)
local due = redis.call('ZRANGE', KEYS[1], '-inf', now, 'BYSCORE', 'LIMIT', 0, tonumber(ARGV[1]))
for _, id in ipairs(due) do
  redis.call('ZREM', KEYS[1], id)
  redis.call('XADD', KEYS[2], '*', 'jobId', id)
end
return #due
retry.pypython
import random, time

BACKOFF = [5, 30, 300, 3600]          # seconds
MAX_ATTEMPTS = len(BACKOFF)

def on_failure(r, job_id: str):
    attempts = r.hincrby(f"job:{{jobs}}:{job_id}", "attempts", 1)
    if attempts > MAX_ATTEMPTS:
        r.zadd("{jobs}:dead", {job_id: time.time()})
        return
    base = BACKOFF[attempts - 1]
    delay = base * random.uniform(0.8, 1.2)       # jitter
    run_at_ms = int((time.time() + delay) * 1000)
    r.zadd("{jobs}:delayed", {job_id: run_at_ms})
scheduler.mermaiddiagram
Rendering diagram…

Interview problem

The problem

Delayed job queue with retries

A job fails and must be retried after 5 s, 30 s, 5 min and 1 h. Design it on Redis, covering worker crashes, duplicate execution, clock issues, retry storms, backoff, jitter and poison messages.

You're given

  • 500K jobs/day
  • Many workers
  • Jobs call external APIs
  • At most 4 retries

The interviewer follows up

01

Why not just use EXPIRE and keyspace notifications to trigger jobs?

When it breaks

Workers claim due jobs with ZRANGE then ZREM in separate calls

What you see

Two workers read the same due jobs before either removes them; both run them.

Fix & prevent

Claim atomically in Lua (or use ZPOPMIN when all popped items are due), then process idempotently.

Explain it without notes

01

Explain why retries need jitter.

Practice

01

Build the delayed queue with the claim script and run 3 workers; kill one mid-job and verify the job completes elsewhere exactly once in effect.

Trade-offs

  • ↔

    Polling intervals trade scheduling precision against Redis load; 100–500 ms is typical.

Done when you can

  • I can build a delayed queue with atomic claiming, backoff, jitter and a DLQ.

  • I can run recurring schedules exactly once in effect across many instances.