Topic 12.6
Delayed Jobs, Retries, and Distributed Scheduling
In one line
A sorted set scored by run-at time is a delayed queue: workers atomically claim jobs whose score is ≤ now. Add exponential backoff with jitter for retries, a dead-letter set for poison jobs, idempotent handlers, and leader election or atomic claiming so schedules run once.
Think of it like this
A hospital ward's medication schedule on a board sorted by time. Nurses take the next due item and write their initials on it so no one else gives the same dose. If a dose can't be given, it's rescheduled later; after too many failures, the doctor is called.
Key ideas
- 01
Scheduling:
ZADD jobs:delayed <runAtMs> <jobId>with job data in a hashjob:{id}. Workers pollZRANGE jobs:delayed -inf <now> BYSCORE LIMIT 0 10and must claim atomically: a Lua script that removes due jobs from the sorted set and pushes them to a ready list or stream, so two workers can't take the same job. - 02
Retries: on failure,
attempts += 1, computedelay = base × 2^attemptscapped (5 s, 30 s, 5 min, 1 h) plus random jitter, andZADDback with the new run-at. After max attempts, move tojobs:deadfor inspection. - 03
Why jitter: if 10,000 jobs fail together (a downstream outage), identical backoff makes them all retry at the same instant, a retry storm that knocks the downstream over again. Jitter spreads them out.
- 04
Worker crash: once claimed jobs move to a ready stream with a consumer group, pending entries and
XAUTOCLAIMrecover them (Topic 5.3). Handlers must be idempotent because a job can run twice. - 05
Clocks: run-at times compared against worker clocks can drift. Use Redis
TIMEinside the claim script as the single clock. - 06
Recurring schedules (cron-like) across many instances: either one elected leader enqueues occurrences (leader lease
SET leader:scheduler <id> NX PX 10000, renewed every few seconds), or every instance triesSET run:{job}:{occurrence} 1 NX EX ...and only the winner enqueues. Record the last run so missed schedules can be caught up after downtime.
Code & diagrams
-- KEYS[1] = delayed zset, KEYS[2] = ready stream (same hash tag); ARGV[1] = batch size
local t = redis.call('TIME')
local now = tonumber(t[1]) * 1000 + math.floor(tonumber(t[2]) / 1000)
local due = redis.call('ZRANGE', KEYS[1], '-inf', now, 'BYSCORE', 'LIMIT', 0, tonumber(ARGV[1]))
for _, id in ipairs(due) do
redis.call('ZREM', KEYS[1], id)
redis.call('XADD', KEYS[2], '*', 'jobId', id)
end
return #dueimport random, time
BACKOFF = [5, 30, 300, 3600] # seconds
MAX_ATTEMPTS = len(BACKOFF)
def on_failure(r, job_id: str):
attempts = r.hincrby(f"job:{{jobs}}:{job_id}", "attempts", 1)
if attempts > MAX_ATTEMPTS:
r.zadd("{jobs}:dead", {job_id: time.time()})
return
base = BACKOFF[attempts - 1]
delay = base * random.uniform(0.8, 1.2) # jitter
run_at_ms = int((time.time() + delay) * 1000)
r.zadd("{jobs}:delayed", {job_id: run_at_ms})Interview problem
The problem
Delayed job queue with retries
A job fails and must be retried after 5 s, 30 s, 5 min and 1 h. Design it on Redis, covering worker crashes, duplicate execution, clock issues, retry storms, backoff, jitter and poison messages.
You're given
- 500K jobs/day
- Many workers
- Jobs call external APIs
- At most 4 retries
The interviewer follows up
Why not just use EXPIRE and keyspace notifications to trigger jobs?
When it breaks
Workers claim due jobs with ZRANGE then ZREM in separate calls
What you see
Two workers read the same due jobs before either removes them; both run them.
Fix & prevent
Claim atomically in Lua (or use ZPOPMIN when all popped items are due), then process idempotently.
Explain it without notes
Explain why retries need jitter.
Practice
Build the delayed queue with the claim script and run 3 workers; kill one mid-job and verify the job completes elsewhere exactly once in effect.
Trade-offs
- ↔
Polling intervals trade scheduling precision against Redis load; 100–500 ms is typical.
Done when you can
I can build a delayed queue with atomic claiming, backoff, jitter and a DLQ.
I can run recurring schedules exactly once in effect across many instances.