Topic 12.5
Event and Request Deduplication
In one line
SET event:{id} 1 NX EX <ttl> is a cheap way to drop repeated events, but it has gaps: the TTL window, marking before processing vs after, and Redis data loss. Durable deduplication belongs in the processing store, with Redis as a fast first filter.
Think of it like this
A receptionist stamping each delivery note "received". A second copy of the same note gets rejected, as long as the stamp book isn't thrown away (TTL) and the stamp is applied at the right moment (before or after the goods are checked in).
Key ideas
- 01
Basic pattern: when an event with ID
earrives,SET dedup:{e} 1 NX EX 86400. OK means new; nil means duplicate, so skip it. - 02
Marking before processing: if processing fails after marking, the event is marked done but never processed; retries are dropped, and the event is lost. Marking after processing: a crash between processing and marking lets a retry process it again (duplicate).
- 03
Better: mark as
PROCESSINGwith a short TTL, process, then setDONEwith a long TTL (the same claim pattern as idempotency keys). Or make processing itself idempotent (upserts, unique constraints) and use Redis only to skip obvious repeats cheaply. - 04
TTL window: duplicates arriving after the TTL are processed again. Size the TTL to the maximum redelivery delay of your source (Kafka consumer rewinds, webhook retry schedules), and keep durable dedup in the database for anything that must never repeat.
- 05
Memory: millions of event IDs per day × ~80 bytes each adds up; use short IDs, TTLs, or a Bloom filter per time window when occasional false "duplicate" answers are acceptable (they aren't for most business events).
Code & diagrams
def handle_webhook(r, db, event):
key = f"dedup:stripe:{event['id']}"
if not r.set(key, "PROCESSING", nx=True, ex=300):
state = r.get(key)
return "duplicate" if state == b"DONE" else "in-progress"
try:
with db.transaction():
# durable guard: UNIQUE(event_id) in processed_events
db.execute("INSERT INTO processed_events(event_id) VALUES (%s)", [event["id"]])
apply_business_logic(db, event)
r.set(key, "DONE", ex=7 * 86400)
return "processed"
except UniqueViolation:
r.set(key, "DONE", ex=7 * 86400)
return "duplicate"
except Exception:
r.delete(key) # allow a retry to process it
raiseInterview problem
The problem
The same event received 10 times
A webhook provider retries deliveries and sometimes sends the same event 10 times over two days. Your handler updates balances. Design deduplication and discuss TTL, races, the false assumption of exactly-once, and durable dedup.
The interviewer follows up
What if event IDs aren't provided?
Explain it without notes
What goes wrong if you mark an event as seen before processing it? After?
Practice
Send the same event concurrently from 5 threads and sequentially 3 times across a restart of Redis without persistence. Which duplicates does Redis catch, and which does the database constraint catch?
Trade-offs
- ↔
Redis dedup is fast and cheap but not durable; database dedup is durable but costs a write per event.
Done when you can
I can implement race-safe dedup with a claim state machine.
I size dedup TTLs to redelivery windows and keep durable dedup for money.