Command Palette

Search for a command to run...

Hectal
PHASE 7Advanced ~9 min· topic 3 of 4

Topic 7.3

Zero-Downtime Rebuilds and Versioned Filters

In one line

Rebuild a saturated or drifting filter without downtime: build v2 from a database scan, replay changes that happened during the build, validate it, switch readers atomically via a pointer or config, keep v1 until the switch completes, then retire it. Dual-write both versions during the transition.

0/4 · 0%

Think of it like this

Replacing a shop's printed price list. You print the new list in the back office, update it with any price changes made while printing, check it, then swap it onto the counter in one move, and only then throw away the old one.

Key ideas

  1. 01

    Full rebuild: scan the source of truth (streaming, paginated, from a replica), add every key to a new filter sized for projected growth. Record the scan's consistent snapshot position (LSN, timestamp, Kafka offsets).

  2. 02

    Catch-up: while building, new inserts keep happening. Either dual-write (updaters add to both v1 and v2 during the build) or replay the change stream from the snapshot position into v2 until caught up.

  3. 03

    Validation: check fill ratio and estimated FPR against expectations, spot-check known members (must all be present), measure FPR on a sample of known non-members.

  4. 04

    Switch: an atomic pointer (SET bf:products:current v2 in Redis, a config flag, or a volatile reference swap in-process). Readers resolve the pointer (cached briefly).

  5. 05

    Retire: after all readers use v2 (and a safety period), stop updating v1 and delete it. Keep the build artefact (snapshot) for fast restarts.

  6. 06

    Memory: during the transition both versions exist (2× filter memory); plan for it.

Code & diagrams

rebuild.mermaiddiagram
Rendering diagram…
rebuild.shbash
# 1. Create v2 sized for projected growth (2x current)
redis-cli BF.RESERVE bf:products:v2 0.001 240000000 NONSCALING
# 2. Updater config: dual-write enabled for v1 and v2
# 3. Bulk load from a replica, in batches
psql -h replica -Atc "COPY (SELECT id FROM products) TO STDOUT" \
  | xargs -n 1000 redis-cli BF.MADD bf:products:v2 > /dev/null
# 4. Validate
redis-cli BF.INFO bf:products:v2
# 5. Atomic switch
redis-cli SET bf:products:current v2
# 6. Later: stop dual-write, then
redis-cli DEL bf:products:v1

Interview problem

The problem

Replace a saturated v1 with v2, no downtime

Filter v1 is too saturated (observed FPR 8% vs a 1% target). Design building v2, validating it, switching atomically and retiring v1 without downtime, and discuss memory, consistency and incremental rebuilds.

When it breaks

Rebuild without dual-write or change replay

What you see

Items inserted between the scan snapshot and the switch are missing from v2; after the switch they get false 404s until the next rebuild.

Fix & prevent

Dual-write during the build or replay changes from the recorded snapshot position before switching.

Explain it without notes

01

Why is dual-writing needed during a rebuild?

Practice

01

Write the runbook for rebuilding a 500M-item filter in Redis with a 30-minute build time.

Trade-offs

  • ↔

    Zero-downtime rebuilds cost 2× memory temporarily and pipeline complexity; in-place rebuilds are simpler but cause outages or false negatives.

Done when you can

  • I can run a zero-downtime, versioned filter rebuild with dual-writes and validation.