Topic 9.2
Failure Modes, Fail-Open vs Fail-Closed, and Failure Injection
In one line
When the filter is unavailable, fail open (skip it and query the source of truth, protected by rate limits) or fail closed (reject), decided by the use case: availability for cache protection, safety for security blocklists. Rehearse fifteen failures, from Redis and Kafka outages to hash changes and duplicate events.
Think of it like this
A building's badge reader breaks. For the cafeteria door you prop it open (fail open: people get in, nobody is harmed). For the server room you keep it locked and call security (fail closed).
Key ideas
- 01
Fail open: treat every query as "maybe" and go to the authoritative store. Keeps the service available; database load rises to the unprotected level, so you need concurrency limits and rate limiting during the outage.
- 02
Fail closed: treat every query as "blocked" or "absent". Safe for security decisions, but can make the service unusable.
- 03
Middle paths: fail to a secondary check (the authoritative blocklist service, slower but correct), degrade features (skip recommendations filtering), or use the last known good local copy (usually available for local filters).
- 04
The questions for every failure: what breaks, what continues, can correctness be affected (false negatives?), can availability be affected, can queries be rejected wrongly, can memory grow, how do we recover, how do we detect it, how do we prevent recurrence.
Code & diagrams
Scenario Wrong rejections? Main effect / response
A Filter unavailable no (fail open) DB load up; limit concurrency, restore from snapshot
B Redis (shared filter) down no (fail open) same as A; local fallback copy if available
C Database down n/a negatives still answered; maybes fail -> degrade
D Kafka down possible (recent) updates stall; watermark bypass for new IDs
E Updater consumer crashes possible (recent) lag grows; alert, bypass recent IDs, restart
F Filter saturated no FPR high, DB load high; rebuild larger
G FPR increases no investigate hashing/capacity/measurement
H Stale filter (missed adds) YES false negatives; reliable updates, rebuild
I Rebuild fails no keep old version; alert; retry
J Partial filter update/load YES never publish partial filters; validate, atomic swap
K Region unavailable possible peer region loads snapshot + replays
L Version mismatch YES (if mixed m/k) header checks; keep both versions during rollout
M Hash algorithm changed YES new version only; never mix hashes
N Data deletion no stale bits -> false positives; rebuild or cuckoo
O Duplicate events no adds are idempotentInterview problem
The problem
Malware blocklist filter unavailable: continue or stop?
A malware-hash blocklist is checked with a Bloom filter on every file upload. The filter becomes unavailable. Should uploads continue or stop? Analyse security, availability, user impact and authoritative lookup cost.
Explain it without notes
Which failure scenarios can cause false negatives, and why are they the most important?
Practice
Inject scenarios E and M in a lab and verify your safeguards catch them.
Trade-offs
- ↔
Fail-open preserves availability at the cost of protection; fail-closed preserves safety at the cost of availability; the business decides.
Done when you can
I can choose fail-open or fail-closed per use case and analyse all fifteen failure scenarios.