Command Palette

Search for a command to run...

Hectal
PHASE 9Advanced ~8 min· topic 2 of 4

Topic 9.2

Failure Modes, Fail-Open vs Fail-Closed, and Failure Injection

In one line

When the filter is unavailable, fail open (skip it and query the source of truth, protected by rate limits) or fail closed (reject), decided by the use case: availability for cache protection, safety for security blocklists. Rehearse fifteen failures, from Redis and Kafka outages to hash changes and duplicate events.

0/4 · 0%

Think of it like this

A building's badge reader breaks. For the cafeteria door you prop it open (fail open: people get in, nobody is harmed). For the server room you keep it locked and call security (fail closed).

Key ideas

  1. 01

    Fail open: treat every query as "maybe" and go to the authoritative store. Keeps the service available; database load rises to the unprotected level, so you need concurrency limits and rate limiting during the outage.

  2. 02

    Fail closed: treat every query as "blocked" or "absent". Safe for security decisions, but can make the service unusable.

  3. 03

    Middle paths: fail to a secondary check (the authoritative blocklist service, slower but correct), degrade features (skip recommendations filtering), or use the last known good local copy (usually available for local filters).

  4. 04

    The questions for every failure: what breaks, what continues, can correctness be affected (false negatives?), can availability be affected, can queries be rejected wrongly, can memory grow, how do we recover, how do we detect it, how do we prevent recurrence.

Code & diagrams

scenarios.txttext
Scenario                       Wrong rejections?   Main effect / response
A Filter unavailable           no (fail open)      DB load up; limit concurrency, restore from snapshot
B Redis (shared filter) down   no (fail open)      same as A; local fallback copy if available
C Database down                n/a                 negatives still answered; maybes fail -> degrade
D Kafka down                   possible (recent)   updates stall; watermark bypass for new IDs
E Updater consumer crashes     possible (recent)   lag grows; alert, bypass recent IDs, restart
F Filter saturated             no                  FPR high, DB load high; rebuild larger
G FPR increases                no                  investigate hashing/capacity/measurement
H Stale filter (missed adds)   YES                 false negatives; reliable updates, rebuild
I Rebuild fails                no                  keep old version; alert; retry
J Partial filter update/load   YES                 never publish partial filters; validate, atomic swap
K Region unavailable           possible            peer region loads snapshot + replays
L Version mismatch             YES (if mixed m/k)  header checks; keep both versions during rollout
M Hash algorithm changed       YES                 new version only; never mix hashes
N Data deletion                no                  stale bits -> false positives; rebuild or cuckoo
O Duplicate events             no                  adds are idempotent

Interview problem

The problem

Malware blocklist filter unavailable: continue or stop?

A malware-hash blocklist is checked with a Bloom filter on every file upload. The filter becomes unavailable. Should uploads continue or stop? Analyse security, availability, user impact and authoritative lookup cost.

Explain it without notes

01

Which failure scenarios can cause false negatives, and why are they the most important?

Practice

01

Inject scenarios E and M in a lab and verify your safeguards catch them.

Trade-offs

  • ↔

    Fail-open preserves availability at the cost of protection; fail-closed preserves safety at the cost of availability; the business decides.

Done when you can

  • I can choose fail-open or fail-closed per use case and analyse all fifteen failure scenarios.