Command Palette

Search for a command to run...

Hectal
Phase 9Advanced10 of 12 in Bloom Filters

Operating Bloom Filters

Metrics that matter and how to measure the real false-positive rate, failure modes and fail-open vs fail-closed, fifteen failure-injection scenarios, security and membership inference, and complete capacity planning from 500M to 10B items.

A filter in production needs the same care as any stateful component: dashboards, alerts, failure plans, security review and capacity planning. This phase turns the earlier theory into an operating manual.

0/4 · 0%
4 topics ~32 min 4 code blocks & diagrams
Start with the first topic
1
9.1

Observability: Measuring What the Filter Is Really Doing

Track inserted count, fill ratio and estimated FPR (fill^k), query, positive and negative counts, and observed FPR (filter said maybe, source of truth said absent). Add filter version and age, rebuild duration and failures, updater lag, and a sampled false-negative check that should always read zero.

8 min 1 code practice

2
9.2

Failure Modes, Fail-Open vs Fail-Closed, and Failure Injection

When the filter is unavailable, fail open (skip it and query the source of truth, protected by rate limits) or fail closed (reject), decided by the use case: availability for cache protection, safety for security blocklists. Rehearse fifteen failures, from Redis and Kafka outages to hash changes and duplicate events.

8 min 1 code practice

3
9.3

Security: Membership Inference, Leakage, and Adversarial Inputs

A Bloom filter reveals membership to anyone who can query it or download it: attackers can test millions of candidate IDs (enumeration) or check whether a specific person is in a sensitive set. Protect filters like the data they summarise: access control, rate limits, no public distribution of sensitive filters, keyed hashing, and careful API design.

8 min 1 code practice

4
9.4

Capacity Planning from 500 Million to 10 Billion Items

Size filters from n (with growth), the target FPR (from the cost of a false positive) and the placement (local copies multiply memory). 500M at 0.1% ≈ 0.9 GB; 1B ≈ 1.8 GB; 10B ≈ 18 GB, which moves the design from a local filter to partitioned, sharded or static-compressed filters.

8 min 1 code practice