Topic 9.1
Observability: Measuring What the Filter Is Really Doing
In one line
Track inserted count, fill ratio and estimated FPR (fill^k), query, positive and negative counts, and observed FPR (filter said maybe, source of truth said absent). Add filter version and age, rebuild duration and failures, updater lag, and a sampled false-negative check that should always read zero.
Think of it like this
A smoke detector's test button and battery light. The device can't tell you by itself whether it's still working well; you need built-in indicators and regular tests.
Key ideas
- 01
The filter alone can't know whether a "maybe" was right. Instrument the verification path: every "maybe" that the authoritative store resolves as absent is an observed false positive. Observed FPR = false positives ÷ non-member queries that reached verification.
- 02
Estimated FPR from the filter: fill^k, computed from the bit popcount (cheap) or from the inserted count. If observed ≫ estimated, suspect bad hashing or skewed query patterns; if both are high, suspect saturation.
- 03
Capacity metrics: inserted count vs design capacity, fill ratio (alert above ~55%), and for scalable filters the number of sub-filters.
- 04
Health metrics: filter version and age, last rebuild time, duration and success, updater consumer lag and watermark age, Redis latency for shared filters, database fallback rate.
- 05
The metric that should always be zero: sampled false negatives (filter said absent, a background check found the item exists). Any non-zero value is a correctness incident.
Code & diagrams
bloom_queries_total{result="negative"} counter
bloom_queries_total{result="maybe"} counter
bloom_verified_absent_total counter (maybe -> DB not found = observed FP)
bloom_observed_fpr = verified_absent / (verified_absent + negatives) # over non-members
bloom_fill_ratio gauge (popcount / m)
bloom_estimated_fpr = fill_ratio ^ k gauge
bloom_inserted_items / bloom_capacity gauge
bloom_filter_version, bloom_filter_age_seconds gauge
bloom_rebuild_duration_seconds, bloom_rebuild_failures_total
bloom_updater_lag_seconds gauge (Kafka time lag / watermark age)
bloom_sampled_false_negatives_total counter MUST STAY 0Interview problem
The problem
Expected FPR 1%, observed 8%
A filter designed for a 1% false-positive rate shows an observed rate of 8%. Investigate saturation, wrong capacity, poor hashing, incorrect insertion assumptions, configuration mismatch, measurement bugs and rebuild or version problems.
Explain it without notes
How do you measure a Bloom filter's real false-positive rate in production?
Practice
Build a Grafana dashboard for your filter with the metrics above and alerts for fill > 55%, observed FPR > 2× target and any sampled false negative.
Trade-offs
- ↔
Verification-based FPR measurement is accurate but only covers queries that reach the store; estimated FPR covers everything but assumes ideal hashing.
Done when you can
I can instrument a filter and measure observed and estimated FPR and false negatives.