Command Palette

Search for a command to run...

Hectal
PHASE 9Advanced ~8 min· topic 1 of 4

Topic 9.1

Observability: Measuring What the Filter Is Really Doing

In one line

Track inserted count, fill ratio and estimated FPR (fill^k), query, positive and negative counts, and observed FPR (filter said maybe, source of truth said absent). Add filter version and age, rebuild duration and failures, updater lag, and a sampled false-negative check that should always read zero.

0/4 · 0%

Think of it like this

A smoke detector's test button and battery light. The device can't tell you by itself whether it's still working well; you need built-in indicators and regular tests.

Key ideas

  1. 01

    The filter alone can't know whether a "maybe" was right. Instrument the verification path: every "maybe" that the authoritative store resolves as absent is an observed false positive. Observed FPR = false positives ÷ non-member queries that reached verification.

  2. 02

    Estimated FPR from the filter: fill^k, computed from the bit popcount (cheap) or from the inserted count. If observed ≫ estimated, suspect bad hashing or skewed query patterns; if both are high, suspect saturation.

  3. 03

    Capacity metrics: inserted count vs design capacity, fill ratio (alert above ~55%), and for scalable filters the number of sub-filters.

  4. 04

    Health metrics: filter version and age, last rebuild time, duration and success, updater consumer lag and watermark age, Redis latency for shared filters, database fallback rate.

  5. 05

    The metric that should always be zero: sampled false negatives (filter said absent, a background check found the item exists). Any non-zero value is a correctness incident.

Code & diagrams

bloom-metrics.txttext
bloom_queries_total{result="negative"}          counter
bloom_queries_total{result="maybe"}             counter
bloom_verified_absent_total                     counter   (maybe -> DB not found = observed FP)
bloom_observed_fpr = verified_absent / (verified_absent + negatives)   # over non-members
bloom_fill_ratio                                gauge     (popcount / m)
bloom_estimated_fpr = fill_ratio ^ k            gauge
bloom_inserted_items / bloom_capacity           gauge
bloom_filter_version, bloom_filter_age_seconds  gauge
bloom_rebuild_duration_seconds, bloom_rebuild_failures_total
bloom_updater_lag_seconds                       gauge     (Kafka time lag / watermark age)
bloom_sampled_false_negatives_total             counter   MUST STAY 0

Interview problem

The problem

Expected FPR 1%, observed 8%

A filter designed for a 1% false-positive rate shows an observed rate of 8%. Investigate saturation, wrong capacity, poor hashing, incorrect insertion assumptions, configuration mismatch, measurement bugs and rebuild or version problems.

Explain it without notes

01

How do you measure a Bloom filter's real false-positive rate in production?

Practice

01

Build a Grafana dashboard for your filter with the metrics above and alerts for fill > 55%, observed FPR > 2× target and any sampled false negative.

Trade-offs

  • ↔

    Verification-based FPR measurement is accurate but only covers queries that reach the store; estimated FPR covers everything but assumes ideal hashing.

Done when you can

  • I can instrument a filter and measure observed and estimated FPR and false negatives.