Topic 11.3
Advanced Questions
In one line
Variants and production internals: cuckoo, blocked, partitioned, quotient and static filters, cache locality, LSM trees, RedisBloom, stale filters, rebuilds, and adversarial inputs.
Think of it like this
An advanced driving course. You're asked about skids, night driving and motorways, the situations where the basics are not enough.
Key ideas
- 01
Frame variant answers as "what problem it solves, what it costs".
- 02
For production questions, always say where the source of truth lives and how the filter stays complete.
Code & diagrams
Need deletes -> cuckoo (or counting, ~4x memory)
Unknown growth -> scalable Bloom
Speed / cache misses -> blocked Bloom (1 cache line per lookup)
Static set, min memory -> xor / binary fuse (~20-30% smaller)
Resizing, merging -> quotient filter
Shared across services -> RedisBloom BF.* / CF.*Explain it without notes
Compare Bloom and cuckoo filters.
What is a blocked Bloom filter?
What is a partitioned Bloom filter?
Why can Bloom filters have poor cache locality?
How do LSM trees use Bloom filters?
How does RocksDB use Bloom filters?
How does RedisBloom work?
How do you handle stale Bloom filters?
How do you rebuild a Bloom filter safely?
How do you protect against adversarial inputs?
Practice
Pick a variant for each: 1B static malware hashes, a user set with deletes, an unknown-growth event stream.
Trade-offs
- ↔
Every variant fixes one Bloom limitation by paying elsewhere: memory, insert failures, or rebuilds.
Done when you can
I can answer all ten advanced questions and justify a variant choice.