Topic 3.4
Libraries: Guava, Redis, RocksDB, and Others
In one line
Use proven implementations: Guava's BloomFilter (Java, thread-safe, serializable, mergeable), RedisBloom's BF.* commands (shared, scalable), storage-engine filters in RocksDB, Cassandra and HBase, and libraries like bits-and-blooms (Go) or pybloom-live (Python). Know their parameters and limits.
Think of it like this
Buying a certified fire extinguisher instead of building one. You still need to know how it works and where to mount it.
Key ideas
- 01
Guava:
BloomFilter.create(Funnels.stringFunnel(UTF_8), expectedInsertions, fpp),put,mightContain,expectedFpp(),approximateElementCount(),putAll(other)for merging compatible filters,writeTo/readFromfor persistence. Thread-safe since Guava 23. - 02
Redis (Redis 8 core, RedisBloom, or valkey-bloom on Valkey):
BF.RESERVE key error_rate capacity [EXPANSION n] [NONSCALING],BF.ADD,BF.MADD,BF.EXISTS,BF.MEXISTS,BF.INSERT,BF.INFO,BF.CARD,BF.SCANDUMP/BF.LOADCHUNKfor backup. Scaling by default (adds sub-filters when full). - 03
Storage engines: RocksDB (
bloom_bits_per_key, often 10, with cache-local Bloom and Ribbon filters), LevelDB (NewBloomFilterPolicy(10)), Cassandra (bloom_filter_fp_chanceper table), HBase (row or row+column blooms per HFile), Parquet (split-block Bloom filters per column chunk). - 04
Other languages: Go (
bits-and-blooms/bloom), Python (pybloom-live,rbloom), Rust (fastbloom,xorffor Xor filters). - 05
Library checks before adopting: hash algorithm and stability across versions, serialization compatibility, thread safety, maximum size, merge support, and whether scalable growth is automatic.
Code & diagrams
BloomFilter<CharSequence> seen = BloomFilter.create(Funnels.stringFunnel(UTF_8), 100_000_000, 0.01);
seen.put("https://example.com/a");
boolean maybe = seen.mightContain("https://example.com/a"); // true
double fpp = seen.expectedFpp(); // current estimate
long approx = seen.approximateElementCount();
BloomFilter<CharSequence> other = BloomFilter.create(Funnels.stringFunnel(UTF_8), 100_000_000, 0.01);
if (seen.isCompatible(other)) seen.putAll(other); // bitwise OR merge127.0.0.1:6379> BF.RESERVE bf:products 0.01 100000000 EXPANSION 2
OK
127.0.0.1:6379> BF.MADD bf:products p:1001 p:1002 p:1003
1) (integer) 1
2) (integer) 1
3) (integer) 1
127.0.0.1:6379> BF.MEXISTS bf:products p:1002 p:999999
1) (integer) 1
2) (integer) 0
127.0.0.1:6379> BF.INFO bf:products
1) Capacity 2) (integer) 100000000
3) Size 4) (integer) 119816312
5) Number of filters 6) (integer) 1
7) Number of items inserted 8) (integer) 3
9) Expansion rate 10) (integer) 2Explain it without notes
What should you verify before adopting a Bloom filter library in production?
Practice
Create the same 10M-item, 1% filter with Guava and RedisBloom and compare memory, lookup latency and measured FPR.
Trade-offs
- ↔
Local libraries are fastest but per-instance; Redis is shared and scalable but adds network latency and a dependency.
Done when you can
I know the main Bloom filter libraries and engine settings and what to check before using one.