Topic 3.3
Persistence, Serialization, and Versioning
In one line
An in-memory filter disappears on restart. Persist it with its metadata (m, k, hash algorithm and seed, encoding, format version, expected capacity, target FPR, build time), store it in Redis or object storage, and version it (filter:v1, filter:v2) so incompatible filters are never mixed.
Think of it like this
A spreadsheet saved without its column headers. The numbers survive, but nobody can tell what they mean. A filter's bits are useless without knowing which m, k and hash produced them.
Key ideas
- 01
Header fields: magic number and format version, m (bits), k, hash algorithm and seed, item encoding (UTF-8, lowercase-normalised), expected insertions and target FPR, inserted count, build timestamp and source snapshot (for example the database LSN or Kafka offset it was built up to).
- 02
Payload: the raw bit words, optionally compressed (a half-full filter compresses poorly; sparse ones compress well). Add a checksum to detect corruption.
- 03
Restart options: rebuild from the database (slow for big datasets, always correct), load from object storage or Redis (fast, then apply updates since the snapshot's position), or keep it in Redis permanently (shared, network lookups).
- 04
Versioning: key names or object paths include the version (
filter:products:v7). Readers check the header and refuse incompatible formats rather than guessing. Changing m, k, hash or encoding always means a new version. - 05
Rolling deployments: during a rollout, old and new app versions may expect different filters. Keep both versions available until the rollout completes, and make each app read the version it was built for, or make new code able to read old versions.
Code & diagrams
offset field value (example)
0 magic "BLMF"
4 formatVersion 2
6 hash "murmur3_128_mitz_64", seed 0
.. encoding "utf8-lowercase-nfkc"
.. numBits (m) 958505838
.. numHashFunctions (k) 7
.. expectedInsertions 100000000
.. targetFpp 0.01
.. insertedCount 98231145
.. builtAt 2026-09-28T02:00:00Z
.. sourcePosition kafka products-changes offsets {0: 991231, 1: 988112, ...}
.. crc32 0x5f1c9a2e
.. words[] long[14976654]// Guava: header (strategy, k, bit count) + bits
try (var out = new BufferedOutputStream(Files.newOutputStream(path))) {
filter.writeTo(out);
}
BloomFilter<CharSequence> restored;
try (var in = new BufferedInputStream(Files.newInputStream(path))) {
restored = BloomFilter.readFrom(in, Funnels.stringFunnel(UTF_8));
}
// Upload to object storage under a versioned key, e.g. s3://acme-filters/products/v7/filter.bin
// plus a small JSON manifest: { "version": 7, "sourcePosition": ..., "expectedInsertions": ... }Interview problem
The problem
Restart, and a rolling deployment with two filter versions
The application restarts and its in-memory filter is lost. Compare rebuilding, persisting locally, loading from Redis and loading from object storage. Separately: during a rollout, 50 of 100 instances use filter v1 and 50 use v2. How do you deploy safely?
When it breaks
Filter bits restored with a different hash seed or encoding
What you see
Lookups probe different positions than inserts used: massive false negatives, rejecting valid items.
Fix & prevent
Store hash, seed and encoding in the header, validate on load, refuse mismatches; round-trip tests in CI.
Explain it without notes
Why must the snapshot record a source position?
Practice
Design the manifest and loading procedure for a 1.2 GB filter shared by 200 instances.
Trade-offs
- ↔
Snapshots speed up starts but add a pipeline to maintain; rebuild-on-start is simple but slow and heavy for big datasets.
Done when you can
I can persist, version and restore filters safely and roll out format changes without mixing versions.