Command Palette

Search for a command to run...

Hectal
PHASE 10Advanced ~9 min· topic 5 of 5

Topic 10.5

Capstone: Bloom + Redis + Kafka + Database, and the Anti-Patterns

In one line

The full architecture: clients → rate limiter → application with a Bloom filter → "absent" rejected, "maybe" → Redis → database; database changes → outbox/CDC → Kafka → Bloom updater → versioned filters; snapshots, rebuilds, metrics and failure plans. Plus the list of when not to use a Bloom filter at all.

0/5 · 0%

Think of it like this

A well-run venue entrance: a queue manager limits the crowd (rate limiter), a door list rules out uninvited guests (filter), a quick desk checks regulars (Redis), the office holds the official register (database), and every new registration is sent to the door within a minute (Kafka updater).

Key ideas

  1. 01

    Build the library: BloomFilter<T>, CountingBloomFilter<T>, ScalableBloomFilter<T> with add, mightContain, remove (where supported), serialize, deserialize, metrics and configuration.

  2. 02

    Redis integration: Spring Boot with POST /items (DB insert + outbox), GET /items/{id} (filter → Redis → DB), DELETE /items/{id} (DB delete; filter left stale or cuckoo delete).

  3. 03

    Kafka updates: event schema (ItemCreated/ItemDeleted with IDs and versions), consumer group, retries, DLT, idempotency, watermark, metrics.

  4. 04

    Dashboard: filter size, inserts, queries, positives, estimated and observed FPR, fill ratio, version, last rebuild and duration, Kafka lag, Redis latency, database fallback rate, sampled false negatives.

  5. 05

    Don't use a Bloom filter when exact membership is required without a verifying store, the dataset is small, memory isn't a concern, frequent deletes are mandatory and no variant fits, false positives are unacceptable, the filter can't be maintained safely, the authoritative lookup is already cheap, or you need values, counts or cardinality.

  6. 06

    Common mistakes: treating "maybe" as "yes", using the filter as the source of truth, ignoring capacity and saturation, no rebuild plan, poor hashes, forgetting deletion limits, stale filters, unreliable updates, inappropriate FPR, exposing sensitive membership.

Code & diagrams

capstone.mermaiddiagram
Rendering diagram…
reasoning-chain.txttext
Requirement -> exact or approximate -> false positives tolerable? -> false negatives tolerable?
-> dataset size -> growth -> target FPR -> memory (m, k) -> hash strategy -> deletion needs
-> variant (standard / counting / cuckoo / scalable / static) -> source of truth
-> update strategy -> rebuild strategy -> distribution (local / shared / per-shard / regions)
-> failure behaviour (fail open / closed) -> observability -> security -> trade-offs

Interview problem

The problem

Design the complete cache penetration protection system

Design cache penetration protection end to end: client → rate limiter → Bloom filter → Redis → database, with Kafka/CDC updates and rebuilds. Explain every layer and the system's behaviour when each component fails.

You're given

  • 100M product IDs
  • 90% of attack traffic invalid
  • False negatives unacceptable

Explain it without notes

01

List five situations where you should not use a Bloom filter.

Practice

01

Build the Spring Boot + Redis Bloom + PostgreSQL + Kafka project with the three APIs and the dashboard, then run failure scenarios A, E, F and H.

Trade-offs

  • ↔

    Each layer adds protection and a moving part; the design is only complete when every failure has a planned behaviour.

Done when you can

  • I can design and present the full Bloom + Redis + Kafka + database architecture with failure behaviour and anti-patterns.