Topic 12.1
Caching in Front of the Database
In one line
Caches (in-process and distributed) cut latency and database load for hot reads. Cache-aside is the default: read the cache, load from the database on a miss, and delete the key after writes commit. Correctness problems come from invalidation races, stampedes on hot keys, penetration by nonexistent keys and synchronized expiry. Each has a standard defence.
Think of it like this
A chef's mise en place. Frequently used ingredients sit chopped on the counter (cache); the pantry (database) is the source. If the recipe changes, the chopped onions on the counter must be thrown out, or dishes come out wrong.
Key ideas
- 01
Patterns: cache-aside (application manages the cache; most common), read-through (the cache library loads on a miss), write-through (write the cache and database together), write-behind (write the cache, flush to the database later; risk of loss), refresh-ahead (refresh popular keys before expiry).
- 02
Invalidation: after a write commits, delete (don't update) the cache key. Updating risks racing writers leaving an older value; deleting lets the next read load the latest. Races remain (a reader loads a stale value just before the delete and writes it back after), so keep TTLs as a safety net or invalidate via CDC with versioned values.
- 03
Local (in-process, e.g. Caffeine) vs distributed (Redis, Memcached): local is nanoseconds with no network but has N copies to invalidate (use short TTLs or pub/sub invalidation); distributed is shared and consistent-ish. A two-level cache combines them for very hot keys.
- 04
Stampede: a hot key expires and thousands of requests hit the database. Fix with request coalescing (single-flight), a short lock per key, stale-while-revalidate, or probabilistic early refresh. Avalanche: many keys expire together, so add TTL jitter. Penetration: requests for nonexistent keys, so use negative caching and Bloom filters (see the Bloom Filters course).
- 05
Eviction: LRU for recency, LFU for frequency (Redis
allkeys-lfu); measure hit ratio per key prefix. A cache that hides a slow query can fail catastrophically when it goes cold, so the database must survive a cold start (warmup, rate limits).
Code & diagrams
public Product get(long id) {
String key = "product:v3:" + id;
String cached = redis.get(key);
if (cached != null) return "NONE".equals(cached) ? null : json.read(cached);
// single-flight: one loader per key per instance
return loaders.computeIfAbsent(key, k -> CompletableFuture.supplyAsync(() -> {
Product p = repo.findById(id).orElse(null);
Duration ttl = Duration.ofMinutes(10).plusSeconds(ThreadLocalRandom.current().nextInt(120)); // jitter
redis.set(key, p == null ? "NONE" : json.write(p), p == null ? Duration.ofSeconds(60) : ttl);
return p;
})).whenComplete((r, e) -> loaders.remove(key)).join();
}
@Transactional
public void update(Product p) {
repo.save(p);
TransactionSynchronizationManager.registerSynchronization(new TransactionSynchronization() {
@Override public void afterCommit() { redis.delete("product:v3:" + p.getId()); } // delete AFTER commit
});
}Interview problem
The problem
Cache the product page for a flash sale
A product page gets 200K reads/sec during a sale; price and stock change every few seconds. The database handles 10K queries/sec. Design caching that keeps the database safe and shows reasonably fresh stock.
When it breaks
Updating the cache before the database transaction commits
What you see
The transaction rolls back but the cache holds the uncommitted value; users see data that never existed until the TTL expires.
Fix & prevent
Invalidate after commit (transaction synchronization or CDC), and prefer delete over set.
Cache cluster restart during peak
What you see
Hit ratio drops from 99% to 0; the database receives 100× its normal load and falls over, and the outage outlasts the cache restart.
Fix & prevent
Warm critical keys before taking traffic, coalesce misses, rate-limit database access and shed load, and use replicated cache nodes.
Explain it without notes
Why delete the cache key after a write instead of updating it?
Practice
Compute database load from a 99% hit ratio at 50K reads/sec, and at 95%.
Trade-offs
- ↔
Caches buy latency and capacity with staleness, invalidation complexity, and dangerous cold-start behaviour.
Done when you can
I can design cache-aside with safe invalidation, stampede protection and TTL jitter.