Topic 6.7
Multi-Level Caching and Hit Ratio
In one line
Layering an in-process L1 cache in front of Redis (L2) in front of the database (L3) cuts latency and removes hot-key load, at the cost of more places for data to go stale. Measure hit ratio per layer, but optimise for user latency and database load, not the ratio itself.
Think of it like this
A chef keeps salt at the stove (L1), a bigger supply in the kitchen cupboard (L2), and a bulk bag in the storeroom (L3). Most reaches are to the stove; the cupboard is refilled from the storeroom; the stove is refilled from the cupboard. Moving the salt to a new recipe means updating all three.
Key ideas
- 01
L1 (in-process, for example Caffeine, a Guava cache, or a Node LRU): nanosecond reads, no network, no serialisation, but per-instance memory, duplicated across instances, and invalidation is harder. Keep L1 small and short-lived (seconds) unless you have invalidation messages.
- 02
L2 (Redis): shared by all instances, sub-millisecond, larger capacity, one place to invalidate. L3: the database or origin service. A CDN in front of it all is L0 for public content.
- 03
Invalidation across layers: delete in Redis and broadcast an invalidation (Pub/Sub channel or
CLIENT TRACKING) so every instance drops its L1 entry. Always keep TTLs on L1 as a backstop for missed messages. - 04
Hit ratio = hits / (hits + misses), per layer and per keyspace. Useful for spotting problems (sudden drops, evictions, bad keys), but don't chase 100%: the goal is latency and cost. A 99% hit ratio with stale prices is worse than 95% with correct ones.
- 05
Questions to ask before tuning: is stale data acceptable here, and for how long? What does a miss cost (latency, database load, money)? What does the memory cost? Does the cache hold the right data (hot set) or just a lot of data?
Code & diagrams
public class TwoLevelCache {
private final Cache<String, String> l1 = Caffeine.newBuilder()
.maximumSize(10_000)
.expireAfterWrite(Duration.ofSeconds(5)) // short: bounds staleness
.build();
private final StringRedisTemplate redis;
public String get(String key, Supplier<String> loader) {
return l1.get(key, k -> { // coalesces concurrent L1 misses
String v = redis.opsForValue().get(k); // L2
if (v != null) return v;
v = loader.get(); // L3
redis.opsForValue().set(k, v, Duration.ofMinutes(10));
return v;
});
}
// Invalidation: delete in Redis, then tell every instance to drop L1.
public void invalidate(String key) {
redis.delete(key);
redis.convertAndSend("cache-invalidate", key);
}
// Registered as a MessageListener on channel "cache-invalidate"
public void onInvalidate(String key) { l1.invalidate(key); }
}Interview problem
The problem
Recommendation widget: which layers?
A recommendation widget on every page reads a per-user list (changes hourly) and a global "trending" list (changes every minute). Traffic is 300K requests/sec across 400 instances. Design the caching layers and explain hit ratio targets and staleness.
The interviewer follows up
When is adding an L1 cache a bad idea?
When it breaks
L1 cache with a long TTL and no invalidation
What you see
After an update, instances serve different versions for minutes; users see values flip back and forth between requests.
Fix & prevent
Short L1 TTLs (seconds), invalidation broadcasts, and flushing L1 on Redis reconnects.
Explain it without notes
Why shouldn't you optimise purely for hit ratio?
Practice
Add a Caffeine (or equivalent) L1 in front of an existing Redis cache for one hot key and measure Redis ops/sec for that key before and after.
Trade-offs
- ↔
L1 gives the lowest latency and removes hot-key load, but multiplies staleness sources and memory across instances.
- ↔
Longer TTLs raise hit ratio and staleness together.
Done when you can
I can design L1/L2/L3 caching with appropriate TTLs and invalidation.
I can compute and interpret hit ratio per layer without over-optimising it.