Command Palette

Search for a command to run...

Hectal
PHASE 6Intermediate ~9 min· topic 7 of 7

Topic 6.7

Multi-Level Caching and Hit Ratio

In one line

Layering an in-process L1 cache in front of Redis (L2) in front of the database (L3) cuts latency and removes hot-key load, at the cost of more places for data to go stale. Measure hit ratio per layer, but optimise for user latency and database load, not the ratio itself.

0/7 · 0%

Think of it like this

A chef keeps salt at the stove (L1), a bigger supply in the kitchen cupboard (L2), and a bulk bag in the storeroom (L3). Most reaches are to the stove; the cupboard is refilled from the storeroom; the stove is refilled from the cupboard. Moving the salt to a new recipe means updating all three.

Key ideas

  1. 01

    L1 (in-process, for example Caffeine, a Guava cache, or a Node LRU): nanosecond reads, no network, no serialisation, but per-instance memory, duplicated across instances, and invalidation is harder. Keep L1 small and short-lived (seconds) unless you have invalidation messages.

  2. 02

    L2 (Redis): shared by all instances, sub-millisecond, larger capacity, one place to invalidate. L3: the database or origin service. A CDN in front of it all is L0 for public content.

  3. 03

    Invalidation across layers: delete in Redis and broadcast an invalidation (Pub/Sub channel or CLIENT TRACKING) so every instance drops its L1 entry. Always keep TTLs on L1 as a backstop for missed messages.

  4. 04

    Hit ratio = hits / (hits + misses), per layer and per keyspace. Useful for spotting problems (sudden drops, evictions, bad keys), but don't chase 100%: the goal is latency and cost. A 99% hit ratio with stale prices is worse than 95% with correct ones.

  5. 05

    Questions to ask before tuning: is stale data acceptable here, and for how long? What does a miss cost (latency, database load, money)? What does the memory cost? Does the cache hold the right data (hot set) or just a lot of data?

Code & diagrams

TwoLevelCache.javajava
public class TwoLevelCache {
    private final Cache<String, String> l1 = Caffeine.newBuilder()
        .maximumSize(10_000)
        .expireAfterWrite(Duration.ofSeconds(5))       // short: bounds staleness
        .build();
    private final StringRedisTemplate redis;

    public String get(String key, Supplier<String> loader) {
        return l1.get(key, k -> {                       // coalesces concurrent L1 misses
            String v = redis.opsForValue().get(k);      // L2
            if (v != null) return v;
            v = loader.get();                           // L3
            redis.opsForValue().set(k, v, Duration.ofMinutes(10));
            return v;
        });
    }

    // Invalidation: delete in Redis, then tell every instance to drop L1.
    public void invalidate(String key) {
        redis.delete(key);
        redis.convertAndSend("cache-invalidate", key);
    }

    // Registered as a MessageListener on channel "cache-invalidate"
    public void onInvalidate(String key) { l1.invalidate(key); }
}
levels.mermaiddiagram
Rendering diagram…

Interview problem

The problem

Recommendation widget: which layers?

A recommendation widget on every page reads a per-user list (changes hourly) and a global "trending" list (changes every minute). Traffic is 300K requests/sec across 400 instances. Design the caching layers and explain hit ratio targets and staleness.

The interviewer follows up

01

When is adding an L1 cache a bad idea?

When it breaks

L1 cache with a long TTL and no invalidation

What you see

After an update, instances serve different versions for minutes; users see values flip back and forth between requests.

Fix & prevent

Short L1 TTLs (seconds), invalidation broadcasts, and flushing L1 on Redis reconnects.

Explain it without notes

01

Why shouldn't you optimise purely for hit ratio?

Practice

01

Add a Caffeine (or equivalent) L1 in front of an existing Redis cache for one hot key and measure Redis ops/sec for that key before and after.

Trade-offs

  • ↔

    L1 gives the lowest latency and removes hot-key load, but multiplies staleness sources and memory across instances.

  • ↔

    Longer TTLs raise hit ratio and staleness together.

Done when you can

  • I can design L1/L2/L3 caching with appropriate TTLs and invalidation.

  • I can compute and interpret hit ratio per layer without over-optimising it.