Topic 15.5
Recommendation Caches and API Gateways
In one line
A recommendation cache precomputes per-user results and serves them from Redis with TTLs, warming and fallbacks; an API gateway uses Redis for rate limits, session and token checks, idempotency and response caching. Both are exercises in key design, hot keys, failure behaviour and cost.
Think of it like this
A personal shopper who prepares a rack of suggestions for each regular customer before they arrive (precomputed recommendations), and a building's security desk that checks badges, limits visitors per company and remembers who already signed in (gateway).
Key ideas
- 01
Recommendation cache: keys
rec:v3:{userId}holding a small list of item IDs (not full items), TTL a few hours, refreshed by the model pipeline; hydrate items from an item cache. Missing entries fall back to popular items per segment (a small, hot key cached locally). - 02
Cost control: store IDs compactly (packed integers or short strings), cache only active users (TTL), and compute on demand for the long tail.
- 03
Gateway: per-request Lua rate limiting, token introspection cache (
auth:token:{hash}with TTL shorter than token expiry), idempotency for POSTs, and response caching for public GETs; all with tight timeouts and fail-open or fail-closed choices per feature. - 04
Multi-region: recommendations computed centrally and replicated to regional caches; gateways use regional Redis only.
Code & diagrams
Interview problem
The problem
Netflix-like recommendation cache for 100M users
Design the recommendation cache: 100M users, rows of recommendations per user on the home page, recomputed daily, 50K home page loads/sec. Cover key structure, TTL, hot keys, cache warming, multi-region, failure behaviour and cost.
Explain it without notes
Why store item IDs rather than full item objects in per-user recommendation keys?
Practice
List the Redis keys a gateway would touch for one authenticated POST request with an idempotency key.
Trade-offs
- ↔
Precomputation gives fast, predictable reads at the cost of batch compute and staleness until the next refresh.
Done when you can
I can design precomputed recommendation caches with TTLs, fallbacks and warming.
I can list the Redis roles in an API gateway and their failure modes.