Topic 6.3
Cache Invalidation and Consistency Races
In one line
Stale values reappear when a slow reader writes an old database value into the cache after a writer already invalidated it. The fixes, in increasing strength: delete after commit, short TTLs, delayed double delete, versioned or compare-and-set cache writes, and change-data-capture invalidation.
Think of it like this
A whiteboard showing today's price. A clerk reads the old price from the ledger, walks slowly to the whiteboard, and meanwhile the manager updates the ledger and wipes the board. Then the slow clerk arrives and writes the old price back up. Everyone sees the wrong price until someone notices.
Key ideas
- 01
The three orderings: (1) delete cache → update database: a reader between the two steps re-caches the old value, which is bad. (2) update database → update cache: concurrent writers can apply cache updates out of order. (3) update database → delete cache: the best default, but still has a rare race.
- 02
The remaining race with (3): the cache is empty; reader R misses and reads the old row (v1) from the database; writer W updates the row to v2 and deletes the cache key; R, which was slow, now
SETs v1 into the cache. Stale until TTL. It needs a miss plus a slow reader overlapping a write, so it's rare, but real under load. - 03
Delayed double delete: after the write, delete the key again after a short delay (longer than a typical read-then-set, for example 500 ms–1 s) to evict any stale value a racing reader put back. Cheap and effective, not airtight.
- 04
Versioned writes: store a version with the value (row version or updated-at) and only let the cache accept newer versions, using a small Lua compare-and-set:
if cached_version < new_version then SET. Readers that loaded an old row can't overwrite a newer one. - 05
Lease or placeholder: on a miss, the reader first sets a short-lived placeholder (
SET key __loading__ NX PX 2000) and only writes the loaded value if the placeholder is still its own; a writer'sDELremoves the placeholder, so the stale reader's write is rejected. Facebook's memcache "leases" paper describes this pattern. - 06
Change data capture: stream database changes (Debezium reading the WAL/binlog → Kafka) to an invalidator that deletes keys. It catches every write path, including batch jobs and manual SQL that application-level invalidation misses, and retries on failure. Combine it with TTLs.
Code & diagrams
Cache writes carry the row version; older versions can never overwrite newer ones.
-- KEYS[1] = cache key, ARGV[1] = version, ARGV[2] = payload, ARGV[3] = ttl seconds
local current = redis.call('HGET', KEYS[1], 'v')
if current and tonumber(current) >= tonumber(ARGV[1]) then
return 0 -- cache already has same or newer
end
redis.call('HSET', KEYS[1], 'v', ARGV[1], 'data', ARGV[2])
redis.call('EXPIRE', KEYS[1], ARGV[3])
return 1Interview problem
The problem
Price updated in the database, Redis still shows the old one
A product's price is updated from 400 to 500 in the database, but Redis still returns 400. Compare "delete cache → update DB", "update DB → delete cache" and "update DB → update cache", walk through the races, and design a safe strategy.
You're given
- Price must be correct at checkout
- Listing pages may be stale for up to 1 minute
- Several services and a nightly batch job write prices
The interviewer follows up
Can you ever get strong consistency between Redis and the database?
When it breaks
Invalidation code added only to the API, not to admin tools or batch jobs
What you see
Prices changed by the nightly job stay stale in the cache until TTL, sometimes for hours if TTLs are long.
Fix & prevent
CDC-based invalidation catches all writers; keep TTLs as a hard upper bound.
Explain it without notes
Explain the stale-read race that remains with "update DB, then delete cache".
How do versioned cache writes prevent stale overwrites?
Practice
Reproduce the stale race in a test: add a sleep between the reader's DB read and cache set, run a concurrent update, and observe the stale value. Then apply delayed double delete.
Trade-offs
- ↔
Stronger invalidation (CDC, versioning) costs infrastructure and complexity; TTLs alone are simple but allow bounded staleness.
- ↔
Deleting instead of updating costs an extra miss after writes but avoids ordering bugs.
Done when you can
I can compare the three write orderings and describe each race.
I can apply delayed double delete, versioned writes, leases and CDC.
I never rely on a cache for money-critical reads.