Command Palette

Search for a command to run...

Hectal
PHASE 12Advanced ~10 min· topic 2 of 6

Topic 12.2

The Lock Expiry Race and Fencing Tokens

In one line

A lock holder can pause (GC, VM stall, network delay) past its TTL, and then act while another holder also acts. No TTL fixes that. Fencing tokens do: each acquisition gets a monotonically increasing number, and the protected resource rejects writes carrying an older token than it has already seen.

0/6 · 0%

Think of it like this

Numbered tickets at a counter. Your ticket was 41, but you stepped out and the clerk moved on to 42. When you come back waving 41, the clerk refuses, because 42 has already been served. The clerk (the resource) enforces order, not your memory of having been first.

Key ideas

  1. 01

    The race: T0 A gets the lock; T10 A pauses (GC, CPU starvation, swap); T30 the lock expires; T31 B gets the lock and starts writing; T32 A wakes up, still believing it holds the lock, and writes too. Both write: corruption or double processing.

  2. 02

    Why TTL tuning doesn't solve it: pauses are unbounded in practice (long GC, VM live migration, a stuck disk), and A can't reliably check "do I still hold it?" right before writing, because a pause can happen between the check and the write.

  3. 03

    Fencing token: on acquire, also get a number that only increases: INCR lock:resource:fence. Every write to the protected resource includes the token. The resource stores the highest token it has seen and rejects lower ones: in SQL, UPDATE ... SET data = ?, fence = ? WHERE id = ? AND fence < ?.

  4. 04

    The key point: correctness is enforced by the resource (the database, the storage service), not by Redis. Redis provides the numbers; the resource checks them.

  5. 05

    If the resource can't check tokens (a third-party API), use its idempotency features instead, or accept that the lock is only an efficiency optimisation.

  6. 06

    Consensus stores provide fencing naturally: etcd revisions and ZooKeeper zxids are monotonic and survive failover. A Redis INCR counter can go backwards after a failover that loses recent increments, so persist the highest seen token at the resource and treat Redis tokens as best effort unless durability is guaranteed.

Code & diagrams

fencing.mermaiddiagram
Rendering diagram…
acquire-with-fence.lualua
-- KEYS[1] = lock key, KEYS[2] = fence counter (same hash tag)
-- ARGV[1] = owner token, ARGV[2] = ttl ms
if redis.call('SET', KEYS[1], ARGV[1], 'NX', 'PX', ARGV[2]) then
  return redis.call('INCR', KEYS[2])     -- fencing token
end
return 0
fenced-write.sqlsql
-- The storage enforces order: stale holders can't overwrite.
UPDATE inventory_snapshot
SET    payload = $1, fence = $2
WHERE  sku = $3 AND fence < $2;
-- 0 rows updated => our token is stale: abort, don't retry blindly.

Interview problem

The problem

Two services believe they hold the same lock

A payment reconciliation job uses a Redis lock. An incident shows two instances processed the same batch concurrently. Explain how that happened (several ways) and how to make the lock safer.

The interviewer follows up

01

How do you make a distributed lock safer in one sentence?

When it breaks

Relying on lock TTL alone for a money-moving job

What you see

A long GC pause causes two instances to move money for the same batch; duplicates reach the payment provider.

Fix & prevent

Idempotency keys at the provider plus fencing or unique constraints in the database; the lock becomes an optimisation only.

Explain it without notes

01

Walk through the lock expiry race and explain why fencing fixes it.

Practice

01

Simulate the race: worker A acquires with token 1 and sleeps past the TTL, B acquires with token 2 and writes; then A writes. Show the fenced update rejects A.

Trade-offs

  • ↔

    Fencing requires changes in the protected resource, but it's the only way to get correctness from a lease-based lock.

Done when you can

  • I can explain the pause-past-TTL race and failover lock loss.

  • I can implement fencing tokens and enforce them in storage.