Command Palette

Search for a command to run...

Hectal
PHASE 12Advanced ~9 min· topic 3 of 6

Topic 12.3

Redlock: Majority Locking and Its Critiques

In one line

Redlock acquires the same lock on a majority of N independent Redis masters within a time limit, to survive the failure of individual nodes. It improves availability over a single instance but relies on timing assumptions, and critics (notably Martin Kleppmann) argue it isn't safe for correctness without fencing.

0/6 · 0%

Think of it like this

Reserving a table by calling five restaurants' shared booking desks and only considering yourself booked if at least three confirm quickly. It tolerates one or two desks being down, but if your phone calls take so long that the early confirmations have already expired, you might believe you have a booking that no longer exists.

Key ideas

  1. 01

    Algorithm (N = 5 independent masters, no replication between them): record the start time; try SET key token NX PX ttl on each with a short per-node timeout; if at least 3 succeed and the elapsed time is less than the TTL, the lock is held for validity = ttl − elapsed − clock drift margin; otherwise release on all nodes and retry after a random delay.

  2. 02

    Why majority: two clients can't both hold a majority at the same time, so it survives up to 2 of 5 nodes failing, unlike a single Redis where failover can lose the lock.

  3. 03

    The critique (Kleppmann, "How to do distributed locking", 2016): Redlock depends on bounded network delay, bounded process pauses and clocks that don't jump. A GC pause after acquiring, or a node's clock jumping forward (expiring the lock early), can give two clients the lock. And Redlock generates no fencing token, so the resource can't detect the violation.

  4. 04

    The response (antirez, "Is Redlock safe?"): the algorithm's assumptions are reasonable in practice, with monotonic clock use and careful configuration (for example delaying restarts of crashed nodes by at least the TTL, or using fsync-always persistence).

  5. 05

    The practical takeaway to say in interviews: for efficiency (avoid duplicate work), a single-instance lock is usually enough, and Redlock adds cost and complexity. For correctness, neither is sufficient alone: use fencing tokens with a resource that checks them, or a consensus system (etcd, ZooKeeper, Consul) that provides linearizable locks and monotonic revisions.

Code & diagrams

redlock.mermaiddiagram
Rendering diagram…
redisson-redlock.javajava

Redisson's RedissonMultiLock can combine locks from independent instances; it still doesn't give fencing tokens.

RLock l1 = redisson1.getLock("lock:settlement");
RLock l2 = redisson2.getLock("lock:settlement");
RLock l3 = redisson3.getLock("lock:settlement");
RLock multi = new RedissonMultiLock(l1, l2, l3);   // majority-style lock across instances
if (multi.tryLock(2, 30, TimeUnit.SECONDS)) {
    try {
        settle();   // still write with a fencing token or idempotency key
    } finally {
        multi.unlock();
    }
}

Interview problem

The problem

Should we use Redlock?

A team proposes Redlock across 5 Redis nodes to guarantee that a ledger settlement job never runs twice. Evaluate the proposal: assumptions, failure scenarios, and what you'd recommend.

The interviewer follows up

01

When is Redlock a reasonable choice?

Explain it without notes

01

Summarise Kleppmann's critique of Redlock.

Practice

01

Write a one-paragraph decision record choosing between a single Redis lock, Redlock, and etcd for three jobs: cache rebuild, nightly email digest, and payout processing.

Trade-offs

  • ↔

    Redlock raises availability of locking but not its correctness guarantees; consensus systems give stronger guarantees at higher latency and operational cost.

Done when you can

  • I can explain the Redlock algorithm and its assumptions.

  • I can argue both sides of the Redlock debate and give a recommendation.