Command Palette

Search for a command to run...

Hectal
PHASE 10Advanced ~10 min· topic 3 of 4

Topic 10.3

Redis Sentinel: Automatic Failover

In one line

Sentinel processes monitor a primary and its replicas, agree that the primary is down (SDOWN then ODOWN by quorum), elect a leader among themselves by majority, promote the best replica, reconfigure the others, and tell clients the new primary's address.

0/4 · 0%

Think of it like this

Three lifeguards watching a pool's head coach. If one lifeguard thinks the coach has fainted, that's a suspicion. If enough of them agree, they declare an emergency, vote on which lifeguard takes charge, and that lifeguard appoints the best assistant coach as the new head, then tells all the swimmers who to listen to now.

Key ideas

  1. 01

    Deploy at least 3 Sentinels on separate failure domains (hosts or AZs), configured with sentinel monitor mymaster 10.0.0.1 6379 2 (quorum 2). Sentinels discover replicas and other Sentinels automatically through the primary and a hello channel.

  2. 02

    SDOWN (subjectively down): one Sentinel hasn't had a valid reply from the primary for down-after-milliseconds (for example 5,000). ODOWN (objectively down): at least quorum Sentinels report SDOWN.

  3. 03

    Leader election: to actually fail over, a Sentinel must be elected by a majority of all Sentinels (not just the quorum). With 3 Sentinels, 2 must be reachable to fail over; with 2 Sentinels, losing one host prevents failover, which is why 2 isn't enough.

  4. 04

    Promotion: the leader picks a replica by replica-priority (0 means never promote), then the largest replication offset (most data), then the lowest run ID; sends it REPLICAOF NO ONE; points the other replicas at it; and publishes +switch-master.

  5. 05

    Clients: connect to Sentinels, ask SENTINEL get-master-addr-by-name mymaster, connect to that address, and re-ask on errors or +switch-master events. Sentinel-aware clients (Lettuce, Jedis JedisSentinelPool, redis-py Sentinel) do this for you.

  6. 06

    Timing: detection (down-after-milliseconds) + election + promotion is typically a few seconds to tens of seconds. Writes fail during that window, and writes acknowledged by the old primary but not replicated are lost.

  7. 07

    Sentinel is for one primary's dataset (no sharding). For sharding with built-in failover, use Redis Cluster (Phase 11).

Code & diagrams

sentinel.conftext
port 26379
sentinel monitor mymaster 10.0.0.1 6379 2
sentinel auth-user mymaster sentinel-user
sentinel auth-pass mymaster <secret>
sentinel down-after-milliseconds mymaster 5000
sentinel failover-timeout mymaster 60000
sentinel parallel-syncs mymaster 1
sentinel resolve-hostnames yes
failover.mermaiddiagram
Rendering diagram…
sentinel-client.pypython
from redis.sentinel import Sentinel

sentinel = Sentinel([("s1", 26379), ("s2", 26379), ("s3", 26379)],
                    socket_timeout=0.2, sentinel_kwargs={"password": "..."})
primary = sentinel.master_for("mymaster", socket_timeout=0.2, password="...")
replica = sentinel.slave_for("mymaster", socket_timeout=0.2, password="...")

primary.set("k", "v")          # re-resolved automatically after a failover
print(sentinel.discover_master("mymaster"))   # ('10.0.0.2', 6379)

Interview problem

The problem

Primary failure with three Sentinels

You run a primary, two replicas and three Sentinels. The primary's host dies. Explain step by step: failure, detection, election, promotion, client redirection. Then: what if only two Sentinels were deployed, and they share a host with the primary?

You're given

  • 1 primary, 2 replicas
  • 3 Sentinels, quorum 2
  • down-after 5 s

The interviewer follows up

01

What's the difference between quorum and majority in Sentinel?

When it breaks

Clients hard-code the primary's IP instead of asking Sentinel

What you see

After a correct failover, clients keep writing to the old address: errors, or worse, writes to a returned old primary that gets demoted and loses them.

Fix & prevent

Use Sentinel-aware clients (or a stable endpoint maintained by a proxy) and test failover in staging.

Sentinels all in one availability zone

What you see

An AZ outage takes out the majority of Sentinels; no failover can happen for Redis nodes in other zones.

Fix & prevent

Spread Sentinels across three zones.

Explain it without notes

01

Explain SDOWN, ODOWN and leader election.

02

How does Sentinel pick which replica to promote?

Practice

01

Run a primary, two replicas and three Sentinels in Docker Compose; stop the primary container and time the failover from Sentinel logs.

Trade-offs

  • ↔

    Shorter down-after-milliseconds means faster failover and more false failovers during brief network glitches or GC-like pauses.

Done when you can

  • I can walk through a Sentinel failover step by step.

  • I deploy 3+ Sentinels across failure domains and use Sentinel-aware clients.