Topic 10.3
Redis Sentinel: Automatic Failover
In one line
Sentinel processes monitor a primary and its replicas, agree that the primary is down (SDOWN then ODOWN by quorum), elect a leader among themselves by majority, promote the best replica, reconfigure the others, and tell clients the new primary's address.
Think of it like this
Three lifeguards watching a pool's head coach. If one lifeguard thinks the coach has fainted, that's a suspicion. If enough of them agree, they declare an emergency, vote on which lifeguard takes charge, and that lifeguard appoints the best assistant coach as the new head, then tells all the swimmers who to listen to now.
Key ideas
- 01
Deploy at least 3 Sentinels on separate failure domains (hosts or AZs), configured with
sentinel monitor mymaster 10.0.0.1 6379 2(quorum 2). Sentinels discover replicas and other Sentinels automatically through the primary and a hello channel. - 02
SDOWN (subjectively down): one Sentinel hasn't had a valid reply from the primary for
down-after-milliseconds(for example 5,000). ODOWN (objectively down): at leastquorumSentinels report SDOWN. - 03
Leader election: to actually fail over, a Sentinel must be elected by a majority of all Sentinels (not just the quorum). With 3 Sentinels, 2 must be reachable to fail over; with 2 Sentinels, losing one host prevents failover, which is why 2 isn't enough.
- 04
Promotion: the leader picks a replica by
replica-priority(0 means never promote), then the largest replication offset (most data), then the lowest run ID; sends itREPLICAOF NO ONE; points the other replicas at it; and publishes+switch-master. - 05
Clients: connect to Sentinels, ask
SENTINEL get-master-addr-by-name mymaster, connect to that address, and re-ask on errors or+switch-masterevents. Sentinel-aware clients (Lettuce, JedisJedisSentinelPool, redis-pySentinel) do this for you. - 06
Timing: detection (
down-after-milliseconds) + election + promotion is typically a few seconds to tens of seconds. Writes fail during that window, and writes acknowledged by the old primary but not replicated are lost. - 07
Sentinel is for one primary's dataset (no sharding). For sharding with built-in failover, use Redis Cluster (Phase 11).
Code & diagrams
port 26379
sentinel monitor mymaster 10.0.0.1 6379 2
sentinel auth-user mymaster sentinel-user
sentinel auth-pass mymaster <secret>
sentinel down-after-milliseconds mymaster 5000
sentinel failover-timeout mymaster 60000
sentinel parallel-syncs mymaster 1
sentinel resolve-hostnames yesfrom redis.sentinel import Sentinel
sentinel = Sentinel([("s1", 26379), ("s2", 26379), ("s3", 26379)],
socket_timeout=0.2, sentinel_kwargs={"password": "..."})
primary = sentinel.master_for("mymaster", socket_timeout=0.2, password="...")
replica = sentinel.slave_for("mymaster", socket_timeout=0.2, password="...")
primary.set("k", "v") # re-resolved automatically after a failover
print(sentinel.discover_master("mymaster")) # ('10.0.0.2', 6379)Interview problem
The problem
Primary failure with three Sentinels
You run a primary, two replicas and three Sentinels. The primary's host dies. Explain step by step: failure, detection, election, promotion, client redirection. Then: what if only two Sentinels were deployed, and they share a host with the primary?
You're given
- 1 primary, 2 replicas
- 3 Sentinels, quorum 2
- down-after 5 s
The interviewer follows up
What's the difference between quorum and majority in Sentinel?
When it breaks
Clients hard-code the primary's IP instead of asking Sentinel
What you see
After a correct failover, clients keep writing to the old address: errors, or worse, writes to a returned old primary that gets demoted and loses them.
Fix & prevent
Use Sentinel-aware clients (or a stable endpoint maintained by a proxy) and test failover in staging.
Sentinels all in one availability zone
What you see
An AZ outage takes out the majority of Sentinels; no failover can happen for Redis nodes in other zones.
Fix & prevent
Spread Sentinels across three zones.
Explain it without notes
Explain SDOWN, ODOWN and leader election.
How does Sentinel pick which replica to promote?
Practice
Run a primary, two replicas and three Sentinels in Docker Compose; stop the primary container and time the failover from Sentinel logs.
Trade-offs
- ↔
Shorter
down-after-millisecondsmeans faster failover and more false failovers during brief network glitches or GC-like pauses.
Done when you can
I can walk through a Sentinel failover step by step.
I deploy 3+ Sentinels across failure domains and use Sentinel-aware clients.