Topic 11.5
Operating Clusters and Choosing a Topology
In one line
Choose between standalone with replicas, Sentinel, Cluster, client-side sharding or a proxy based on data size, throughput, multi-key needs and client support. Then operate it: balanced slots, zone-aware placement, per-node monitoring, safe upgrades and capacity headroom.
Think of it like this
Choosing how to run a restaurant as it grows: one kitchen with a backup cook (primary + replica with Sentinel), several kitchens each making different dishes (Cluster), or a central order desk that routes orders to kitchens (a proxy). Bigger setups serve more but need more coordination.
Key ideas
- 01
Standalone + replica + Sentinel: data fits one node (tens of GB), throughput fits one core, you want full multi-key freedom and simple clients. Most applications live here happily.
- 02
Cluster: data or throughput exceed one node, you can live with per-slot multi-key operations, and clients are cluster-aware. Scales out by adding shards.
- 03
Proxies (Envoy Redis proxy, Twemproxy, managed service endpoints): one address for simple clients, sharding hidden behind it, sometimes with command restrictions. Client-side sharding (consistent hashing in the app) is the old way; it lacks automatic failover and resharding.
- 04
Day-2 operations:
redis-cli --cluster checkin monitoring, alerts oncluster_state, per-node memory, ops and latency (skew reveals hot slots), replica count per primary, and zone placement after failovers. - 05
Upgrades: roll replicas first, then fail over each primary to its upgraded replica with
CLUSTER FAILOVER, then upgrade the old primary. No downtime and no data loss for planned maintenance. - 06
Capacity: add shards before nodes pass ~70–80% memory; rebalancing needs headroom. Keep shards similar in size so one hot or full node doesn't limit the cluster.
Code & diagrams
# For each shard:
# 1. Upgrade the replica (restart with new version), wait for sync
redis-cli -h replica-A1 INFO replication | grep master_link_status # up
# 2. Swap roles safely
redis-cli -h replica-A1 CLUSTER FAILOVER
# 3. Upgrade the old primary (now a replica), wait for sync
# 4. Optionally fail back to restore zone placement
redis-cli --cluster check primary-B:6379Interview problem
The problem
Sentinel or Cluster?
Two teams ask you to pick an architecture. Team 1: 12 GB of sessions and caches, 80K ops/sec, heavy use of Lua across arbitrary keys. Team 2: 400 GB of feature data for ML serving, 1.5M ops/sec, simple key lookups. Recommend and justify.
The interviewer follows up
When would you choose a managed service instead?
Explain it without notes
Explain how to upgrade a Redis Cluster with no downtime.
Practice
Write a topology decision for a system you know, including shard count, replicas, zones and persistence.
Trade-offs
- ↔
Cluster buys scale with operational and application complexity; Sentinel keeps things simple until one node is no longer enough.
Done when you can
I can choose between standalone, Sentinel, Cluster and proxies.
I can run rolling upgrades and monitor cluster health.