Topic 11.3
MOVED, ASK, and Resharding
In one line
MOVED tells a client a slot now lives on another node permanently, so update the slot map. ASK tells it that during a migration this one key has already moved, so try once on the target with ASKING. Resharding moves slots key by key while the cluster keeps serving.
Think of it like this
A company moving departments between floors. "Accounts has moved to floor 5" is permanent, so you update your directory (MOVED). "Accounts is moving today; this particular file has already been carried up, ask on floor 5 just for this one" is temporary (ASK); tomorrow the directory will say floor 5 anyway.
Key ideas
- 01
MOVED 9189 10.0.0.2:6379: the slot is owned by another node. Clients retry there and refresh their slot map (often the whole map). Frequent MOVED errors mean clients have stale maps, for example after a failover or rebalance. - 02
Slot migration steps: the target is marked
IMPORTINGfor the slot and the sourceMIGRATING; keys are moved in batches withMIGRATE(atomic per batch: copy to target, delete from source); finally both nodes and the rest of the cluster are told the new owner withCLUSTER SETSLOT <slot> NODE <target>. - 03
ASK 9189 10.0.0.4:6379: during migration, if the source no longer has the requested key, it redirects the client to the target for this command only. The client sendsASKINGthen the command to the target, but doesn't update its slot map, since the slot officially still belongs to the source until migration ends. - 04
Multi-key commands on a migrating slot can fail with
TRYAGAINwhen some keys are on the source and others on the target; clients retry after a short delay. - 05
Tooling:
redis-cli --cluster reshard,--cluster rebalance(moves slots to even out ownership, optionally weighted),--cluster add-nodeanddel-node,--cluster checkand--cluster fix. Big keys make migrations slow and can block (eachMIGRATEof a huge key serialises it on the main thread), another reason to avoid them. Newer Valkey versions introduce atomic slot migration to simplify this process.
Code & diagrams
# Add an empty primary, then move 2000 slots to it
redis-cli --cluster add-node 10.0.0.7:6379 10.0.0.1:6379
redis-cli --cluster reshard 10.0.0.1:6379 \
--cluster-to <new-node-id> --cluster-slots 2000 \
--cluster-from all --cluster-yes --cluster-pipeline 100
# Or even out ownership automatically
redis-cli --cluster rebalance 10.0.0.1:6379 --cluster-use-empty-masters
redis-cli --cluster check 10.0.0.1:6379
[OK] All nodes agree about slots configuration.
[OK] All 16384 slots covered.Interview problem
The problem
Resharding a live cluster
Your 3-primary cluster (A, B, C) is at 80% memory. You add node D and move slots to it during business hours. Explain migration, availability, client behaviour, ASK vs MOVED, and the risks.
You're given
- ~15 GB per node
- 200K ops/sec total
- Some keys up to 50 MB
The interviewer follows up
What's the difference between MOVED and ASK in one sentence?
When it breaks
Migration of a slot containing a 2 GB key
What you see
MIGRATE blocks the source for seconds; clients time out and the cluster may even mark the node as failing.
Fix & prevent
Find and split big keys before resharding (--bigkeys); increase timeouts for the migration tool; migrate off-peak.
Explain it without notes
Explain what happens to reads and writes for a slot while it's being migrated.
Practice
Add a fourth primary to your lab cluster and reshard 1,000 slots while a load script runs; count ASK and MOVED redirects in client logs.
Trade-offs
- ↔
Online resharding avoids downtime but adds load and latency while it runs.
Done when you can
I can explain MOVED vs ASK and the migration states.
I can reshard and rebalance a cluster safely.