Command Palette

Search for a command to run...

Hectal
PHASE 14Advanced ~9 min· topic 5 of 5

Topic 14.5

Upgrades, Backups, and Disaster Recovery

In one line

Upgrade Redis by rolling through replicas and failing over, never by restarting a primary in place. Back up RDB files off-host and test restores. Define RPO and RTO per keyspace and rehearse recovering from a lost node, a lost zone, a bad FLUSHALL, and a lost region.

0/5 · 0%

Think of it like this

Changing the tyres on a bus route. You swap buses one at a time with a spare ready, keep a copy of the timetable somewhere else, and practise what happens if a depot burns down.

Key ideas

  1. 01

    Rolling upgrade (Sentinel or Cluster): upgrade replicas first and let them resync; fail over to an upgraded replica (SENTINEL FAILOVER mymaster or CLUSTER FAILOVER on the replica); upgrade the old primary. Check release notes for config changes and incompatible changes, and keep the RDB format in mind: newer versions can read older RDB files, not always the reverse, so rollback may require restoring from a replica still on the old version.

  2. 02

    Redis to Valkey (or version jumps): Valkey 8.x reads Redis ≤7.2 RDB files and can replicate from them, so you can add Valkey replicas, fail over, and retire the old nodes. Going the other way, or from Redis 7.4+/8 formats, may not be supported: plan and test.

  3. 03

    Backups: scheduled RDB snapshots from replicas copied to object storage (encrypted, versioned, cross-region for DR), with retention matching compliance. Managed services offer automated snapshots; still test restores.

  4. 04

    Human error: FLUSHALL or a bad script deletes data and replicates the deletion instantly. Replicas don't protect you; backups do. Restrict dangerous commands with ACLs to prevent it.

  5. 05

    RPO/RTO per keyspace: caches (RPO irrelevant, RTO = warm-up time), sessions (RPO seconds, RTO minutes), queues (RPO near zero, backed by the database). Write them down and test them.

  6. 06

    DR drills: restore last night's backup into a scratch environment monthly, measure how long it takes (RTO), check data freshness (RPO), and document the steps.

Code & diagrams

restore-drill.shbash
# Monthly restore drill
aws s3 cp s3://acme-redis-backups/prod/dump-20260927T020000Z.rdb ./dump.rdb
docker run -d --name restore-test -v "$PWD":/data redis:8 redis-server --dir /data --dbfilename dump.rdb --appendonly no
sleep 5
docker exec restore-test redis-cli DBSIZE
(integer) 18244213                     # compare with the production DBSIZE at backup time
docker exec restore-test redis-cli INFO persistence | grep loading
loading:0
docker exec restore-test redis-cli --scan --pattern 'sess:*' --count 1000 | head -3
# Record: restore time, key count, spot checks -> DR report

Interview problem

The problem

Someone ran FLUSHALL in production

At 14:05 an engineer accidentally ran FLUSHALL on the production session store (primary with two replicas, AOF everysec, nightly RDB backups). Everyone is logged out. What happens, how do you recover, and how do you prevent it?

The interviewer follows up

01

Why don't replicas protect against a mistaken FLUSHALL?

Explain it without notes

01

Explain how to upgrade a Sentinel-managed Redis with no downtime.

Practice

01

Run the restore drill script against one of your lab backups and record restore time and key count.

Trade-offs

  • ↔

    More frequent backups improve RPO but add fork and storage costs; take them from replicas.

Done when you can

  • I can upgrade Redis with rolling failovers and plan version and engine migrations.

  • I have tested backups, per-keyspace RPO/RTO, and a plan for human error.