Topic 9.3
Choosing Durability: RDB, AOF, Both, or None
In one line
Pick persistence per workload: none for pure caches, RDB for fast restarts with minutes of acceptable loss, AOF everysec (usually with an RDB preamble) for about a second of loss, plus replicas and off-host backups. For data that must never lose an acknowledged write, Redis alone isn't enough.
Think of it like this
How you protect documents. Scratch notes need no backup; a novel draft gets a copy every evening (RDB); a legal contract is saved after every paragraph and copied to another building (AOF + replicas + offsite backups). You pick protection by what losing it would cost.
Key ideas
- 01
Pure cache: persistence off (or RDB only to warm faster after restarts). Loss means misses, not data loss.
- 02
Sessions, rate limits, leaderboards: AOF
everysec+ a replica, or RDB every few minutes if a few minutes of loss is fine. Decide based on the user-visible impact. - 03
Queues, streams, idempotency keys, locks: AOF
everysecat least, replicas with automatic failover, and designs that tolerate the remaining loss window (idempotency, reconciliation against the database). - 04
Restart time matters: loading 50 GB takes time (minutes for big AOFs), during which the node doesn't serve. RDB or an RDB-preamble AOF loads fastest. Replicas let you fail over instead of waiting for a reload.
- 05
Backups: copy RDB files (or
BGSAVEoutput) to object storage on a schedule, encrypt them, keep several generations, and test restores regularly. A backup you've never restored is a hope, not a backup. - 06
Beyond Redis: Amazon MemoryDB stores writes in a multi-AZ transaction log before acknowledging, giving durable, strongly consistent writes with Redis/Valkey compatibility at higher write latency. For money and orders, the database remains the source of truth.
Code & diagrams
Config Loss on crash Restart speed Write overhead
none everything instant (empty) none
RDB every 5 min up to ~5 min fast fork per snapshot
AOF everysec (+RDB preamble) ~1-2 s medium small
AOF always ~current batch medium high (disk sync)
AOF everysec + replica ~1-2 s, or replica lag on failover small
MemoryDB (durable log) none acknowledged managed higher write latency#!/usr/bin/env bash
set -euo pipefail
# Run against a replica to keep fork cost off the primary
redis-cli -h redis-replica-1 BGSAVE
while [ "$(redis-cli -h redis-replica-1 INFO persistence | grep -c 'rdb_bgsave_in_progress:1')" = "1" ]; do sleep 2; done
ts=$(date -u +%Y%m%dT%H%M%SZ)
redis-cli -h redis-replica-1 --rdb /backups/dump-$ts.rdb # streams the RDB to this host
aws s3 cp /backups/dump-$ts.rdb s3://acme-redis-backups/prod/ --sse aws:kms
# Restore test (weekly): start redis:8 with the file in a scratch container and compare DBSIZEInterview problem
The problem
Durable session system
User sessions must survive a Redis restart. Discuss RDB, AOF, fsync policy, recovery, the data-loss window, replication and backups, and state exactly what can still be lost.
You're given
- 10M sessions, ~6 GB
- Mass logout is unacceptable
- Losing a few seconds of session updates is fine
The interviewer follows up
Redis primary fails immediately after acknowledging a write. What happens?
The AOF has grown to 200 GB. What happens and what do you do?
When it breaks
Persistence disabled on a primary with automatic restart and replicas
What you see
The primary restarts empty and quickly; replicas sync from it and wipe their own data too. Everything is gone.
Fix & prevent
Enable persistence on primaries, or disable automatic restarts so a replica is promoted instead of an empty primary coming back.
Explain it without notes
"Redis restarts. How much data can be lost?" Answer for four configurations.
Practice
Write a persistence policy for three keyspaces in your system: cache, sessions, and a job queue.
Trade-offs
- ↔
Stronger durability costs write latency, disk I/O and restart time; the right level depends on the cost of losing each keyspace.
Done when you can
I can choose persistence per workload and state its exact loss window.
I back up RDB files off-host and test restores.
I know when Redis durability isn't enough and what to use instead.