Topic 9.1
RDB Snapshots: Fork, Copy-on-Write, Recovery
In one line
An RDB file is a compact point-in-time snapshot of the whole dataset, written by a forked child process (BGSAVE). It's great for backups, fast restarts and replica full syncs, but anything written after the last snapshot is lost on a crash.
Think of it like this
Taking a photo of a whiteboard every hour. The photo is small and quick to restore from, but if the board is wiped at 10:59, everything written since the 10:00 photo is gone.
Key ideas
- 01
BGSAVEforks the process; the child writes the dataset to a temp file and atomically renames it todump.rdb, while the parent keeps serving clients.SAVEdoes it on the main thread and blocks everything; never use it in production. - 02
Schedules:
save 3600 1 300 100 60 10000(the default in Redis 7) means snapshot if at least 1 change in an hour, 100 in 5 minutes, or 10,000 in a minute.save ""disables automatic snapshots. - 03
Fork cost: forking copies the page tables, which takes time proportional to memory (roughly 10–20 ms per GB on typical hardware, worse on some VMs), and blocks the main thread during the fork.
INFO stats→latest_fork_usecshows it. Copy-on-write then duplicates any page the parent modifies during the snapshot, so heavy writes can add up to the dataset size in extra memory. - 04
Failures: if a background save fails (disk full, permissions),
stop-writes-on-bgsave-error yes(default) makes Redis reject writes, so you notice. Monitorrdb_last_bgsave_status. - 05
Recovery: on startup Redis loads the RDB file (fast, binary, compressed with LZF), typically much faster than replaying an AOF. RDB files are also the unit of backup: copy them off the host regularly and test restores.
Code & diagrams
127.0.0.1:6379> CONFIG GET save
1) "save"
2) "3600 1 300 100 60 10000"
127.0.0.1:6379> BGSAVE
Background saving started
127.0.0.1:6379> INFO persistence
rdb_changes_since_last_save:0
rdb_bgsave_in_progress:0
rdb_last_save_time:1727520000
rdb_last_bgsave_status:ok
rdb_last_bgsave_time_sec:4
rdb_last_cow_size:268435456 # 256 MB copied during the last snapshot
127.0.0.1:6379> INFO stats
latest_fork_usec:152034 # 152 ms main-thread pause for forkWhen it breaks
BGSAVE on a 60 GB instance with transparent huge pages
What you see
The fork pauses Redis for over a second, and copy-on-write of 2 MB pages balloons memory until the OOM killer ends the process.
Fix & prevent
Disable THP, keep shards smaller, leave memory headroom, and run snapshots on replicas where possible.
Disk fills up
What you see
BGSAVE fails; with stop-writes-on-bgsave-error yes, every write returns MISCONF errors.
Fix & prevent
Alert on disk usage and rdb_last_bgsave_status; fix disk space, then writes resume after a successful save.
Explain it without notes
Why doesn't BGSAVE block Redis while writing the snapshot, and what does it still cost?
Practice
Load 2 GB of data in the lab, run BGSAVE while writing heavily, and read latest_fork_usec and rdb_last_cow_size.
Trade-offs
- ↔
RDB: compact, fast restarts, easy backups; loses everything since the last snapshot.
Done when you can
I can explain BGSAVE, fork, copy-on-write and their costs.
I monitor snapshot status and fork time.