Command Palette

Search for a command to run...

Hectal
PHASE 9Intermediate ~8 min· topic 1 of 3

Topic 9.1

RDB Snapshots: Fork, Copy-on-Write, Recovery

In one line

An RDB file is a compact point-in-time snapshot of the whole dataset, written by a forked child process (BGSAVE). It's great for backups, fast restarts and replica full syncs, but anything written after the last snapshot is lost on a crash.

0/3 · 0%

Think of it like this

Taking a photo of a whiteboard every hour. The photo is small and quick to restore from, but if the board is wiped at 10:59, everything written since the 10:00 photo is gone.

Key ideas

  1. 01

    BGSAVE forks the process; the child writes the dataset to a temp file and atomically renames it to dump.rdb, while the parent keeps serving clients. SAVE does it on the main thread and blocks everything; never use it in production.

  2. 02

    Schedules: save 3600 1 300 100 60 10000 (the default in Redis 7) means snapshot if at least 1 change in an hour, 100 in 5 minutes, or 10,000 in a minute. save "" disables automatic snapshots.

  3. 03

    Fork cost: forking copies the page tables, which takes time proportional to memory (roughly 10–20 ms per GB on typical hardware, worse on some VMs), and blocks the main thread during the fork. INFO stats → latest_fork_usec shows it. Copy-on-write then duplicates any page the parent modifies during the snapshot, so heavy writes can add up to the dataset size in extra memory.

  4. 04

    Failures: if a background save fails (disk full, permissions), stop-writes-on-bgsave-error yes (default) makes Redis reject writes, so you notice. Monitor rdb_last_bgsave_status.

  5. 05

    Recovery: on startup Redis loads the RDB file (fast, binary, compressed with LZF), typically much faster than replaying an AOF. RDB files are also the unit of backup: copy them off the host regularly and test restores.

Code & diagrams

rdb.redisredis
127.0.0.1:6379> CONFIG GET save
1) "save"
2) "3600 1 300 100 60 10000"
127.0.0.1:6379> BGSAVE
Background saving started
127.0.0.1:6379> INFO persistence
rdb_changes_since_last_save:0
rdb_bgsave_in_progress:0
rdb_last_save_time:1727520000
rdb_last_bgsave_status:ok
rdb_last_bgsave_time_sec:4
rdb_last_cow_size:268435456          # 256 MB copied during the last snapshot
127.0.0.1:6379> INFO stats
latest_fork_usec:152034              # 152 ms main-thread pause for fork
fork-cow.mermaiddiagram
Rendering diagram…

When it breaks

BGSAVE on a 60 GB instance with transparent huge pages

What you see

The fork pauses Redis for over a second, and copy-on-write of 2 MB pages balloons memory until the OOM killer ends the process.

Fix & prevent

Disable THP, keep shards smaller, leave memory headroom, and run snapshots on replicas where possible.

Disk fills up

What you see

BGSAVE fails; with stop-writes-on-bgsave-error yes, every write returns MISCONF errors.

Fix & prevent

Alert on disk usage and rdb_last_bgsave_status; fix disk space, then writes resume after a successful save.

Explain it without notes

01

Why doesn't BGSAVE block Redis while writing the snapshot, and what does it still cost?

Practice

01

Load 2 GB of data in the lab, run BGSAVE while writing heavily, and read latest_fork_usec and rdb_last_cow_size.

Trade-offs

  • ↔

    RDB: compact, fast restarts, easy backups; loses everything since the last snapshot.

Done when you can

  • I can explain BGSAVE, fork, copy-on-write and their costs.

  • I monitor snapshot status and fork time.