Topic 9.2
AOF: The Append-Only File and fsync
In one line
The AOF logs every write command; on restart Redis replays it. appendfsync decides how often it's flushed to disk (always, every second, or left to the OS), which sets the data-loss window. Rewrites compact the log in the background; Redis 7 stores it as multi-part files with a manifest.
Think of it like this
A cashier writing every sale in a notebook as it happens. If the till is stolen, you can reconstruct the day by re-reading the notebook. How much you lose depends on how often the notebook is locked in the safe (fsync).
Key ideas
- 01
Enable with
appendonly yes. Every write command is appended to an in-memory buffer and written to the AOF file;fsyncforces it onto disk. - 02
appendfsync always: fsync on every write batch; the safest (loses at most the current batch) and the slowest, since write latency includes disk sync.everysec(default): fsync once per second in a background thread; loses up to about 1–2 seconds on a crash, with little overhead.no: the OS decides (often ~30 s on Linux); fastest, largest loss window. - 03
Rewrite: the log grows forever as keys are updated, so Redis periodically rewrites it (
BGREWRITEAOF, or automatically viaauto-aof-rewrite-percentage 100andauto-aof-rewrite-min-size 64mb). A forked child writes a compact base; new writes during the rewrite go to a new incremental file. - 04
Redis 7 multi-part AOF: files live in
appenddirname(defaultappendonlydir): a base file (RDB-format whenaof-use-rdb-preamble yes), incremental AOF files, and a manifest. Restart loads the base (fast) then replays increments. - 05
Corruption: a crash mid-write can leave a truncated last command;
aof-load-truncated yes(default) loads what's valid and logs a warning. Real corruption in the middle needsredis-check-aof --fix, after backing up the file. - 06
Durability is still not a database's: an acknowledged write may be in memory but not yet fsynced (everysec), and replication is asynchronous.
WAITAOF(7.2) lets a client wait until its writes are fsynced locally and/or on N replicas, trading latency for durability per operation.
Code & diagrams
127.0.0.1:6379> CONFIG SET appendonly yes
OK
127.0.0.1:6379> CONFIG GET appendfsync
1) "appendfsync"
2) "everysec"
127.0.0.1:6379> INFO persistence
aof_enabled:1
aof_rewrite_in_progress:0
aof_last_bgrewrite_status:ok
aof_last_write_status:ok
aof_delayed_fsync:0 # >0 means the disk couldn't keep up
127.0.0.1:6379> SET payment:9 captured
OK
127.0.0.1:6379> WAITAOF 1 1 100 # fsynced locally AND on 1 replica, 100 ms timeout
1) (integer) 1
2) (integer) 1$ ls appendonlydir/
appendonly.aof.3.base.rdb # compact base (RDB preamble)
appendonly.aof.5.incr.aof # writes since the last rewrite
appendonly.aof.manifest
$ cat appendonlydir/appendonly.aof.manifest
file appendonly.aof.3.base.rdb seq 3 type b
file appendonly.aof.5.incr.aof seq 5 type iWhen it breaks
appendfsync always on a busy instance with slow network disks
What you see
Every write waits for disk sync; throughput collapses and latency rises to disk latency (ms instead of µs).
Fix & prevent
Use everysec and replicas, or WAITAOF only for the few writes that need it; use fast local SSDs.
Disk slower than the write rate with everysec
What you see
aof_delayed_fsync grows; Redis may delay writes to keep the one-second guarantee, causing latency spikes.
Fix & prevent
Faster disks, move persistence to replicas, or reduce write volume.
Explain it without notes
Compare the three appendfsync settings by data-loss window and performance.
Practice
Enable AOF, write keys in a loop, kill -9 the Redis process, restart, and check how many writes survived with everysec vs always.
Trade-offs
- ↔
AOF: small loss window and durable logs, at the cost of disk writes, slower restarts, and rewrite overhead.
Done when you can
I can configure AOF and explain fsync policies and rewrites.
I know the Redis 7 multi-part AOF layout and how to repair a truncated AOF.