PostgresStorageExhaustionPredicted — orders-db free storage 4.1 GB, projected full in ~3 hours
The database is 95% full at 2 a.m.
Primary on-call, woken at 02:07.
Briefing
A predictive alert (the pattern from the Observability course, Case 1.4) says storage will run out in about three hours. When Postgres runs out of disk, it stops accepting writes, so checkout would fail completely.
Your shift
Decide before you look. Each choice is locked once made; the next decision unlocks after it. There's no perfect path in real incidents either, only better and worse judgement under uncertainty.
Free storage has been falling fast since about midnight.
- orders-db FreeStorageSpace: 4.1 GB (was 38 GB at 23:30) · max_allocated_storage: disabled
- TransactionLogsDiskUsage: 31 GB and rising
What's your first action?
Debrief
With a resource approaching exhaustion, first buy time with a safe, reversible change, then find the consumer, then fix it in coordination with its owner. Predictive alerts are what make this calm: three hours of warning turned a potential outage into an overnight task.
The postmortem
Blameless postmortem · draft
Summary. An abandoned logical replication slot caused Postgres to retain WAL, filling storage; mitigated by increasing storage and dropping the slot.
Impact. No customer impact; the database would have stopped accepting writes in ~3 hours without intervention.
Timeline
- 00:02CDC pipeline crashes; slot becomes inactive
- 02:07Predictive storage alert pages on-call
- 02:15Allocated storage increased
- 02:35Slot dropped after confirming with analytics on-call; WAL released
Root cause. An inactive logical replication slot prevented WAL cleanup.
Contributing factors
- No alert on inactive slots or retained WAL.
- Storage autoscaling disabled.
- New replication consumer added without database-owner review.
Action items
- detectAlert on inactive replication slots and retained WAL > 10 GBsre
- mitigateEnable RDS storage autoscaling with a maxplatform
- preventSet
max_slot_wal_keep_sizeto cap WAL retention per slotdba
The practice behind it
Buy time, then fix
When a resource is heading to exhaustion, the first move is usually a safe way to extend the deadline: add storage, raise a limit, scale a pool, pause a batch job. That converts a crisis into a normal investigation. The fix comes after, with the right people involved.
Your turn
Name two Postgres settings or features that could have prevented or limited this incident.
Interview questions
A production database is about to run out of disk. What do you do?