Command Palette

Search for a command to run...

Hectal
on-callShift 2.4 · Capacity & Resilience
SEV2

PostgresStorageExhaustionPredicted — orders-db free storage 4.1 GB, projected full in ~3 hours

The database is 95% full at 2 a.m.

Primary on-call, woken at 02:07.

capacity alerts that give you timesafe emergency changesfinding what's consuming resourcesnot making it worse

Briefing

A predictive alert (the pattern from the Observability course, Case 1.4) says storage will run out in about three hours. When Postgres runs out of disk, it stops accepting writes, so checkout would fail completely.

Your shift

Decide before you look. Each choice is locked once made; the next decision unlocks after it. There's no perfect path in real incidents either, only better and worse judgement under uncertainty.

02:07· decision 1 of 3

Free storage has been falling fast since about midnight.

  • orders-db FreeStorageSpace: 4.1 GB (was 38 GB at 23:30) · max_allocated_storage: disabled
  • TransactionLogsDiskUsage: 31 GB and rising

What's your first action?

Debrief

With a resource approaching exhaustion, first buy time with a safe, reversible change, then find the consumer, then fix it in coordination with its owner. Predictive alerts are what make this calm: three hours of warning turned a potential outage into an overnight task.

The postmortem

Blameless postmortem · draft

Summary. An abandoned logical replication slot caused Postgres to retain WAL, filling storage; mitigated by increasing storage and dropping the slot.

Impact. No customer impact; the database would have stopped accepting writes in ~3 hours without intervention.

Timeline

  • 00:02CDC pipeline crashes; slot becomes inactive
  • 02:07Predictive storage alert pages on-call
  • 02:15Allocated storage increased
  • 02:35Slot dropped after confirming with analytics on-call; WAL released

Root cause. An inactive logical replication slot prevented WAL cleanup.

Contributing factors

  • No alert on inactive slots or retained WAL.
  • Storage autoscaling disabled.
  • New replication consumer added without database-owner review.

Action items

  • detectAlert on inactive replication slots and retained WAL > 10 GBsre
  • mitigateEnable RDS storage autoscaling with a maxplatform
  • preventSet max_slot_wal_keep_size to cap WAL retention per slotdba

The practice behind it

Buy time, then fix

When a resource is heading to exhaustion, the first move is usually a safe way to extend the deadline: add storage, raise a limit, scale a pool, pause a batch job. That converts a crisis into a normal investigation. The fix comes after, with the right people involved.

Your turn

01

Name two Postgres settings or features that could have prevented or limited this incident.

Interview questions

01

A production database is about to run out of disk. What do you do?

0/4 · 0%