Postmortem review — last Tuesday's 2-hour outage started when an engineer ran a migration against production by mistake
The engineer who dropped the table
You're facilitating the postmortem review meeting.
Briefing
Last Tuesday, Karan ran a database migration script intended for staging against production. It dropped and recreated the stock table; restoring from backup took two hours. Karan is in the meeting, clearly upset. So are his manager, two SREs, and a director.
Your shift
Decide before you look. Each choice is locked once made; the next decision unlocks after it. There's no perfect path in real incidents either, only better and worse judgement under uncertainty.
The director opens with: 'So, Karan, walk us through why you ran it on prod.'
As facilitator, how do you steer the meeting?
Debrief
Blameless postmortems aren't about being nice; they're about getting accurate information. People only describe what really happened if they're confident it won't be used against them. Look past 'human error' to the conditions that made the error easy and costly, and turn them into a small number of owned, tracked changes.
The postmortem
Blameless postmortem · draft
Summary. A migration intended for staging was run against production, dropping the stock table; restored from backup after 2 hours.
Impact. Checkout unavailable for 2 h 04 min; 14 minutes of stock updates had to be reconciled manually.
Timeline
- 15:12Migration run from a laptop with cached production credentials
- 15:13Checkout errors; incident declared 15:16
- 15:25Cause identified; restore from snapshot started
- 17:17Restore complete; checkout recovers
Root cause. Destructive migration executed against production because production credentials were active in the operator's terminal and the tool gave no indication of its target.
Contributing factors
- Standing production database credentials on engineer laptops.
- Migration tool had no environment banner or confirmation for destructive operations.
- Staging and production hostnames differed by one word.
- Restore procedure hadn't been rehearsed in 8 months, which is why recovery took 2 hours.
Action items
- preventRemove standing prod DB access; just-in-time access via SSO with 1-hour expiryplatform
- preventMigrations run only through CI with an explicit environment input and approval for prodkaran
- mitigateQuarterly restore drill with measured time-to-restoresre
The practice behind it
Blameless postmortems
A blameless postmortem assumes everyone acted reasonably given the information and tools they had, and asks why the system allowed a reasonable action to cause harm. It focuses on contributing factors, not a single root cause, and produces system changes, not personal commitments. Publishing postmortems widely, including the embarrassing ones, builds a culture where problems surface early.
Your turn
Rewrite 'Engineer error: ran migration on the wrong database' as a set of contributing factors.
Interview questions
What makes a good postmortem?