Operating at Scale
The practices that keep reliability improving as the team and system grow: rehearsing failure, changing things safely, keeping on-call humane, and handling the incidents customers can see.
The last four shifts step back from single incidents. They're about the habits that decide how often incidents happen and how well a team copes with them: deliberate failure testing, progressive delivery, sustainable on-call, and honest communication when customers are directly affected.
- SHIFT 3.1PLANNING3 decisions
Breaking things on purpose
Game day plan review — 'kill the payments service in production and see what happens'
- SHIFT 3.2SEV23 decisions
A config change at peak hour
Change request — raise the ShopLite DB connection pool from 10 to 50 on all tasks, now, during the evening peak
- SHIFT 3.3PLANNING3 decisions
Your teammate is burning out
On-call review — Ravi was paged 31 times last week, 9 of them between midnight and 6 a.m.; he's said he wants to leave the rotation
- SHIFT 3.4SEV13 decisions
Customers were charged twice
SEV1 — support reports customers seeing duplicate charges on their cards for single orders; social media posts appearing