SRE · an on-call simulator
It's 3:12 a.m.
What do you do first?
Reliability engineering is mostly judgement: what to do first, what to leave alone, when to roll back, who to wake up, what to say to customers, and what to change afterwards. You can't learn judgement by reading a list of best practices. In these 16 shifts you're on call for ShopLite: you make 48 decisions under pressure and see the consequences of each one.
How a shift works
- 01
You get paged
A real-feeling situation with the signals you'd actually have: alerts, graphs, messages from colleagues.
- 02
You decide
Several decision points, each with four plausible options. Choices lock in, and the clock keeps running.
- 03
Consequences
See what your choice did, then what the other options would have done, and why.
- 04
Postmortem
A blameless postmortem of the incident, with action items, and the SRE practice it teaches.
Chapters
Incident Response
SEV2SEV2SEV1PLANYour first pages as ShopLite's on-call: assess, declare, mitigate, coordinate, and learn — the loop every other SRE practice builds on.
Reliability by Design
PLANPLANSEV3PLANThe decisions made before anything breaks: how reliable to be, what to do when the budget runs out, which work to automate, and whether a launch is ready.
Capacity & Resilience
SEV2SEV1SEV1SEV2When load or failure exceeds what the system was built for: surges, cascades, lost availability zones, and a database running out of room at 2 a.m.
Operating at Scale
PLANSEV2PLANSEV1The practices that keep reliability improving as the team and system grow: rehearsing failure, changing things safely, keeping on-call humane, and handling the incidents customers can see.