Incident Response
Your first pages as ShopLite's on-call: assess, declare, mitigate, coordinate, and learn — the loop every other SRE practice builds on.
Most incident damage comes not from the original failure but from the first thirty minutes of the response: acting before understanding, debugging when you should be mitigating, working alone when you should be coordinating. These four shifts build the habits that separate a 15-minute blip from a 3-hour outage.
- SHIFT 0.1SEV23 decisions
Your first page
CheckoutErrorBudgetFastBurn — shoplite /checkout error ratio 18% (threshold 7.2%)
- SHIFT 0.2SEV23 decisions
The VP wants a root cause. Now.
SEV2 ongoing — product pages 40% slower since a config change; VP Engineering has joined the channel
- SHIFT 0.3SEV13 decisions
Fourteen people in the channel
SEV1 — shoplite, payments, and search all returning 5xx; status page shows full outage
- SHIFT 0.4PLANNING3 decisions
The engineer who dropped the table
Postmortem review — last Tuesday's 2-hour outage started when an engineer ran a migration against production by mistake