4 chapters · 16 shifts · 48 decisions
The on-call rota
Each shift puts you on call for ShopLite. You make the calls, see the consequences, and then read the postmortem and the SRE practice behind it.
Incident Response
Your first pages as ShopLite's on-call: assess, declare, mitigate, coordinate, and learn — the loop every other SRE practice builds on.
Your first page
CheckoutErrorBudgetFastBurn — shoplite /checkout error ratio 18% (threshold 7.2%)
The VP wants a root cause. Now.
SEV2 ongoing — product pages 40% slower since a config change; VP Engineering has joined the channel
Fourteen people in the channel
SEV1 — shoplite, payments, and search all returning 5xx; status page shows full outage
The engineer who dropped the table
Postmortem review — last Tuesday's 2-hour outage started when an engineer ran a migration against production by mistake
Reliability by Design
The decisions made before anything breaks: how reliable to be, what to do when the budget runs out, which work to automate, and whether a launch is ready.
Product wants 100%
Planning meeting — Product asks for a '100% uptime' target for checkout in the new customer contract
Out of budget, a week before launch
Error budget report — checkout SLO at 99.71% for the window; budget exhausted; big feature launch scheduled next Tuesday
40% of the week on restarts
On-call report — 23 manual interventions last week; 14 were 'restart the worker when the queue backs up'
Is the Diwali sale ready?
Launch readiness review — Diwali sale in 10 days; marketing expects 8–10× normal traffic for 6 hours
Capacity & Resilience
When load or failure exceeds what the system was built for: surges, cascades, lost availability zones, and a database running out of room at 2 a.m.
An influencer posted a link
ShopLite p99 3.8s and rising — traffic at 7× normal, unannounced; autoscaling at max tasks
One slow dependency, everything down
SEV1 — shoplite, search, and account pages all timing out; only payments was 'slightly slow' ten minutes ago
An availability zone goes dark
SEV1 — ap-south-1a network degradation; 50% of ShopLite requests failing; RDS primary is in 1a
The database is 95% full at 2 a.m.
PostgresStorageExhaustionPredicted — orders-db free storage 4.1 GB, projected full in ~3 hours
Operating at Scale
The practices that keep reliability improving as the team and system grow: rehearsing failure, changing things safely, keeping on-call humane, and handling the incidents customers can see.
Breaking things on purpose
Game day plan review — 'kill the payments service in production and see what happens'
A config change at peak hour
Change request — raise the ShopLite DB connection pool from 10 to 50 on all tasks, now, during the evening peak
Your teammate is burning out
On-call review — Ravi was paged 31 times last week, 9 of them between midnight and 6 a.m.; he's said he wants to leave the rotation
Customers were charged twice
SEV1 — support reports customers seeing duplicate charges on their cards for single orders; social media posts appearing