Chapter 4 / 5
Alerts & SLOs
Four incidents about the alerting itself: pages nobody should get, reliability nobody could agree on, thresholds that are both too noisy and too slow, and the alert that couldn't fire.
ShopLite now has a handful of alert rules, written one incident at a time. That's how most teams end up with hundreds of alerts and a tired on-call. This chapter adds Alertmanager for routing, then replaces ad-hoc thresholds with Service Level Objectives and burn-rate alerts, the approach from Google's SRE workbook, and finishes by testing alert rules like code.
0/4 · 0%
- CASE 4.1SEV2alertingA week of cleanup
“On-call got paged 64 times last night. None of the pages needed a human. At 06:10 a real one was acknowledged and ignored.”
64 pages, zero actions
- CASE 4.2LABshoplite checkoutOne SLO document and a set of recording rules
“Product says reliability is fine at 99.4%. Engineering says we're on fire. Both are looking at the same graph.”
'99.4% is fine' — or is it?
- CASE 4.3SEV2alerting · shoplite checkoutRules rewritten in an afternoon
“The 5% threshold paged us for a 90-second blip — and stayed silent while a 1% error rate ate the whole month's budget in three days.”
Too noisy and too slow at the same time
- CASE 4.4SEV1shoplite-api · alerting25 min outage; alert rules fixed the same day
“ShopLite was completely down for 25 minutes. Not one alert fired. Every rule looked correct.”
The alert that couldn't fire