Command Palette

Search for a command to run...

Hectal
Chapter 4 / 5

Alerts & SLOs

Four incidents about the alerting itself: pages nobody should get, reliability nobody could agree on, thresholds that are both too noisy and too slow, and the alert that couldn't fire.

ShopLite now has a handful of alert rules, written one incident at a time. That's how most teams end up with hundreds of alerts and a tired on-call. This chapter adds Alertmanager for routing, then replaces ad-hoc thresholds with Service Level Objectives and burn-rate alerts, the approach from Google's SRE workbook, and finishes by testing alert rules like code.

0/4 · 0%
  1. CASE 4.1SEV2alertingA week of cleanup

    “On-call got paged 64 times last night. None of the pages needed a human. At 06:10 a real one was acknowledged and ignored.”

    64 pages, zero actions

  2. CASE 4.2LABshoplite checkoutOne SLO document and a set of recording rules

    “Product says reliability is fine at 99.4%. Engineering says we're on fire. Both are looking at the same graph.”

    '99.4% is fine' — or is it?

  3. CASE 4.3SEV2alerting · shoplite checkoutRules rewritten in an afternoon

    “The 5% threshold paged us for a 90-second blip — and stayed silent while a 1% error rate ate the whole month's budget in three days.”

    Too noisy and too slow at the same time

  4. CASE 4.4SEV1shoplite-api · alerting25 min outage; alert rules fixed the same day

    “ShopLite was completely down for 25 minutes. Not one alert fired. Every rule looked correct.”

    The alert that couldn't fire