Command Palette

Search for a command to run...

Hectal

Casebook · 6 chapters · 24 incidents

Every incident, on one board

Each card is a real class of production incident. Work them in order: every case assumes the tools and habits from the ones before it.

CH 0

Seeing Anything at All

Four incidents that teach the three signals: metrics tell you THAT something is wrong, logs tell you WHAT, traces tell you WHERE.

CH 1

Metrics in Anger

Four incidents where the metrics were there, but reading them wrong, or collecting them wrong, made things worse.

CH 2

Logs at Scale

Four incidents about logs as a system in their own right: their volume, what they must never contain, how to line them up with change, and how to parse the ones you didn't write.

CH 3

Tracing Deep Dive

Four incidents that only make sense as a trace: death by a thousand queries, retries that amplify an outage, a trace that silently breaks at a queue, and the trace you needed but didn't keep.

CH 4

Alerts & SLOs

Four incidents about the alerting itself: pages nobody should get, reliability nobody could agree on, thresholds that are both too noisy and too slow, and the alert that couldn't fire.

CH 5

Production Observability

Four incidents about the observability platform itself: the monitor that silently stopped monitoring, the data you needed but no longer had, the outage only an outsider could see, and the dashboard nobody could find.