Chapter 5 / 5
Production Observability
Four incidents about the observability platform itself: the monitor that silently stopped monitoring, the data you needed but no longer had, the outage only an outsider could see, and the dashboard nobody could find.
Every case so far assumed the observability stack itself was healthy, complete, and easy to use. In production it's a system like any other: it breaks, it forgets, it has blind spots, and it sprawls. This final chapter makes the lab production-grade and ends by tearing it all down.
0/4 · 0%
- CASE 5.1SEV2prometheus · alertmanager6 days undetected
“Last Tuesday's new alert never fired. It was never loaded — the config reload had been failing silently for six days.”
Who monitors the monitor?
- CASE 5.2LABmetrics storageAn afternoon to set up remote storage
“Finance wants Q3's checkout latency for the quarterly review. Prometheus keeps 15 days. And last week's restart left a 20-minute hole.”
The data we no longer had
- CASE 5.3SEV1public endpoint · TLS50 min
“Every dashboard was green for 50 minutes while every customer saw a certificate error.”
Green inside, broken outside
- CASE 5.4SEV3grafanaA sprint of cleanup
“During the outage, three engineers opened three different 'ShopLite' dashboards. Two were broken, and the third used the old metric names.”
214 dashboards and none of them right