Command Palette

Search for a command to run...

Hectal
on-callShift 1.2 · Reliability by Design
PLANNING

Error budget report — checkout SLO at 99.71% for the window; budget exhausted; big feature launch scheduled next Tuesday

Out of budget, a week before launch

You're the SRE representing reliability in the weekly release planning meeting.

applying an error budget policynegotiating trade-offsreliability work prioritisationrisk-reducing a launch

Briefing

The error budget policy (agreed last quarter) says: when the budget is exhausted, feature releases pause and reliability work takes priority. The 'one-click reorder' launch is next Tuesday, with a marketing campaign already booked.

Your shift

Decide before you look. Each choice is locked once made; the next decision unlocks after it. There's no perfect path in real incidents either, only better and worse judgement under uncertainty.

Wed 14:00· decision 1 of 3

The feature team lead: 'The policy is guidance, right? This launch is too important to delay.'

What's your position?

Debrief

Error budgets work when both sides treat them as a shared decision tool. Apply the policy, then solve the problem together: prioritise the fix that spent the budget, reduce the launch's blast radius, and improve the policy based on experience.

The postmortem

Blameless postmortem · draft

Summary. Planning outcome: launch proceeds as a staged rollout after the reliability fix that consumed most of the budget.

Impact. N/A — planning session.

Timeline

  • WedBudget exhausted; release planning discussion
  • FriConnection-pool fix deployed
  • Tue–ThuFeature rolled out 5% → 25% → 100% with automated rollback

Root cause. N/A

Contributing factors

  • Two incidents from one unfixed connection-pool bug consumed ~80% of the budget.

Action items

  • processClarify error budget policy: definitions, exception process, staged-launch rulessre + product
  • mitigateAutomated rollback on burn rate for flag-gated launchesplatform

The practice behind it

Error budget policies

A written agreement, made in advance by product and engineering, about what happens at different budget levels: normal velocity when there's plenty left, extra caution as it shrinks, reliability focus when it's gone. Its value is that the decision was made calmly, before the pressure of a specific launch. Use it to start conversations, not to end them.

Your turn

01

Which work should count as 'reliability work' when the budget is exhausted?

Interview questions

01

What happens when a service exhausts its error budget?

0/4 · 0%