← Back to search results

Error Budgets: Turning Reliability Into a Number You Can Spend

A service level objective (say, 99.9% availability over 30 days) implies an error budget: the 0.1% of allowed unavailability, which at 99.9% over 30 days works out to about 43 minutes. The budget's real purpose isn't the number itself, it's the policy attached to spending it: while budget remains, the team can ship features and take reasonable risk, including deploys that carry some chance of a brief incident; once the budget is exhausted, the policy shifts toward reliability work — freezing risky launches, prioritizing the backlog of flaky-test and infra-debt tickets — until the budget recovers. This reframes reliability from an open-ended, unbounded goal ('always be more reliable') into a bounded resource with an explicit tradeoff against feature velocity, which is what makes it something a team can actually make decisions against instead of a vague aspiration everyone nods at. The number itself matters less than actually enforcing the policy: teams that track error budget burn but never actually slow down feature work when it's exhausted have an SLO in name only, and the metric becomes decoration rather than a real constraint. Setting the SLO number itself should come from what users actually need and notice, not from what's technically achievable — a 99.99% target on a service where users can't perceive the difference from 99.9% just makes the budget artificially tight for no real benefit.

Related documents

A queue that grows without bound isn't absorbing load, it's deferring an outage. Real backpressure means the system tells upstream producers to slow down before that happens.

Most postmortems get written once and never referenced again. The ones that change how a team operates share a specific structure: timeline, contributing factors, and action items with owners.

Team size and deploy friction are better predictors of when to split a service than technical elegance. Splitting too early adds distributed-systems cost before the org is big enough to need it.