Structuring a Postmortem People Actually Read
A postmortem that lists a single root cause is usually hiding the real story. Real incidents almost always have multiple contributing factors that combined at the wrong moment: a deploy that was individually safe, a monitoring gap that delayed detection, an on-call runbook that was out of date. Structure the document around a timeline built from timestamps and logs, not memory, since memory reorders events to make the story cleaner than it was. Separate 'what happened' from 'what we're doing about it' into distinct sections, and write action items as owned, dated, and specific ('add a p99 latency alert on the checkout service, owner: X, due: date') rather than vague aspirations like 'improve monitoring.' The single biggest predictor of whether a postmortem actually prevents a repeat incident is whether its action items get tracked to completion in the same system as regular engineering work, not left in a document nobody revisits. Blamelessness isn't about avoiding names in the timeline, it's about the review asking 'what made this the reasonable decision at the time' instead of 'who should have known better.'
Team size and deploy friction are better predictors of when to split a service than technical elegance. Splitting too early adds distributed-systems cost before the org is big enough to need it.
A queue that grows without bound isn't absorbing load, it's deferring an outage. Real backpressure means the system tells upstream producers to slow down before that happens.
An SLO without an error budget policy is just a target nobody's accountable to. The budget is what turns 'be reliable' into an actual decision-making tool for shipping velocity.