← Back to search results

Handling Backpressure in Event-Driven Systems

It's tempting to treat a message queue as infinite slack that decouples a slow consumer from a fast producer, but an unbounded queue just delays the failure and makes it worse when it arrives: by the time consumers fall far enough behind that anyone notices, there can be hours of backlog to drain, often processed against now-stale data. Real backpressure means the system actively signals upstream when it can't keep up, rather than silently absorbing unlimited work. Concretely, that means bounded queue depth with an explicit policy for what happens at the bound (reject new messages, apply load shedding to lower-priority ones, or block the producer), consumer-side concurrency limits tied to actual downstream capacity rather than an arbitrary worker count, and monitoring queue depth and consumer lag as first-class metrics with alerts, not just consumer error rate. For systems built on log-based queues specifically, consumer lag (how far behind the latest offset a consumer group is) is the single most useful backpressure signal, since it directly measures the gap between production and consumption rate rather than inferring it indirectly from downstream symptoms. The deeper principle: backpressure should propagate to whoever can actually do something about the load, which is usually further upstream than the immediate producer — a spike in write traffic often has a real cause (a client retry storm, a batch job that just started) that's worth surfacing rather than just absorbing.

Related documents

An SLO without an error budget policy is just a target nobody's accountable to. The budget is what turns 'be reliable' into an actual decision-making tool for shipping velocity.

An N+1 query pattern looks completely reasonable in a diff — one clean loop over a list, one query inside it. The problem only shows up under realistic data volume, which most review environments don't have.

Most postmortems get written once and never referenced again. The ones that change how a team operates share a specific structure: timeline, contributing factors, and action items with owners.