← Back to search results

Diagnosing Connection Pool Exhaustion Under Load

A service that responds fine under normal traffic but starts hanging or returning 504s the moment concurrent requests spike is exhibiting a classic symptom of connection pool exhaustion. CPU and memory graphs stay flat because the process isn't doing more work, it's waiting: every incoming request blocks on acquiring a database connection, and once the pool's max size is reached, new requests queue behind whichever request currently holds a connection open. Confirm it by checking your pool's active/idle/waiting metrics during the incident window, not just database-side connection counts, since the bottleneck is almost always the application-side pool limit rather than the database's own max_connections setting. The fix is rarely 'just raise the pool size' — a larger pool just moves the bottleneck to the database, which has its own connection ceiling and per-connection memory cost. Instead, profile what's holding connections open longest (usually a slow query or a transaction that does non-database work while a connection is checked out), fix that first, and only then tune pool size against your actual concurrency ceiling, using a formula like connections = ((core_count * 2) + effective_spindle_count) as a starting point, not a target.

Related documents

A queue that grows without bound isn't absorbing load, it's deferring an outage. Real backpressure means the system tells upstream producers to slow down before that happens.

Rotating a database password or API key is simple in isolation. Doing it without an outage across every process holding the old credential in memory is the actual hard part.

Team size and deploy friction are better predictors of when to split a service than technical elegance. Splitting too early adds distributed-systems cost before the org is big enough to need it.