N+1 Queries and Why They Pass Code Review
The N+1 pattern — fetch a list of N records, then issue one additional query per record to fetch related data — reads as clean, idiomatic code in a pull request: a loop, a call to fetch related data, nothing that looks obviously wrong. It fails to show up in code review because reviewers evaluate correctness and style, not query count, and it often fails to show up in local testing too, because dev and staging databases typically have a handful of seed rows where twenty extra queries add a few milliseconds nobody notices. Under production data volume — a list of 500 orders instead of 5 — the same code issues 501 queries instead of 6, and latency degrades roughly linearly with list size instead of staying flat, which is the actual signature that distinguishes an N+1 problem from general slowness: response time that scales with result-set size rather than staying roughly constant. The fix is almost always eager loading — fetching the related data in a single batched query (a JOIN, or an ORM's explicit eager-load / includes / prefetch call) instead of one query per parent row — but the harder problem is catching it before production, since it's invisible in small-data environments. The most reliable catch is a query-count assertion in tests (assert this endpoint issues at most N queries) run against a seeded dataset large enough to expose the pattern, combined with query-count-per-request logging in staging under realistic data volume, rather than relying on code review to spot it by inspection.
GraphQL solves a real problem — overfetching and underfetching across many client shapes. Most internal, single-client APIs don't have that problem, and adopt the complexity without the benefit.
A dashboard showing a healthy average response time can hide a bad experience for one in a hundred requests. For any system with real concurrency, that's not a rare edge case.
When a service starts timing out under traffic spikes even though CPU and memory look fine, the culprit is usually a saturated connection pool, not the database itself. Here is how to confirm it and size the pool correctly.