Why p99 Latency Matters More Than the Average
Average latency is dominated by the bulk of fast requests and hides the tail, which is exactly the part users notice: nobody complains that the site is '40ms slower on average,' they complain that it 'randomly hangs sometimes.' The reason the tail matters disproportionately is compounding: if a single request touches five downstream services, each with a 1% chance of hitting their own p99, the probability that the overall request hits at least one slow dependency is much higher than 1% — with five independent calls each with a 1-in-100 chance of a slow response, roughly one in twenty user-facing requests will experience at least one slow downstream call. This is why service-level objectives are usually written against p95 or p99, not average, and why a single slow dependency in a request's fan-out can dominate the parent request's tail even if it's fast on average. Investigating tail latency requires different tools than investigating averages: aggregate metrics tell you the tail exists, but finding the cause usually requires looking at individual slow traces to find what they have in common — a specific downstream host, a specific cache-miss pattern, a specific query plan that only triggers for certain input shapes.
An N+1 query pattern looks completely reasonable in a diff — one clean loop over a list, one query inside it. The problem only shows up under realistic data volume, which most review environments don't have.
Teams new to observability tend to over-invest in one pillar and under-invest in the other two. Each answers a different kind of question, and they're not substitutes for each other.
GraphQL solves a real problem — overfetching and underfetching across many client shapes. Most internal, single-client APIs don't have that problem, and adopt the complexity without the benefit.