When to Reach for Logs, Metrics, or Traces
Metrics answer 'is something wrong right now, and how wrong' — they're cheap to store at high resolution and are what should page someone, but they can't tell you why. Logs answer 'what exactly happened on this one request or process,' with full context, but they're expensive to search at scale and useless for spotting a trend across millions of events. Traces answer 'where did the time go across a distributed call chain' — which service, which downstream call, which retry — and are the only one of the three that reconstructs causality across service boundaries. A common failure mode is alerting directly off logs (grepping for ERROR and paging on count), which produces noisy, low-precision alerts because logs weren't designed to be aggregated that way; use metrics for alerting and reserve logs for the investigation that follows. Another is adding distributed tracing last, after logs and metrics, when it's actually the fastest path to finding the actual bottleneck in a multi-service request — a single trace showing 800ms spent in a downstream auth call answers in seconds what could take an hour of log correlation across services.
Teams that route all three through one flag system tend to end up with flags nobody's sure are safe to delete. Each has a different lifecycle and deserves different tooling.
A dashboard showing a healthy average response time can hide a bad experience for one in a hundred requests. For any system with real concurrency, that's not a rare edge case.
An N+1 query pattern looks completely reasonable in a diff — one clean loop over a list, one query inside it. The problem only shows up under realistic data volume, which most review environments don't have.