← Back to search results

Rate Limiting Strategies at the API Gateway

Fixed-window rate limiting (e.g. 1000 requests per minute, counted from the top of each minute) is the cheapest to implement but has a boundary problem: a client can send 1000 requests in the last second of one window and another 1000 in the first second of the next, doubling the effective limit in a two-second span. Sliding window log fixes this by tracking exact timestamps of every request in the trailing window, which is precise but memory-expensive at scale. Sliding window counter approximates the sliding log by weighting the previous window's count, giving most of the accuracy at a fraction of the memory cost, and is what most production API gateways actually implement. Token bucket takes a different approach entirely: a bucket refills at a steady rate and each request consumes a token, which naturally allows bursts up to the bucket size while still enforcing a long-run average rate — this is usually the right choice for APIs where legitimate clients are bursty (a dashboard that fires ten requests on page load, then goes quiet) rather than steady. Whichever algorithm you pick, rate limit on the dimension that actually maps to abuse (per-API-key or per-user, not per-IP, once you have authenticated clients, since IPs are shared behind NATs and cheap to rotate) and return a Retry-After header so well-behaved clients back off correctly instead of hammering the limit.

Related documents

Teams that route all three through one flag system tend to end up with flags nobody's sure are safe to delete. Each has a different lifecycle and deserves different tooling.

Team size and deploy friction are better predictors of when to split a service than technical elegance. Splitting too early adds distributed-systems cost before the org is big enough to need it.

A dashboard showing a healthy average response time can hide a bad experience for one in a hundred requests. For any system with real concurrency, that's not a rare edge case.