Rate Limiting Algorithms: Token Bucket vs Sliding Window Explained
September 30, 2026
5 min read
A practical comparison of the two most common rate limiting algorithms, token bucket and sliding window, and guidance on which one fits a given API.
Ayodele Miracle JohnSOFTWARE ENGINEER
Every public API eventually needs a way to say "not so fast." Without it, one misbehaving client, a retry loop gone wrong, or a scraper can consume enough capacity to slow the service down for everyone else. Rate limiting is the mechanism that enforces a cap on how many requests a client can make in a given period, and the algorithm you pick decides how fair, how bursty, and how expensive that enforcement actually is.
Two algorithms show up in almost every real system: token bucket and sliding window. They solve the same problem but make different tradeoffs, and picking the wrong one for the situation shows up later as either angry customers getting blocked unfairly or a service that never really protects itself.
Why rate limiting exists in the first place
A rate limit protects three things at once: the backend's own capacity, other tenants sharing that capacity, and, less obviously, the calling client itself. A client that's about to exceed a limit is often about to make things worse for itself too, retrying failed requests in a loop, burning through a budget, or hammering a downstream service it doesn't realize is struggling.
Limits are usually expressed as "N requests per period," like 100 requests per minute. The simplest possible implementation, a fixed window counter that resets every 60 seconds, has an obvious flaw: a client can send 100 requests in the last second of one window and another 100 in the first second of the next, producing 200 requests in roughly two seconds even though the stated limit is 100 per minute. Both token bucket and sliding window exist to close that gap.
The token bucket algorithm
A token bucket holds a fixed number of tokens, refilled at a steady rate, say one token every 600 milliseconds for a 100-per-minute limit. Every incoming request consumes one token. If a token is available, the request goes through and the count drops by one. If the bucket is empty, the request is rejected or delayed until the next token arrives.
The useful property here is that the bucket can hold unused tokens up to its capacity, which means a client that's been idle can send a short burst of requests immediately, then settle into the steady refill rate. This matches how real traffic actually behaves: a dashboard that loads ten resources at once on page load, then goes quiet. Token bucket tolerates that burst without needing a special case for it.
The tradeoff is that token bucket is a smoothing model, not a precise historical count. It doesn't answer "how many requests happened in the last 60 seconds," it answers "is there capacity available right now," which is a different (and usually more useful) question for protecting a backend, but a worse fit if the actual requirement is a hard, auditable cap like "this API key gets exactly 10,000 calls a day."
The sliding window algorithm
Sliding window fixes the boundary problem in fixed windows by looking at a rolling period instead of a fixed calendar-aligned one. Instead of asking "how many requests since minute 0," it asks "how many requests happened in the last 60 seconds, starting from right now."
There are two common ways to implement it. A true sliding log keeps a timestamp for every request and counts how many fall inside the trailing window, which is accurate but can get memory-heavy under high request volume. A sliding window counter approximates this cheaply: it keeps two fixed window counters (the current one and the previous one) and computes a weighted count based on how far into the current window the request landed. That weighted approach is what most production systems, including common Redis-based rate limiters, actually implement, because it gets very close to the true sliding log at a fraction of the storage cost.
Sliding window's strength is precision and fairness across the boundary, it closes the fixed-window loophole directly and gives an answer that matches what a person means by "100 requests per minute." Its weakness is that it doesn't naturally allow the same kind of burst tolerance that token bucket gives for free. A client that's been well under its limit for an hour gets no credit for that when a burst arrives.
Comparing the two in practice
Token bucket tends to fit situations where short bursts are normal and desirable, API gateways protecting infrastructure, client SDKs pacing their own outbound calls, or anywhere the goal is smoothing traffic rather than producing an exact count. Sliding window tends to fit situations where the limit is a contractual or billing-relevant number, like a per-plan API quota that needs to be enforceable and explainable to a customer who asks why they got a 429.
Neither algorithm is inherently "more correct." A payments API and an internal caching layer within the same company might reasonably choose different ones for the same nominal limit, because they're actually solving different problems: one is protecting shared infrastructure from spikes, the other is enforcing a plan boundary a customer agreed to.
Where enforcement actually lives
Rate limiting can sit in a few different places, and where it sits changes what's practical. At the API gateway or edge (a service like Cloudflare, or a reverse proxy), it protects the backend before a request even reaches application code, which is the right place for coarse, infrastructure-protecting limits. Inside the application, backed by a shared store like Redis, it can enforce per-user or per-API-key limits with either algorithm, since Redis makes atomic increment-and-check operations cheap and the same counters are visible across every application instance. Purely in-memory, per-process limiting is the least reliable option in any system running more than one instance, because each instance ends up enforcing its own separate limit rather than a shared one.
Whichever layer enforces the limit, the response matters as much as the algorithm. A 429 status code with a Retry-After header, or the increasingly standard RateLimit response headers, tells a well-behaved client exactly when to try again instead of forcing it to guess and retry blindly, which is itself a good way to accidentally cause the exact overload the limit was meant to prevent.
Choosing between them
A reasonable default: use token bucket when the goal is protecting a system from spiky traffic and bursts are actually fine, and use sliding window when the limit needs to be precise, explainable, and tied to something like a pricing plan or contractual quota. Plenty of systems end up combining the two: a fast token bucket at the edge to absorb spikes, and a sliding window count further in to enforce the account-level quota that shows up on an invoice.

Sep 25, 2026
Why Your App Shows Stale Data Right After a Write: Understanding Read Replica Lag
Read replicas help databases scale, but they introduce a subtle bug where users don't see their own recent changes. Here's why that happens and how teams actually fix it.

Sep 21, 2026
4 min read
Why Serverless Functions Keep Exhausting Your Postgres Connections
Serverless functions open far more Postgres connections than a traditional server ever would, and raising max_connections only delays the failure. Here's why connection pooling is the actual fix, and what it costs you.

Sep 20, 2026
5 min read
Idempotency Keys: Making Payment APIs Safe to Retry
How idempotency keys stop a flaky network from turning one API request into two duplicate charges, and the practical details of implementing them correctly.