The club that almost closed
Imagine a popular nightclub with room for 200 people. On its first big Saturday, 2,000 people show up at once. Everyone rushes in, the music stops, the bar runs dry, and the fire marshal shuts the place down.
The next weekend, the owner hires a bouncer with a clicker. When the club is full, he says "come back in a few minutes." Those inside enjoy the night, and those outside wait their turn.
That bouncer is a rate limiter. Your API is the club. Without a bouncer, one buggy script, one scraper or one traffic spike can take it down for everyone.
Why rate limit at all?
- Availability: one noisy client shouldn't starve the others.
- Cost: every request uses CPU, database calls and sometimes paid third-party APIs.
- Security: it slows brute-force logins, OTP abuse and credential stuffing.
- Fairness: free and paid users can get different limits.
- Stability: it prevents retry storms (see the restaurant-waiter post).
The bouncer's five rulebooks
1. Fixed window
Count requests per fixed slot, such as one minute, and reset at the boundary. It's simple and cheap, but it has the boundary problem: 100 requests at 12:00:59 and 100 more at 12:01:00 means 200 in two seconds.
2. Sliding window log
Store the timestamp of every request and count those in the last 60 seconds. It's perfectly accurate, but memory-heavy because you store one entry per request.
3. Sliding window counter
Keep counts for the current and previous window and blend them. It's cheap and smooth, though only an approximation.
estimated = previous_count × (1 − fraction_of_current_window_elapsed) + current_count
4. Token bucket
Each client has a bucket of up to N tokens. Each request takes one, and tokens refill at a steady rate. An empty bucket means rejection. It allows short bursts while enforcing a long-term average, which is why many public APIs prefer it. The cost is two settings to tune: bucket size and refill rate.
5. Leaky bucket
Requests enter a queue and leave at a constant rate, like water dripping from a hole. The output is very smooth, which suits protecting fragile downstream systems, but bursts get delayed or dropped.
Quick comparison
| Algorithm | Bursts | Memory | Accuracy | Best for |
|---|---|---|---|---|
| Fixed window | Yes (boundary spike) | Very low | Low | Simple internal limits |
| Sliding log | No | High | Exact | Small scale, strict rules |
| Sliding counter | Slight | Low | Good | General public APIs |
| Token bucket | Yes, controlled | Low | Good | User-facing APIs |
| Leaky bucket | No | Low | Good | Protecting downstream services |
Whom does the bouncer count?
- Per IP: easy, but shared office or mobile NAT addresses punish innocent users.
- Per user or API key: fairest for authenticated APIs.
- Per endpoint:
/loginneeds a much tighter limit than/products. - Global: a last line of defence for the whole service.
In practice, layer them: strict per-IP on login, per-user on the API, and a global cap as a safety net.
Doing it in ASP.NET Core
Since .NET 7, rate limiting is built in with System.Threading.RateLimiting. The built-in limiters are FixedWindow, SlidingWindow, TokenBucket and Concurrency.
builder.Services.AddRateLimiter(options =>
{
options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;
options.AddPolicy("per-user", ctx =>
RateLimitPartition.GetTokenBucketLimiter(
ctx.User.Identity?.Name
?? ctx.Connection.RemoteIpAddress?.ToString()
?? "anonymous",
_ => new TokenBucketRateLimiterOptions
{
TokenLimit = 20, // burst size
TokensPerPeriod = 5, // refill amount
ReplenishmentPeriod = TimeSpan.FromSeconds(1),
QueueLimit = 0, // reject, don't queue
AutoReplenishment = true
}));
options.OnRejected = async (context, token) =>
{
if (context.Lease.TryGetMetadata(MetadataName.RetryAfter, out var ra))
context.HttpContext.Response.Headers.RetryAfter =
((int)ra.TotalSeconds).ToString();
await context.HttpContext.Response.WriteAsync(
"Too many requests. Please slow down.", token);
};
});
app.UseRateLimiter();
app.MapGet("/orders", () => Results.Ok()).RequireRateLimiting("per-user");
The concurrency limiter is different: it caps how many requests run at the same time, not how many arrive per second.
The hard part: many servers
The built-in limiter keeps counters in memory, per instance. With 4 instances behind a load balancer, "100 per minute" can quietly become 400.
The fix is a shared store, usually Redis. The trap is a race condition:
INCR key ← server crashes here EXPIRE key 60 ← never runs, key lives forever
Make both steps atomic with a Lua script, since Redis runs scripts atomically:
local current = redis.call('INCR', KEYS[1])
if current == 1 then
redis.call('EXPIRE', KEYS[1], ARGV[1])
end
return current
Other options are enforcing limits at the edge (API gateway, NGINX, YARP, Cloudflare) or syncing counters periodically and accepting slight over-limit. Decide what happens if Redis is down: fail open (allow traffic) suits most APIs, while login endpoints may deserve fail closed.
Being a polite API
- Return HTTP 429 Too Many Requests, not 500 or 403.
- Include a
Retry-Afterheader. - Optionally expose remaining-request and reset headers (an IETF draft is standardising
RateLimitheaders). - Write a clear error body.
Being a polite client: backoff with jitter
If 1,000 clients get a 429 and all retry after exactly one second, you get a thundering herd. Use exponential backoff with jitter:
wait = random(0, min(cap, base × 2^attempt))
In .NET, Polly handles this well and pairs nicely with a circuit breaker.
Common mistakes
- Limiting only by IP and blocking whole offices.
- Using in-memory limits across multiple instances.
- One limit for every endpoint, when login and search cost very different amounts.
- No
Retry-After, so clients guess and hammer harder. - Forgetting internal callers, so your own services limit each other.
- Never monitoring 429s. A spike may be an attack or a limit that's too tight.
- Treating rate limiting as DDoS defence. Real DDoS protection happens at the network edge.
Wrapping up
A rate limiter isn't there to punish users. It's the bouncer who keeps the club open all night. Choose the algorithm for your traffic (token bucket is a good default), choose your partition key carefully, share state across instances, and always tell clients when to come back.
Next in the series: Circuit Breakers, the electrician who cuts the power before the house burns down.





