Rate Limiting Strategies Explained: Algorithms, Redis, and Distributed Enforcement
Learn how rate limiting works with fixed windows, sliding windows, token buckets, leaky buckets, Redis, concurrency limits, and distributed systems.
33 min read
One API Usually Needs More Than One Limit
Imagine we have a Reports API:
GET /projects
POST /reports
POST /login
POST /webhooks/testThese endpoints do very different work.
For example:
GET /projects
→ simple database readPOST /reports
→ expensive CPU + database workPOST /login
→ security-sensitive authenticationPOST /webhooks/test
→ calls an external providerNow imagine we have these rules:
Free tenant
→ 60 normal API requests/minute
Paid tenant
→ 5,000 normal API requests/minute
Report generation
→ maximum 2 active reports per tenant
Login
→ maximum 5 failed attempts per 15 minutes
Webhook provider
→ maximum 20 outgoing calls/secondThese are not the same kind of limit.
One Redis counter cannot correctly protect all of them.
The important question is not:
Which Redis command should I use?The first question should be:
What resource am I protecting, from whom, and how quickly can that resource be consumed?
That question determines the correct limiting strategy.
Rate Limiting Is Not the Same as Everything Else
Several terms are commonly mixed together:
rate limiting
quota
concurrency limiting
throttling
backpressure
load sheddingThey are related, but they solve different problems.
Rate Limit
A rate limit answers:
How many operations can happen during a period of time?
Example:
60 requests/minuteor:
10 requests/secondFor example:
Tenant A
12:00–12:01
→ maximum 60 API requestsThis protects the system from too much traffic over time.
Quota
A quota usually covers a longer product period.
For example:
10,000 reports/monthThis is more about:
product usage
billing
plan limitsthan immediate server protection.
For example:
Free plan
→ 10 reports/month
Pro plan
→ 1,000 reports/monthA tenant may stay below its rate limit but eventually hit its monthly quota.
Concurrency Limit
A concurrency limit asks:
How many operations can be running at the same time?
Example:
maximum 2 active reportsSuppose:
Report A
→ running
Report B
→ running
Report C
→ arrivesReport C may be:
rejected
or
queuedeven if the tenant has made only three requests this hour.
Why?
Because the problem is not request rate.
The problem is:
too much work running simultaneouslyThrottling
Throttling means intentionally slowing or shaping accepted work.
Example:
100 webhook requests arrive immediatelybut the external provider allows only:
20 requests/secondWe can queue them and drain the queue at:
20/secondThat is throttling.
Backpressure
Backpressure happens when a slow consumer tells producers:
Slow down.Example:
API
→ sends jobs
Workers
→ process jobsIf workers are falling behind:
queue depth risesThe system may:
pause intake
reduce producer speed
reject new workThat is backpressure.
Load Shedding
Load shedding means rejecting lower-priority work when the system is overloaded.
For example:
System CPU = 95%
Database connections nearly exhaustedInstead of allowing everything to fail, we may reject:
optional analytics exportswhile preserving capacity for:
login
payments
critical readsEasy Comparison
| Mechanism | Main question |
|---|---|
| Rate limit | How many operations over time? |
| Quota | How much usage over a longer period? |
| Concurrency | How many operations are active now? |
| Throttling | How fast should accepted work be processed? |
| Backpressure | How should producers slow down? |
| Load shedding | What should be rejected when capacity is exhausted? |
Production systems often use several at once.
For example:
Request
↓
IP rate limit
↓
Tenant rate limit
↓
Monthly quota
↓
Concurrency check
↓
Actual workDefine the Policy Before Choosing an Algorithm
Before implementing Redis logic, define the rule clearly.
For example:
type RateLimitPolicy = {
name: string;
identity: "ip" | "user" | "tenant" | "api-key" | "global";
route: string;
capacity: number;
refillPerSecond: number;
cost: number;
failureMode: "open" | "closed" | "local-fallback";
};This policy should answer:
Who are we limiting?
What route or operation is limited?
How much traffic is allowed?
Are bursts allowed?
How expensive is one request?
What happens when Redis fails?Do not choose:
100 requests/minutejust because it looks reasonable.
A good limit should come from:
database capacity
CPU capacity
provider limits
product tiers
security requirements
load testing
traffic patternsChoosing the Identity
A rate limiter needs to know:
Who owns this traffic?Possible identities include:
IP address
user ID
tenant ID
API key
global serviceFor example:
rl:tenant:acme:projectsor:
rl:user:user-42:loginor:
rl:provider:webhook-vendorPrefer Trusted Server-Side Identity
For authenticated APIs:
tenant ID
user ID
API key IDshould come from trusted authentication state.
For example:
request.auth.tenantId;not:
request.query.tenantId;Otherwise a client could try:
tenantId=someone-elseand bypass its own limit.
Avoid Raw URLs in Redis Keys
Suppose requests look like:
/projects/123
/projects/456
/projects/789Do not create:
rl:/projects/123
rl:/projects/456
rl:/projects/789for every unique ID.
Use the route template:
/projects/:idinstead.
Example:
rl:tenant:acme:route:projects-by-idThis prevents uncontrolled key growth.
Be Careful With IP-Based Limits
IP rate limits are useful, especially before authentication.
For example:
100 login requests/minute/IPBut IP addresses are imperfect identities.
Several legitimate users may share one address because of:
office NAT
mobile networks
VPNs
university networksOn the other hand, attackers can distribute traffic across many IPs.
So IP limiting should normally be:
one defensive layernot the only limiter for authenticated application usage.
Never Blindly Trust X-Forwarded-For
Suppose a client sends:
X-Forwarded-For: 1.2.3.4If your application blindly trusts this header, the attacker can simply change it on every request.
Only trust forwarded IP headers when the request comes through a trusted proxy such as:
Nginx
Cloudflare
AWS ALB
API Gatewayand the proxy is configured to sanitize client-provided forwarding headers.
Fixed Window Rate Limiting
The simplest algorithm is the fixed window counter.
Suppose the policy is:
100 requests/minuteTime is divided into windows:
12:00:00 → 12:00:59
12:01:00 → 12:01:59
12:02:00 → 12:02:59Each window gets its own counter.
For example:
rl:tenant:acme:projects:29859840where the final value identifies the current time window.
Fixed Window Example
Suppose the limit is:
100 requests/minuteA client sends:
100 requests at 12:00:59All are allowed.
One second later:
new minute startsThe client sends:
100 requests at 12:01:00They are also allowed.
So the server receives:
200 requestswithin roughly:
1 secondeven though the configured limit is:
100/minuteThis is called the:
window boundary problemWhy Fixed Window Is Still Useful
Fixed window is:
simple
cheap
easy to understand
low memoryIt can be perfectly good for:
coarse quotas
admin APIs
low-risk endpointswhere short bursts around the boundary are acceptable.
Redis Implementation Problem
A common implementation is:
const count = await redis.incr(key);
if (count === 1) {
await redis.expire(key, 60);
}This looks fine.
But imagine:
INCR succeeds
↓
process crashes
↓
EXPIRE never runsNow the Redis key may never expire.
The user can remain permanently limited.
The operation:
increment
+
set expirationneeds to be atomic.
A Lua script is one common solution.
Sliding Window Log
Fixed windows have rough boundaries.
Sliding window logs solve that problem by storing each accepted request timestamp.
Suppose the policy is:
5 attempts / 15 minutesAt:
12:20we care about requests since:
12:05not:
12:15or some fixed clock boundary.
This creates a true rolling window.
Redis Sorted Set Approach
A Redis sorted set works well.
Each request is stored as:
score
→ timestamp
member
→ unique request identifierThe logic is:
1. Remove requests older than the window
2. Count remaining requests
3. If count >= limit
reject
4. Otherwise add current request
5. Refresh TTLCommon Redis commands:
ZREMRANGEBYSCORE
ZCARD
ZADD
PEXPIREThese steps should be performed atomically.
Why the Member Must Be Unique
Suppose two requests arrive in the same millisecond.
If we use only:
timestampas the sorted-set member, one request may overwrite the other.
Instead use something like:
timestamp + requestIdor another unique identifier.
Sliding Log Trade-Off
Sliding logs are accurate.
But they store:
one entry per accepted requestFor login attempts:
5 attempts / 15 minutesthat is tiny.
For:
100,000 API requests/minutethat becomes expensive.
So sliding logs are a strong fit for:
login attempts
password reset limits
security-sensitive operationsbut often too expensive for every request on a high-throughput public API.
Sliding Window Counter
The sliding counter tries to get much of the fairness of a sliding window without storing every request.
It usually keeps:
previous window count
current window countand weights the previous one.
Sliding Counter Example
Suppose the limit is:
100 requests/minuteThe current minute is:
25% completePrevious window:
80 requestsCurrent window:
20 requestsBecause 75% of the previous minute still overlaps our rolling 60-second window:
previous weight = 0.75Estimated request count:
80 × 0.75 + 20which gives:
60 + 20
= 80If the current window reaches:
40then:
80 × 0.75 + 40
= 100The next request gets rejected.
Why Use Sliding Counters?
They use almost constant storage:
two counters per identityinstead of:
one entry per requestThey also reduce the boundary burst problem of fixed windows.
Trade-off:
they are approximatebecause they do not know exactly when every request occurred inside each bucket.
This makes them useful for:
general API limitswhere exact per-request history is unnecessary.
Token Bucket
Token bucket is one of the most useful algorithms for public APIs.
It allows:
controlled bursts
+
steady long-term trafficImagine a bucket containing tokens.
Each request consumes tokens.
Tokens refill over time.
Token Bucket Example
Suppose:
capacity = 20 tokens
refill rate = 1 token/second
normal request cost = 1 tokenIf a tenant has been idle for a while:
bucket = 20 tokensThey can immediately send:
20 requestsThat burst is allowed.
After the bucket is empty, it refills at:
1 request/secondSo sustained traffic becomes roughly:
1 request/secondWhy Token Bucket Is Useful
It lets us separately define:
maximum burstand:
long-term rateFor example:
capacity = 20
refill = 1 token/secondmeans:
burst up to 20but long-term:
~60 requests/minuteThis is often a better user experience than rejecting every small burst.
Token Refill Example
Suppose:
capacity = 20
refill = 1 token/secondCurrent bucket:
4 tokensThe last update was:
3.5 seconds agoNew tokens:
3.5 × 1
= 3.5Available:
4 + 3.5
= 7.5Maximum is still:
20so:
available = 7.5If the next request costs:
5 tokensit is accepted.
Remaining:
2.5 tokensWeighted Request Costs
Not every request costs the server the same amount.
For example:
GET /projects
→ small indexed read
GET /projects?search=...
→ search query
POST /reports/preview
→ expensive report generationInstead of counting every request as one:
GET /projects
cost = 1
GET /projects?search=...
cost = 2
POST /reports/preview
cost = 10Now the limiter better reflects actual resource consumption.
Keep Costs Simple
Do not create a pricing system where every query has a dynamically calculated cost based on:
query plan
row count
CPU milliseconds
cache hit rateClients will not understand the contract.
Prefer:
small stable categoriessuch as:
normal = 1
heavy = 5
very-heavy = 10Atomic Token Bucket With Redis
When many application instances exist:
Instance A
Instance B
Instance Call of them must update the same bucket safely.
This cannot happen with separate:
GET
calculate
SETcalls because requests can race.
We need one atomic operation.
Lua scripts are a common solution.
Redis Token Bucket Script
local key = KEYS[1]
local capacity = tonumber(ARGV[1])
local refill_per_second = tonumber(ARGV[2])
local cost = tonumber(ARGV[3])
local ttl_ms = tonumber(ARGV[4])
local redis_time = redis.call("TIME")
local now_ms =
redis_time[1] * 1000 +
math.floor(redis_time[2] / 1000)
local bucket =
redis.call(
"HMGET",
key,
"tokens",
"updated_at_ms"
)
local tokens =
tonumber(bucket[1])
local updated_at_ms =
tonumber(bucket[2])
if tokens == nil then
tokens = capacity
updated_at_ms = now_ms
end
if now_ms < updated_at_ms then
updated_at_ms = now_ms
end
local elapsed_ms =
now_ms - updated_at_ms
local refill_per_ms =
refill_per_second / 1000
tokens = math.min(
capacity,
tokens +
elapsed_ms * refill_per_ms
)
local allowed = 0
local retry_after_ms = 0
if tokens >= cost then
tokens = tokens - cost
allowed = 1
else
retry_after_ms =
math.ceil(
(cost - tokens) /
refill_per_ms
)
end
redis.call(
"HSET",
key,
"tokens",
tostring(tokens),
"updated_at_ms",
tostring(now_ms)
)
redis.call(
"PEXPIRE",
key,
ttl_ms
)
return {
allowed,
math.floor(tokens),
retry_after_ms,
now_ms
}This script performs:
read bucket
↓
calculate refill
↓
check capacity
↓
consume tokens
↓
save new state
↓
refresh TTLas one atomic Redis operation.
Why Use Redis Server Time?
Imagine:
Server A clock
→ 12:00:01
Server B clock
→ 12:00:04If every server calculates token refill using its own clock, the bucket may refill inconsistently.
Using:
redis.call("TIME")gives all application instances the same time source for that Redis limiter.
Validate the Policy Before Redis
Do not let invalid configuration reach the Lua script.
For example:
capacity = -10or:
refill rate = 0or:
cost = 100
capacity = 20These should be rejected earlier.
Token Bucket TTL
Rate-limit keys should expire when inactive.
Otherwise Redis slowly fills with old buckets.
Suppose:
capacity = 20
refill = 1 token/secAn empty bucket takes:
20 secondsto become full again.
After enough idle time, the old stored state is no longer useful.
We can choose a TTL longer than refill-to-full time.
Example:
function bucketTtlMs(capacity: number, refillPerSecond: number) {
const refillToFullMs = Math.ceil(capacity / refillPerSecond) * 1000;
return Math.max(60_000, refillToFullMs * 2);
}So Redis automatically removes inactive buckets.
TypeScript Token Bucket Wrapper
The rest of our application should not need to understand Lua details.
Define a simple result:
type TokenBucketPolicy = {
capacity: number;
refillPerSecond: number;
};
type RateLimitDecision = {
allowed: boolean;
limit: number;
remaining: number;
retryAfterMs: number;
};Then:
async function consumeTokenBucket(input: {
redis: RedisClient;
key: string;
policy: TokenBucketPolicy;
cost?: number;
}): Promise<RateLimitDecision> {
const cost = input.cost ?? 1;
const { capacity, refillPerSecond } = input.policy;
if (!Number.isFinite(capacity) || capacity <= 0) {
throw new Error("Token bucket capacity must be positive");
}
if (!Number.isFinite(refillPerSecond) || refillPerSecond <= 0) {
throw new Error("Token bucket refill rate must be positive");
}
if (!Number.isFinite(cost) || cost <= 0 || cost > capacity) {
throw new Error("Token cost must be between zero and bucket capacity");
}
const raw = await input.redis.eval(TOKEN_BUCKET_LUA, {
keys: [input.key],
arguments: [
String(capacity),
String(refillPerSecond),
String(cost),
String(bucketTtlMs(capacity, refillPerSecond)),
],
});
const [allowed, remaining, retryAfterMs] = raw as [number, number, number];
return {
allowed: allowed === 1,
limit: capacity,
remaining,
retryAfterMs,
};
}Now application code gets a simple decision:
allowed?
remaining?
retry after?Different Product Tiers
We can define:
const tierPolicies = {
free: {
capacity: 20,
refillPerSecond: 1,
},
paid: {
capacity: 100,
refillPerSecond: 10,
},
} satisfies Record<string, TokenBucketPolicy>;Then:
const decision = await consumeTokenBucket({
redis,
key: `rl:tenant:${tenant.id}:route:projects`,
policy: tierPolicies[tenant.tier],
cost: requestIsHeavySearch ? 2 : 1,
});Important:
tenant.id
tenant.tiermust come from authenticated application state.
Do not trust a request containing:
?tier=paidWhat Should the API Return When Limited?
Use:
429 Too Many RequestsDo not return:
500 Internal Server Errorbecause rate limiting is intentional behavior.
The client needs to know:
I am limitednot:
the server unexpectedly failedUseful 429 Response
if (!decision.allowed) {
const retryAfterSeconds = Math.max(
1,
Math.ceil(decision.retryAfterMs / 1000),
);
return Response.json(
{
error: {
code: "rate_limit_exceeded",
message: "Too many requests",
retryAfterSeconds,
requestId,
},
},
{
status: 429,
headers: {
"Retry-After": String(retryAfterSeconds),
"RateLimit-Limit": String(decision.limit),
"RateLimit-Remaining": String(decision.remaining),
},
},
);
}Retry-After Matters
Suppose we just return:
429 Too Many RequestsThe client does not know when to retry.
It may retry immediately:
request
→ 429
→ retry
→ 429
→ retry
→ 429This makes the problem worse.
Instead return:
Retry-After: 5meaning:
retry after approximately 5 secondsClients can then back off properly.
Leaky Bucket
Token bucket allows controlled bursts.
Leaky bucket focuses on smoothing output.
Imagine:
100 webhook test requests
arrive instantlybut the provider supports:
20 requests/secondWe can create:
Incoming requests
↓
bounded queue
↓
worker
↓
20 provider calls/secondThe burst enters quickly.
Output leaves smoothly.
That is the basic idea behind a leaky bucket.
Token Bucket vs Leaky Bucket
Token bucket:
Allows burst
then enforces average rateLeaky bucket:
Smooths burst into steady outputExample:
Token bucket
20 requests
→ may all execute immediatelywhile:
Leaky bucket
20 requests
→ execute steadily over timeBe Careful Queuing Synchronous HTTP Requests
Suppose an HTTP request has:
client timeout = 3 secondsbut the queue makes it wait:
8 secondsThe work may eventually execute, but the client is already gone.
That creates wasted work.
For asynchronous tasks, a better design may be:
POST /reports
↓
accept job
↓
202 Accepted
↓
queue
↓
worker processes laterFor synchronous APIs, it may be safer to:
reject immediatelywhen capacity is unavailable.
GCRA
Another algorithm you may encounter is:
GCRA
Generic Cell Rate AlgorithmInstead of storing request timestamps or tokens, GCRA tracks:
theoretical next allowed arrival timeConceptually:
allowed rate
↓
expected spacing between requests
↓
track theoretical arrival timeIf a request arrives too far ahead of schedule:
rejectGCRA provides behavior similar to token bucket while using constant state.
It is useful for precise limits.
But it is harder to explain and operate than token bucket.
For many APIs:
token bucket
or
sliding counteris easier to start with.
Concurrency Limits Are Different
Imagine report generation takes:
2 minutesA tenant sends:
2 requests/minuteThat sounds low.
But after a few minutes, many report jobs could still be running.
So a normal rate limit may not protect:
CPU
database connections
worker capacityWe need a concurrency limit.
Concurrency Example
Policy:
maximum active reports per tenant = 2State:
Report A
→ running
Report B
→ running
Report C
→ arrivesReport C must:
wait
or
be rejectedregardless of the per-minute request count.
Local Semaphore
If the application runs on one server:
one Node.js processwe can use an in-memory semaphore.
But imagine:
Server A
Server B
Server CEach server thinks:
2 reports allowedTotal could become:
6 running reportsSo distributed concurrency needs shared coordination.
Distributed Concurrency Limit
A robust distributed implementation often needs:
unique lease ID
atomic acquisition
expiration
release
lease renewalConceptually:
worker acquires lease
↓
does work
↓
releases leaseIf the worker crashes:
lease expiresso capacity is eventually recovered.
Why Expiry Is Not Perfect
Suppose:
lease TTL = 60 secondsbut a report takes:
90 secondsIf renewal fails, the lease can expire while the original worker is still running.
Another worker may then acquire the slot.
For a short period:
actual concurrency > configured concurrencyThis is why critical systems may need:
lease renewal
fencing tokens
durable job ownershipnot just an expiring counter.
Hierarchical Rate Limits
Real systems often apply several limits to one request.
For example:
Incoming request
↓
IP abuse limit
↓
Global service limit
↓
Tenant plan limit
↓
User limit
↓
Endpoint limit
↓
Concurrency limitEach layer protects something different.
Example
Suppose:
IP limit
→ 500 requests/minuteprotects against basic abuse.
Then:
tenant limit
→ 5,000/minuteenforces the paid plan.
Then:
report concurrency
→ 2 active reportsprotects workers.
One limiter cannot replace all three.
Multi-Bucket Atomicity Is Difficult
Suppose one request must consume from:
global bucket
tenant bucket
user bucketA simple implementation may do:
consume global
↓
consume tenant
↓
consume userBut imagine:
global accepted
tenant accepted
user rejectedThe earlier buckets already lost tokens.
That creates conservative accounting.
One Lua script could potentially update all buckets atomically.
But Redis Cluster complicates this because multi-key scripts usually require keys to be in the same hash slot.
So there are trade-offs between:
atomicity
key distribution
hot spots
complexityOften it is cleaner to enforce limits at different layers.
Example Layered Design
CDN / Edge
→ coarse IP abuse protection
Application
→ tenant rate limits
Worker system
→ report concurrency
Integration service
→ provider quotaThis keeps each policy close to the resource it protects.
Redis Key Cardinality
Every new limiter identity can create Redis state.
Imagine an attacker sends millions of fake usernames:
random-user-1
random-user-2
random-user-3
...If each creates:
rl:login:random-user-XRedis memory can grow dramatically.
This is a cardinality problem.
Control Redis Key Growth
Use strategies such as:
validate identities
authenticate before creating tenant keys
normalize IP addresses
use route templates
set TTL on every key
avoid unnecessary key dimensions
monitor key countFor example, bad:
rl:user:123:path:/projects/92271?search=helloBetter:
rl:user:123:route:projects-searchTTL Is Part of Capacity Management
TTL is not only cleanup.
It controls how much limiter state can remain in Redis.
Different algorithms need different TTL rules.
Examples:
Fixed window
→ slightly longer than the window
Sliding log
→ at least as long as the rolling window
Token bucket
→ long enough for the bucket to become full
Concurrency lease
→ based on expected operation durationHot Keys
Suppose every webhook request uses:
rl:provider:webhook-vendorNow every application instance hits one Redis key.
At very high throughput:
one Redis shardmay become a bottleneck.
The limiter itself can become the system bottleneck.
This is called a:
hot keyPossible Hot-Key Solutions
Depending on how strict the limit must be:
enforce coarse limits at the gateway
allocate budgets per region
allocate token leases to workers
batch calls
use local sub-limitsFor example:
Global provider limit:
20,000 requests/sec
Region A:
12,000/sec
Region B:
7,000/sec
Reserve:
1,000/secNow not every request needs to hit one global key.
Exact Global Limits vs Approximate Limits
If the rate limit protects:
billingor:
strict external provider quotayou may need strong coordination.
If the limit protects only:
general overloadthen small temporary overages may be acceptable.
This distinction matters a lot in distributed systems.
What Happens When Redis Fails?
If Redis is on the request path, Redis can fail.
Examples:
timeout
network partition
Redis restart
high latency
cluster failureYou must decide what the application does.
There are three common approaches.
Fail Open
If Redis is unavailable:
allow requestThis may fit:
low-risk public reads
fairness limits
non-critical throttlingBenefit:
Redis outage does not become full API outageRisk:
protected backend receives unlimited trafficFail Closed
If Redis is unavailable:
reject requestPossible fit:
login abuse protection
strict provider quota
billing-sensitive operation
critical admin operationBenefit:
policy remains protectedRisk:
Redis outage becomes operation outageLocal Fallback Limiter
Another option:
Redis unavailable
↓
use in-memory limiter temporarilyFor example each Node.js instance gets a small local token bucket.
This provides some protection.
But:
Instance A
→ 20 local tokens
Instance B
→ 20 local tokens
Instance C
→ 20 local tokensTotal effective burst becomes:
60So local fallback is approximate.
Still, it can be useful during a short Redis outage.
Failure Policy Should Depend on the Endpoint
You do not need one global rule.
For example:
GET /projects
→ fail openwhile:
POST /login
→ fail closedand:
POST /reports
→ local fallbackDifferent resources have different risk.
Use Short Redis Timeouts
Do not let every request wait:
30 secondsfor Redis.
Rate limiting is part of the request path.
Use a short deadline.
For example:
Redis limiter timeout
→ 20–100 msdepending on architecture and environment.
Then apply the configured fallback.
Do not retry forever inside the request.
Multiple Application Instances
Suppose we deploy:
10 Node.js instancesIf each instance has its own in-memory limiter:
100 requests/minute per instancethen total system capacity becomes:
1,000 requests/minutedepending on load balancing.
For a shared tenant policy, instances need shared state.
Redis is commonly used because all instances see:
one logical bucketMultiple Regions Are Harder
Now imagine:
Region A
→ Karachi / Mumbai region
Region B
→ Europe
Region C
→ USWe want one strict global limit:
5,000 requests/minuteEvery request would need globally coordinated state.
Cross-region coordination adds:
latency
availability problems
network dependencySo strict global limiting has a cost.
Regional Budget Allocation
Instead, divide the global allowance.
Example:
Global allowance:
5,000/minute
Region A:
3,000
Region B:
1,500
Reserve:
500Each region can enforce its own limit locally.
Trade-off:
one region may have spare capacity
while another region is already fullBudgets can be rebalanced over time.
Bounded Overage
Another approach:
each region has local limitand we accept that total global usage may exceed the target slightly.
Example:
Target global limit
→ 5,000/minute
Possible temporary actual usage
→ 5,500/minuteIf this is only overload protection, that may be acceptable.
For strict billing or provider contracts, it may not be.
Centralize Only Strict Limits
A practical design may be:
normal API traffic
→ regional limit
login abuse
→ regional security limit
provider billing quota
→ centralized strict coordinatorOnly operations requiring global precision pay the cross-region coordination cost.
Observability
A rate limiter should be observable like any other production dependency.
Useful metrics include:
allowed requests
rejected requests
429 rate
Redis decision latency
Redis timeout rate
fallback usage
remaining tokens
concurrency saturation
queue age
limiter key count
Redis memory
hot-key behaviorUseful Dimensions
Good metric labels:
policy
route template
tier
resultExamples:
policy=tenant-projects
tier=free
result=rejectedAvoid labels such as:
tenantId
userId
IP
API key
requestIdbecause they can create millions of metric series.
Use logs or traces for exact identity when needed.
Record the Policy Version
Suppose someone changes:
Free tier
20 tokens
→ 10 tokensand suddenly:
429 rate doublesIt helps if metrics or logs include something like:
policyVersion=v3Then incidents can be correlated with configuration changes.
Testing Rate Limiters
Do not test only application wrapper code.
Redis atomicity should be tested against:
real Redisor a compatible Redis test container.
A mocked Redis client cannot prove that concurrent updates are safe.
Important Token Bucket Tests
Test:
1. New bucket starts full.
2. Requests consume tokens.
3. Weighted requests consume correct cost.
4. Last available token succeeds.
5. Next request fails.
6. Retry time is reasonable.
7. Tokens refill over time.
8. Tokens never exceed capacity.
9. TTL is set.
10. TTL refreshes.
11. Invalid policies are rejected.
12. Concurrent requests do not exceed burst.
13. Multiple app instances share state.
14. Redis timeout triggers correct fallback.
15. Metrics avoid high-cardinality labels.Concurrency Test Example
Suppose:
capacity = 20Launch:
100 requests concurrentlyAt most:
20should consume the initial burst when there has been no refill.
If:
35are accepted because requests raced, the implementation is not atomic.
Load Test the Limiter
A limiter should be tested under:
bursts
sustained traffic
Redis latency
Redis outage
application autoscaling
recoveryA limiter that works only under:
average trafficdoes not protect production.
The dangerous moments are usually bursts and failures.
Algorithm Comparison
| Strategy | State | Burst behavior | Accuracy | Good use |
|---|---|---|---|---|
| Fixed window | One counter/window | Large boundary burst | Exact per window | Simple quotas |
| Sliding log | One entry/request | Strict rolling limit | Very high | Login/security |
| Sliding counter | Two counters | Smaller boundary burst | Approximate | General APIs |
| Token bucket | Tokens + timestamp | Controlled burst | High | Public APIs |
| Leaky bucket | Queue/schedule | Smooth output | High | Provider throttling |
| GCRA | One theoretical time | Controlled | High | Precise limits |
| Concurrency limit | Active leases | Limits in-flight work | Different problem | Heavy jobs |
When to Use Fixed Window
Use it when:
simplicity matters
boundary bursts are acceptable
traffic is modestExample:
Free plan:
1,000 API calls/dayWhen to Use Sliding Log
Use it when:
strict rolling behavior matters
traffic per identity is lowExample:
5 failed logins / 15 minutesWhen to Use Sliding Counter
Use it when:
you want smoother rate enforcement
without storing every requestExample:
1,000 API calls/minutefor ordinary application traffic.
When to Use Token Bucket
Use it when:
bursts are acceptable
but long-term rate must be controlledExample:
API clients
mobile apps
tenant plans
weighted endpointsThis is one of the best general-purpose algorithms.
When to Use Leaky Bucket
Use it when:
output rate must stay smoothExample:
third-party provider
allows 20 calls/secondUse a queue and drain at the required rate.
When to Use Concurrency Limits
Use them when:
each operation remains active
for a long timeand consumes scarce resources.
Examples:
AI inference
report generation
video processing
database-heavy exports
large file conversionA Practical Reports API Design
For our original API, we might use:
GET /projects
→ token bucket per tenantbecause normal API traffic benefits from burst allowance.
POST /reports
→ tenant token bucket
+
2-report concurrency limitbecause reports consume expensive resources.
POST /login
→ sliding log per account
+
coarse IP limiterbecause security needs a stricter rolling window.
POST /webhooks/test
→ tenant request limit
+
global provider leaky bucketbecause the third-party provider has its own outgoing quota.
There is no single best rate-limiting algorithm.
The correct algorithm depends on the resource being protected.
A Useful Request Flow
A production request may look like:
Incoming request
↓
Edge IP protection
↓
Authenticate
↓
Resolve:
user
tenant
tier
↓
Tenant rate-limit policy
↓
Endpoint-specific policy
↓
Concurrency check
if required
↓
Perform work
↓
Release concurrency leaseThat is much closer to real production architecture than:
redis.incr(ip)in front of every endpoint.
Common Mistake: One Global Limit
Bad:
Every request
→ 100 requests/minuteProblems:
cheap GET request
and
expensive reportare treated equally.
Authenticated customers and attackers are treated equally.
Free and paid tenants are treated equally.
One number rarely protects every resource correctly.
Common Mistake: Rate Limiting Only by IP
Bad:
Paid tenant quota
→ IP addressA company with 500 employees behind one NAT may appear as:
one IPwhile one attacker with a botnet may use:
10,000 IPsUse authenticated tenant/user identity for product quotas.
Use IP as an additional abuse signal.
Common Mistake: Non-Atomic Redis Operations
Bad:
GET tokens
↓
calculate locally
↓
SET tokensTwo servers can read the same token count and both accept traffic.
Use:
Lua
Redis transaction
or another atomic primitivewhen admission decisions share state.
Common Mistake: No TTL
If every tenant or IP creates a limiter key forever:
Redis memory
→ keeps growingInactive limiter state should expire.
Common Mistake: No Failure Policy
If Redis fails and nobody decided what to do:
developers improvise
during the outageDefine beforehand:
fail open
fail closed
local fallbackper policy.
Common Mistake: Treating Rate Limiting as Security by Itself
A login rate limit helps.
But it is not a replacement for:
secure passwords
MFA
account lockout policy
credential breach detection
bot protection
IP reputation
session securityRate limiting is one security control among many.
Common Mistake: Ignoring Redis Capacity
A limiter can become extremely expensive when there are:
millions of identities
many dimensions
high request volume
global hot keysMonitor:
Redis CPU
memory
latency
key count
commands/sec
hot shardsA protection mechanism should not become the new bottleneck.
Simple Mental Model
When designing a limiter, ask five questions.
1. WHAT are we protecting?Example:
database
CPU
login endpoint
external provider
product planThen:
2. WHO owns the traffic?Example:
IP
user
tenant
API key
global serviceThen:
3. WHAT behavior do we want?Example:
allow bursts
strict rolling limit
smooth output
limit simultaneous workThen:
4. WHERE should enforcement happen?Example:
edge
API server
worker
integration serviceFinally:
5. WHAT happens if the limiter fails?Example:
allow
reject
use local fallbackOnly after answering these questions should you choose the Redis algorithm.
Production Checklist
Before shipping a rate limiter:
□ Define the protected resource.
□ Decide whether this is:
rate
quota
concurrency
throttling
backpressure
or load shedding.
□ Choose a trusted identity.
□ Use authenticated tenant/user
identity for product limits.
□ Use route templates in keys.
□ Normalize IP/network identities.
□ Choose the algorithm based on
burst and accuracy needs.
□ Make shared-state decisions atomic.
□ Use Redis server time
when appropriate.
□ Validate policies before Redis.
□ Add TTL to every limiter key.
□ Bound Redis key cardinality.
□ Support weighted costs
when operations differ greatly.
□ Return HTTP 429
for intentional rate rejection.
□ Return Retry-After.
□ Define fail-open,
fail-closed,
or fallback behavior.
□ Add short Redis deadlines.
□ Add concurrency limits
for long-running work.
□ Think about hot keys.
□ Plan for multiple instances.
□ Plan for multiple regions.
□ Keep metric labels bounded.
□ Monitor limiter latency.
□ Monitor 429 rate.
□ Monitor Redis failures.
□ Test real Redis atomicity.
□ Test bursts.
□ Test concurrent callers.
□ Test Redis outages.
□ Test recovery.
□ Revisit limits when traffic
or infrastructure changes.Conclusion
Rate limiting is not:
one Redis counter
+
one number
+
every endpointDifferent problems require different tools.
Fixed window
→ simple coarse limits
Sliding log
→ strict rolling security limits
Sliding counter
→ efficient rolling approximation
Token bucket
→ controlled bursts + long-term rate
Leaky bucket
→ smooth output
GCRA
→ precise constant-state limiting
Concurrency limit
→ protect work already runningThe most important part is not memorizing every algorithm.
It is understanding what each one protects.
For example:
Normal API traffic
→ rate limit
Monthly product usage
→ quota
Expensive reports
→ concurrency limit
Third-party provider
→ throttling
Overloaded workers
→ backpressure
System at maximum capacity
→ load sheddingThen, in a distributed system:
Application instances
↓
shared Redis state
↓
atomic decision
↓
allow or rejectmust be designed carefully.
A production-grade limiter also needs:
trusted identity
TTL
atomicity
failure policy
429 responses
Retry-After
metrics
testing
multi-instance behavior
multi-region strategyThe final rule is simple:
Do not start with Redis. Start with the resource you are trying to protect.
Once that is clear, choosing the correct rate-limiting algorithm becomes much easier.
Rate limiting is only one overload control. The
backpressure guide explains when
to bound concurrency, queue work, return 429, return 503, or shed optional
work across the whole request path.
References
Related
Written by
Faisal
Software engineer writing about backend systems, Node.js, system design, scalable applications, and modern web and mobile development.