Skip to content

Rate Limiting & Throttling

Rate limiting controls how many requests a client can make in a time window. It protects your system from abuse, accidental spikes, and DDoS attacks.


flowchart TB
subgraph Bucket["Token Bucket"]
T1["🪙 Token"]
T2["🪙 Token"]
T3["🪙 Token"]
Empty["⬜ Empty slot"]
end
Add["⏰ Tokens added at steady rate<br/>(e.g., 10 tokens/sec)"] --> Bucket
Request["📱 Request arrives"] --> Check{"Token available?"}
Check -->|"✅ Yes: Take token"| Allow["Allow request"]
Check -->|"❌ No: No tokens"| Deny["429 Too Many Requests"]
Deny --> Retry["Client waits & retries"]
style Allow fill:#059669,color:#fff
style Deny fill:#dc2626,color:#fff
style Bucket fill:#7c3aed,color:#fff

AlgorithmHow It WorksProsCons
Token BucketTokens added at fixed rate. Request consumes a token.Handles bursts, simpleCan allow short-term over-limit
Leaky BucketRequests processed at fixed rate. Excess queues/discards.Smooth output, predictableNo burst support
Fixed WindowCount requests per minute/hour. Reset counter.Simple, memory efficientBurst at window boundaries
Sliding Window LogTrack timestamps of recent requests. Count in window.Accurate, no boundary burstMore memory per user

DimensionExampleWhy
User ID100 requests/minute per userFair usage across users
IP address1000 requests/minute per IPProtect against DDoS
API endpoint10 requests/second for /searchProtect expensive queries
Global100K requests/second totalProtect overall system capacity

CodeMeaningWhat to Do
429Too Many RequestsClient should retry after Retry-After header
429 + Retry-After: 60Retry in 60 secondsClient MUST wait before retrying

// Pseudocode: token bucket per user
function rateLimit(userId, requestCost = 1) {
const bucket = redis.get(`ratelimit:${userId}`);
if (!bucket) {
// First request — create bucket with max tokens
redis.set(`ratelimit:${userId}`, JSON.stringify({
tokens: 10 - requestCost,
lastRefill: Date.now()
}), { ttl: 60 });
return true; // allow
}
// Refill tokens based on elapsed time
const elapsed = (Date.now() - bucket.lastRefill) / 1000;
bucket.tokens = Math.min(10, bucket.tokens + elapsed * rate);
if (bucket.tokens >= requestCost) {
bucket.tokens -= requestCost;
// Save updated bucket
return true; // allow
}
return false; // rate limited (429)
}

  • Rate limiting protects the system but can frustrate legitimate users if too strict.
  • Token bucket is the most popular algorithm — handles bursts while enforcing average rate.
  • Distributed rate limiting (across multiple servers) needs a shared store (Redis) and adds latency.
  • Always return a Retry-After header so clients know when to retry.

  • Rate limiting = saying “you’re going too fast, slow down” to clients.
  • Token bucket is the most common algorithm: tokens drip in, requests use them up.
  • Return 429 with a Retry-After header when a client exceeds the limit.