Imagine your API suddenly receives 1 million requests in a few seconds . Maybe it's a viral product launch. Maybe a client accidentally created an infinite loop. Or maybe someone is deliberately trying to overwhelm your system. Without protection, your servers could become overloaded, response times could increase, and legitimate users could start receiving errors. This is where rate limiting comes in. Rate limiting controls how many requests a client can make within a specific period of time . In this article, we'll understand how rate limiting works, the most common algorithms, where to implement it, how distributed systems handle it, and how to design rate limiting for large-scale APIs. What Is Rate Limiting? Rate limiting is a mechanism that restricts the number of requests a client can make during a given time window. For example: 100 requests / minute A client can make up to 100 requests during that period. After reaching the limit: Request 101 ↓ Rate Limiter ↓ ❌ Too Many Requests The API typically responds with: HTTP / 1.1 429 Too Many Requests Rate limiting protects your infrastructure while ensuring that one client cannot consume an unreasonable amount of resources. Why Do We Need Rate Limiting? Without rate limiting, an API might look like: Client 1 ───────┐ Client 2 ───────┤ Client 3 ───────┤ Client 4 ───────┼──→ API Servers Client 5 ───────┤ Client 6 ───────┘ Now imagine one client starts sending thousands of requests per second. Attacker │ ├── Request ├── Request ├── Request ├── Request ├── Request └── ... ↓ API Servers This can lead to: Increased CPU usage Higher memory consumption Database overload Increased network traffic Higher infrastructure costs Slow responses Service outages Rate limiting adds a protective layer: Clients ↓ Rate Limiter ↓ Allowed Requests ↓ API Servers Rate Limiting vs Throttling These terms are often used together, but they can represent different behaviors. Rate Limiting Restricts how many requests can be made. 100 requests / minute Requests beyond the limit may be rejected. Throttling Controls the rate at which requests are processed. For example, instead of rejecting everything immediately, the system may slow down processing or queue requests. A simple distinction: Rate limiting controls how much traffic is allowed. Throttling controls how quickly traffic is processed. A Simple Rate Limiting Example Suppose an API has this rule: 100 requests / minute / user A user sends: Request 1 Request 2 Request 3 ... Request 100 All requests are allowed. The next request: Request 101 ↓ ❌ 429 Too Many Requests After the limit resets, requests can be accepted again. What Should We Rate Limit? Not every endpoint needs the same limit. For example: GET /products → 1000 requests/minute
POST /login → 10 requests/minute
POST /payments → 30 requests/minute
POST /password-reset → 5 requests/hour Expensive or security-sensitive endpoints generally need stricter limits. Rate Limiting Strategies There are several common algorithms used to implement rate limiting. The most important ones are: Fixed Window Sliding Window Token Bucket Leaky Bucket Let's understand each. 1. Fixed Window The fixed-window algorithm divides time into fixed intervals. For example: Limit: 100 requests / minute The system creates windows: 12:00:00 ───────── 12:01:00 12:01:00 ───────── 12:02:00 12:02:00 ───────── 12:03:00 Within each window: Counter = 0
Request → Counter = 1
Request → Counter = 2
Request → Counter = 3
...
Request → Counter = 100 Once the counter reaches the limit: Request 101
↓
Rejected At the next window: Counter → 0 Advantages Simple to implement Easy to understand Low memory usage Problem Fixed windows can create boundary spikes . For example: 12:00:59 → 100 requests
12:01:00 → 100 requests A client could potentially send 200 requests in a very short period while staying within two different windows. 2. Sliding Window A sliding-window algorithm evaluates requests over a continuously moving time period. For example: Last 60 seconds Instead of resetting everything at exactly 12:01:00, the system continuously checks the previous 60 seconds. Conceptually: Current Time
↓
───────●────●────●────●────●──────
←── Last 60 seconds ──→ The system counts requests within that moving window. Advantages More accurate Reduces boundary spikes Better traffic control Trade-off It can require more memory or more sophisticated data structures depending on the implementation. 3. Token Bucket The token bucket algorithm is one of the most widely used approaches. Imagine a bucket that holds tokens. Token Generator
↓
┌─────────────┐
│ ● ● ● ● ● │
│ Token │
│ Bucket │
└─────────────┘
↓
API Request Each request consumes a token. Request
↓
Take 1 Token
↓
Token Available?
/
Yes No
↓ ↓
Allow Reject Tokens are added to the bucket at a fixed rate. For example: Bucket capacity: 100 tokens
Refill rate: 10 tokens/second A client can temporarily make a burst of requests if tokens have accumulated. But once the bucket is empty, additional requests are rejected or delayed. Why Token Bucket Is Useful It supports both: Sustained traffic Controlled bursts This makes it a popular choice for APIs. 4. Leaky Bucket The leaky-bucket algorithm behaves more like a queue. Imagine requests entering a bucket: Requests
↓
┌─────────────┐
│ Request │
│ Request │
│ Request │
│ Request │
└──────┬──────┘
↓
Fixed Rate
↓
API Requests are processed at a relatively consistent rate. If the queue becomes full, new requests may be rejected. Token Bucket vs Leaky Bucket A useful distinction is: Token Bucket Allows controlled bursts Leaky Bucket Produces a smoother processing rate The right choice depends on your traffic pattern. Comparing Rate Limiting Algorithms Algorithm Main Idea Burst Support Complexity Fixed Window Fixed time counters Yes Low Sliding Window Moving time window Limited Medium Token Bucket Tokens refill over time Yes Medium Leaky Bucket Process requests at fixed rate Limited Medium There is no universally best algorithm. The choice depends on whether your system needs simplicity, burst handling, accuracy, or smooth traffic. Where Should Rate Limiting Be Implemented? Rate limiting can be implemented at different layers. Application Layer Client
↓
Application
↓
Rate Limiter
↓
Database This gives the application detailed control over users and endpoints. API Gateway Client
↓
API Gateway
↓
Rate Limiter
↓
Services This is useful when multiple services need consistent protection. CDN / Edge User
↓
Edge
↓
Rate Limiter
↓
Origin Rate limiting at the edge can block unwanted traffic before it reaches your infrastructure. Multiple Layers Large systems may use several layers: User
↓
CDN
↓
API Gateway
↓
Service
↓
Database Each layer can have different limits. What Should the Rate Limit Key Be? A rate limiter needs to determine who or what is being limited . Possible keys include: IP Address Rate limit per IP Useful for unauthenticated endpoints. But many users can share the same public IP. User ID Rate limit per user Useful for authenticated APIs. API Key Rate limit per API key Common for developer-facing APIs. Endpoint Rate limit per endpoint Useful when different APIs have different costs. Combined Key For example: user_id + endpoint This provides more granular control. Distributed Rate Limiting Here's where system design becomes interesting. Imagine your API has multiple servers: Load Balancer
/ |
↓ ↓ ↓
Server 1 Server 2 Server 3 If each server maintains its own counter: Server 1 → 40 requests
Server 2 → 40 requests
Server 3 → 40 requests A user could effectively make: 120 requests even if the intended limit was 100. The counters are not shared. Shared Rate Limiter A common solution is to use a shared, fast data store. For example: Load Balancer
/ |
▼ ▼ ▼
Server 1 Server 2 Server 3
\ | /
\ | /
▼ ▼
Redis Now all servers can use the same rate-limit state. Request
↓
Any API Server
↓
Shared Counter
↓
Redis This allows the system to enforce a global limit across multiple application instances. Rate Limiting With Redis A simplified concept might look like: Key:
rate:user:123
Value: 87
TTL: 60 seconds Every request increments the counter. Request ↓ Increment counter ↓ Counter <= Limit? │ ┌─┴───┐ Yes No ↓ ↓ Allow Reject For more sophisticated algorithms, Redis can store timestamps, token counts, or other state. Atomic operations are important because multiple requests can arrive simultaneously. Race Conditions Imagine two requests arrive at exactly the same time. Both servers read: Counter = 99 Both think: 99 < 100 Both increment it. Now the actual count could become inconsistent. This is why distributed rate limiting needs atomic operations or carefully designed server-side logic. The rate limiter itself must be safe under high concurrency. What Response Should the API Return? When a client exceeds the limit, the standard response is: 429 Too Many Requests The API can also communicate useful information about when the client can retry. For example: Retry-After: 30 This tells the client to wait before trying again. Well-designed clients should respect these signals instead of immediately retrying. Rate Limiting and Retry Storms Imagine a service becomes overloaded. Clients receive errors and immediately retry. Those retries create even more traffic. Service Overloaded ↓ Errors ↓ Clients ↓ Immediate Retry ↓ More Traffic ↓ More Errors This can become a retry storm . Rate limiting works best alongside strategies such as: Exponential backoff Jitter Retry limits Circuit breakers Timeouts The goal is to prevent failures from amplifying themselves. Rate Limiting vs Authentication These systems often work together. For example: Request ↓ Authentication ↓ Rate Limiting ↓ Authorization ↓ Application Or at an API gateway: Client ↓ API Gateway ├── Authentication ├── Rate Limiting └── Routing ↓ Service Authentication identifies the client. Rate limiting controls how much traffic that client can generate. Authorization determines what the client is allowed to access. Rate Limiting Expensive Operations Not every request costs the same. Consider: GET /health versus: POST /generate-report The second request may consume significantly more CPU, memory, database resources, or external API calls. Using the same rate limit for both endpoints may not make sense. You can assign different limits based on cost: Health Check → High limit
Read API → Medium limit
Heavy Computation → Low limit This is sometimes called cost-based rate limiting . Rate Limiting at Large Scale A large production architecture might look like: Users │ ▼ CDN │ ▼ API Gateway │ ┌───────┴───────┐ │ Rate Limiter │ └───────┬───────┘ │ Load Balancer │ ┌───────────┼───────────┐ ▼ ▼ ▼ Server 1 Server 2 Server 3 │ │ │ └───────────┼───────────┘ ▼ Database A shared store can maintain rate-limit state: Servers │ ▼ Redis │ ▼ Rate Limit State This allows rate limiting to work across the entire application cluster. Common Rate Limiting Mistakes Some common mistakes include: Using only in-memory counters with multiple servers Applying the same limit to every endpoint Ignoring shared IP addresses Not handling bursts correctly Forgetting about concurrent requests Returning errors without retry information Allowing clients to retry immediately Using a single global limit for different user tiers Not monitoring rejected requests Making the rate limiter itself a single point of failure Rate limiting is infrastructure. It needs to scale alongside the system it protects. Rate Limiting for Different Users Not every user needs the same limits. For example: Free Plan → 100 requests/minute
Pro Plan → 1,000 requests/minute
Enterprise → Custom limits This allows APIs to support different usage levels. The rate-limit key could include: user_id + plan + endpoint This gives the system more granular control. The Big Picture Rate limiting is essentially a traffic-control mechanism for your APIs . Without it: Clients ↓ API ↓ Overload With it: Clients ↓ Rate Limiter ↓ Allowed Traffic ↓ API ↓ Database The rate limiter acts as a protective boundary between unpredictable client traffic and your infrastructure. Key Takeaway Rate limiting isn't just about blocking excessive requests. It's about making your system more predictable, resilient, and fair . A good rate-limiting strategy can: Protect APIs from abuse Handle traffic spikes Prevent resource exhaustion Protect databases Improve system stability Provide fair usage across clients Reduce infrastructure costs For small systems, a simple fixed-window limiter may be enough. For larger distributed systems, you may need token buckets, shared Redis state, API gateways, edge protection, and carefully designed retry behavior. The goal isn't to stop traffic. The goal is to control traffic so your system can keep serving legitimate users reliably.


