Rate limiting controls how many requests a client can make in a time window. Effective rate limiting prevents abuse without blocking legitimate users. Rate limits must balance security against user experience.
Why rate limiting matters
Rate limiting protects against abuse: brute-force attacks, scraping, denial-of-service, and resource exhaustion. Without rate limits, attackers can overwhelm services or guess credentials.
Attack vectors rate limiting addresses
- Brute-force attacks: Repeated login attempts to guess passwords
- Credential stuffing: Testing stolen username-password pairs
- API abuse: Excessive requests that degrade service performance
- Data scraping: Automated extraction of data at scale
Rate limiting algorithms
Different algorithms trade off precision, memory usage, and implementation complexity.
Fixed window
Fixed window counts requests within fixed time intervals (e.g., per minute). Simple to implement but allows bursts at interval boundaries.
// Example: fixed window counter
const window = Math.floor(Date.now() / 60000); // 1-minute window
const key = `ratelimit:${userId}:${window}`;
const count = await redis.incr(key);
await redis.expire(key, 120); // Keep for 2 windows
if (count > 100) {
throw new RateLimitError();
}
Sliding window
Sliding window tracks requests over a rolling time period. More accurate than fixed windows but requires storing timestamps.
Token bucket
Token bucket allows bursts while enforcing average rate. Tokens are added at a fixed rate. Requests consume tokens. When tokens are exhausted, requests are rejected.
Leaky bucket
Leaky bucket smooths traffic by processing requests at a constant rate. Incoming requests queue. If the queue is full, new requests are rejected.
Implementation patterns
Where to enforce rate limits
Rate limits can be enforced at multiple layers:
- Edge (CDN/API Gateway): Protects infrastructure from reaching application servers
- Application middleware: Enforces business logic limits
- Per-endpoint: Different limits for different operations
Client identification
Rate limits require identifying clients. Options include:
- API key: Most accurate for authenticated APIs
- User ID: For logged-in users
- IP address: Fallback for unauthenticated requests, but shared IPs cause collateral blocking
Response headers
Include rate limit information in HTTP headers so clients can self-regulate:
HTTP/1.1 200 OK
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 47
X-RateLimit-Reset: 1699876543
Setting appropriate limits
Limits must reflect actual usage patterns. Too strict limits frustrate users. Too lenient limits fail to prevent abuse.
Tiered limits
Different endpoints need different limits. Authentication endpoints require stricter limits than read operations.
// Example: per-endpoint limits
const limits = {
'POST /auth/login': { requests: 5, window: 60 },
'GET /api/posts': { requests: 100, window: 60 },
'POST /api/posts': { requests: 10, window: 60 }
};
Authenticated vs anonymous limits
Authenticated users typically get higher limits than anonymous requests. This incentivizes registration while protecting against anonymous abuse.
Handling rate limit violations
HTTP 429 status
Return 429 Too Many Requests with a Retry-After header indicating when the client can retry.
HTTP/1.1 429 Too Many Requests
Retry-After: 60
Content-Type: application/json
{
"error": "rate_limit_exceeded",
"message": "Too many requests. Try again in 60 seconds."
}
Progressive penalties
Repeated violations can trigger escalating responses: temporary blocks, longer cooldowns, or account review.
Avoiding false positives
Shared IPs
Corporate networks and VPNs share IP addresses. IP-based limiting can block entire organizations. Prefer user-based or API key-based limiting when possible.
Allowlisting
Trusted clients (internal services, partners) should bypass rate limits or have higher thresholds.
Monitoring rate limits
Track rate limit metrics to tune limits and detect attacks.
Key metrics
- Rate limit hit rate (percentage of requests rejected)
- Unique clients hitting limits
- Endpoints with highest rejection rates
Rate limiting is a balance
Effective rate limiting protects services without degrading user experience. Set limits based on actual usage, provide clear feedback, and monitor for both attacks and false positives. For related security patterns, see Security-First API Design.
Published by the DSSS Engineering Team. For corrections or topic requests, use the contact page.