Rate limiting algorithms and strategies that protect APIs without blocking legitimate users
Security 9 min read

Rate Limiting That Protects Without Blocking Legitimate Users

Rate limiting controls how many requests a client can make in a time window. Effective rate limiting prevents abuse without blocking legitimate users. Rate limits must balance security against user experience.

Why rate limiting matters

Rate limiting protects against abuse: brute-force attacks, scraping, denial-of-service, and resource exhaustion. Without rate limits, attackers can overwhelm services or guess credentials.

Attack vectors rate limiting addresses

  • Brute-force attacks: Repeated login attempts to guess passwords
  • Credential stuffing: Testing stolen username-password pairs
  • API abuse: Excessive requests that degrade service performance
  • Data scraping: Automated extraction of data at scale

Rate limiting algorithms

Different algorithms trade off precision, memory usage, and implementation complexity.

Fixed window

Fixed window counts requests within fixed time intervals (e.g., per minute). Simple to implement but allows bursts at interval boundaries.

// Example: fixed window counter
const window = Math.floor(Date.now() / 60000); // 1-minute window
const key = `ratelimit:${userId}:${window}`;
const count = await redis.incr(key);
await redis.expire(key, 120); // Keep for 2 windows

if (count > 100) {
  throw new RateLimitError();
}

Sliding window

Sliding window tracks requests over a rolling time period. More accurate than fixed windows but requires storing timestamps.

Token bucket

Token bucket allows bursts while enforcing average rate. Tokens are added at a fixed rate. Requests consume tokens. When tokens are exhausted, requests are rejected.

Leaky bucket

Leaky bucket smooths traffic by processing requests at a constant rate. Incoming requests queue. If the queue is full, new requests are rejected.

Implementation patterns

Where to enforce rate limits

Rate limits can be enforced at multiple layers:

  • Edge (CDN/API Gateway): Protects infrastructure from reaching application servers
  • Application middleware: Enforces business logic limits
  • Per-endpoint: Different limits for different operations

Client identification

Rate limits require identifying clients. Options include:

  • API key: Most accurate for authenticated APIs
  • User ID: For logged-in users
  • IP address: Fallback for unauthenticated requests, but shared IPs cause collateral blocking

Response headers

Include rate limit information in HTTP headers so clients can self-regulate:

HTTP/1.1 200 OK
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 47
X-RateLimit-Reset: 1699876543

Setting appropriate limits

Limits must reflect actual usage patterns. Too strict limits frustrate users. Too lenient limits fail to prevent abuse.

Tiered limits

Different endpoints need different limits. Authentication endpoints require stricter limits than read operations.

// Example: per-endpoint limits
const limits = {
  'POST /auth/login': { requests: 5, window: 60 },
  'GET /api/posts': { requests: 100, window: 60 },
  'POST /api/posts': { requests: 10, window: 60 }
};

Authenticated vs anonymous limits

Authenticated users typically get higher limits than anonymous requests. This incentivizes registration while protecting against anonymous abuse.

Handling rate limit violations

HTTP 429 status

Return 429 Too Many Requests with a Retry-After header indicating when the client can retry.

HTTP/1.1 429 Too Many Requests
Retry-After: 60
Content-Type: application/json

{
  "error": "rate_limit_exceeded",
  "message": "Too many requests. Try again in 60 seconds."
}

Progressive penalties

Repeated violations can trigger escalating responses: temporary blocks, longer cooldowns, or account review.

Avoiding false positives

Shared IPs

Corporate networks and VPNs share IP addresses. IP-based limiting can block entire organizations. Prefer user-based or API key-based limiting when possible.

Allowlisting

Trusted clients (internal services, partners) should bypass rate limits or have higher thresholds.

Monitoring rate limits

Track rate limit metrics to tune limits and detect attacks.

Key metrics

  • Rate limit hit rate (percentage of requests rejected)
  • Unique clients hitting limits
  • Endpoints with highest rejection rates

Rate limiting is a balance

Effective rate limiting protects services without degrading user experience. Set limits based on actual usage, provide clear feedback, and monitor for both attacks and false positives. For related security patterns, see Security-First API Design.


Published by the DSSS Engineering Team. For corrections or topic requests, use the contact page.