Rate Limiting
In short
Rate limiting is a technique that caps how many requests a client can make to a server or API within a time window, protecting it from abuse and overload.
What is rate limiting?
Rate limiting controls how often someone can call a service. A rule such as 100 requests per minute per API key, or 5 login attempts per minute per IP address, is checked on every request, and once a client goes over the limit, further requests are rejected until the time window resets. On the web, the server usually responds with the HTTP status 429 Too Many Requests.
It works like a bouncer at a club who lets in only so many people per minute, however long the line gets. Rate limits protect servers from traffic spikes, buggy clients stuck in a loop, aggressive scraping, and brute-force password guessing, and they help share capacity fairly between customers. They are enforced in API gateways, reverse proxies, CDNs, or middleware inside the application, often with counters kept in a fast shared store such as Redis so every server sees the same count.
Common algorithms include the fixed window, which counts requests per calendar minute; the sliding window, which counts requests in the last 60 seconds; and the token bucket, where each client has a bucket that refills with tokens at a steady rate and every request spends one token. The token bucket is popular because it allows short bursts while still enforcing an average rate.
Rate limiting is closely related to throttling, and the two terms are often used interchangeably, but throttling sometimes means slowing requests down or queuing them instead of rejecting them. Well-behaved APIs tell clients about their limits with response headers such as Retry-After or X-RateLimit-Remaining, and well-behaved clients respond to a 429 by waiting and retrying with exponential backoff, meaning longer pauses after each failed attempt.
At a glance
Key takeaways
- Rate limiting caps how many requests a client can make in a period of time.
- Clients over the limit usually receive HTTP
429 Too Many Requests. - Limits are typically applied per API key, user, or IP address.
- Common algorithms are fixed window, sliding window, and token bucket.
- Clients should respect
Retry-Afterand retry with exponential backoff.
Example
// Fixed-window limiter: 100 requests per minute per IP (in memory, one server)
const LIMIT = 100;
const WINDOW_MS = 60_000;
const hits = new Map(); // ip -> { count, start }
function rateLimit(req, res, next) {
const now = Date.now();
let entry = hits.get(req.ip);
if (!entry || now - entry.start >= WINDOW_MS) {
entry = { count: 0, start: now }; // start a new window
hits.set(req.ip, entry);
}
if (++entry.count > LIMIT) return res.status(429).send("Too Many Requests");
next();
}Readers ask
What does HTTP 429 Too Many Requests mean?
It means you have sent more requests than the server allows in a given time. Wait before retrying, ideally for the number of seconds given in the Retry-After header if the response includes one.
What is the difference between rate limiting and throttling?
They are often used as synonyms. When a distinction is made, rate limiting rejects requests over the limit, while throttling slows them down or queues them so they are processed at a controlled pace.
What is the token bucket algorithm?
Each client gets a bucket that refills with tokens at a fixed rate, and every request uses one token. When the bucket is empty, requests are rejected, so short bursts are allowed but the average rate stays under the limit.
See also
- APIBackend & APIs, p. 2An API is a set of rules that lets one piece of software request data or actions from another in a predictable, documented way.
- API GatewayBackend & APIs, p. 3An API gateway is a server that sits in front of a group of backend services and acts as the single entry point that receives, checks, and routes API requests.
- MiddlewareBackend & APIs, p. 30Middleware is software that sits between two layers of a system, most often code that runs between an incoming request and the final response in a web server.
- HTTPWeb Development, p. 19HTTP is the protocol that browsers, apps, and servers use to exchange web pages and data through a simple cycle of requests and responses.
- Load BalancerDevOps & Cloud, p. 34A load balancer is a server or service that spreads incoming traffic across several backend servers so no single one is overloaded and the app stays available.
- API KeySecurity, p. 1An API key is a unique secret string that identifies an application or project when it calls an API, used to control access, track usage, and apply rate limits.
- ThrottlingWeb Development, p. 47Throttling makes a function run at most once per interval, however often it is called, so scroll and resize handlers update steadily without running every time.
Sources
Spotted a mistake or something missing on this page?Suggest an edit