Skip to main content

Rate Limiting

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/rate-limiting

In short

Rate limiting is a technique that caps how many requests a client can make to a server or API within a time window, protecting it from abuse and overload.

What is rate limiting?

Rate limiting controls how often someone can call a service. A rule such as 100 requests per minute per API key, or 5 login attempts per minute per IP address, is checked on every request, and once a client goes over the limit, further requests are rejected until the time window resets. On the web, the server usually responds with the HTTP status 429 Too Many Requests.

It works like a bouncer at a club who lets in only so many people per minute, however long the line gets. Rate limits protect servers from traffic spikes, buggy clients stuck in a loop, aggressive scraping, and brute-force password guessing, and they help share capacity fairly between customers. They are enforced in API gateways, reverse proxies, CDNs, or middleware inside the application, often with counters kept in a fast shared store such as Redis so every server sees the same count.

Common algorithms include the fixed window, which counts requests per calendar minute; the sliding window, which counts requests in the last 60 seconds; and the token bucket, where each client has a bucket that refills with tokens at a steady rate and every request spends one token. The token bucket is popular because it allows short bursts while still enforcing an average rate.

Rate limiting is closely related to throttling, and the two terms are often used interchangeably, but throttling sometimes means slowing requests down or queuing them instead of rejecting them. Well-behaved APIs tell clients about their limits with response headers such as Retry-After or X-RateLimit-Remaining, and well-behaved clients respond to a 429 by waiting and retrying with exponential backoff, meaning longer pauses after each failed attempt.

At a glance

A token bucket that refills with one token every second holds three tokens when four requests arrive at once: requests 1 to 3 each spend a token and get 200 OK, and request 4 finds the bucket empty and gets 429 Too Many Requests.+1 token every secondToken bucketRequest 1token200 OKRequest 2token200 OKRequest 3token200 OKRequest 4empty429 Too Many Requests
Each request spends one token and tokens come back at a steady rate, so short bursts pass but the average rate can never exceed the refill rate.

Key takeaways

  • Rate limiting caps how many requests a client can make in a period of time.
  • Clients over the limit usually receive HTTP 429 Too Many Requests.
  • Limits are typically applied per API key, user, or IP address.
  • Common algorithms are fixed window, sliding window, and token bucket.
  • Clients should respect Retry-After and retry with exponential backoff.

Example

A simple rate-limiting middleware in Express.jsjavascript
// Fixed-window limiter: 100 requests per minute per IP (in memory, one server)
const LIMIT = 100;
const WINDOW_MS = 60_000;
const hits = new Map(); // ip -> { count, start }

function rateLimit(req, res, next) {
  const now = Date.now();
  let entry = hits.get(req.ip);
  if (!entry || now - entry.start >= WINDOW_MS) {
    entry = { count: 0, start: now }; // start a new window
    hits.set(req.ip, entry);
  }
  if (++entry.count > LIMIT) return res.status(429).send("Too Many Requests");
  next();
}

Readers ask

What does HTTP 429 Too Many Requests mean?

It means you have sent more requests than the server allows in a given time. Wait before retrying, ideally for the number of seconds given in the Retry-After header if the response includes one.

What is the difference between rate limiting and throttling?

They are often used as synonyms. When a distinction is made, rate limiting rejects requests over the limit, while throttling slows them down or queues them so they are processed at a controlled pace.

What is the token bucket algorithm?

Each client gets a bucket that refills with tokens at a fixed rate, and every request uses one token. When the bucket is empty, requests are rejected, so short bursts are allowed but the average rate stays under the limit.

See also

Sources

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings