Backpressure
In short
Backpressure is a mechanism that lets a slow consumer signal a fast producer to slow down, so data doesn't pile up faster than it can be processed.
What is backpressure?
Backpressure is resistance that flows backward through a system: when one part produces data faster than the next part can handle it, the slower part pushes back and tells the producer to slow down or pause. Without it, the excess has to go somewhere, usually into ever-growing memory buffers, which leads to rising latency, out-of-memory crashes, or data being dropped at random.
Systems handle overload in a few ways. The consumer can control the flow by pulling data only when it is ready, as in pull-based streams and reactive libraries. A bounded buffer or queue can make the producer wait when it is full. Or the system can shed load on purpose by rejecting excess work, for example with 429 Too Many Requests or 503 Service Unavailable. In Node.js, stream.write() returns false when the internal buffer is full, and a well-behaved producer waits for the drain event before writing more, while pipeline() handles this automatically. TCP has backpressure built in, since its flow control shrinks the sender's window when the receiver's buffer fills up.
Picture a restaurant during the dinner rush: if the cooks are overwhelmed, the host stops seating new guests for a while instead of letting order tickets pile up until every meal is late. Backpressure matters wherever data moves between stages that run at different speeds, such as file and network streams, message queue consumers, event streaming pipelines, logging systems, and microservices calling each other.
Backpressure is often confused with rate limiting and buffering. Rate limiting enforces a fixed maximum, such as 100 requests per minute per client, no matter how busy the server is, while backpressure reacts to the consumer's actual capacity at that moment. A buffer only absorbs short bursts, and an unbounded buffer hides the problem until memory runs out, which is why buffers in a backpressure-aware system always have a limit. A message queue that grows forever is a sign that consumers need to scale up or producers need to slow down.
Key takeaways
- Backpressure lets a slow consumer tell a fast producer to slow down or pause.
- Without it, unbounded buffers grow until latency spikes or memory runs out.
- Common strategies are pull-based consumption, bounded buffers that block, and load shedding.
- Node.js streams signal backpressure when
write()returnsfalse. - Rate limiting is a fixed cap, while backpressure adapts to real-time capacity.
Example
import { createReadStream, createWriteStream } from "node:fs";
import { pipeline } from "node:stream/promises";
// pipeline() handles backpressure: reading pauses whenever
// the slower destination's buffer is full, then resumes
await pipeline(createReadStream("huge.log"), createWriteStream("copy.log"));
// Doing it by hand: respect the return value of write()
function writeChunk(stream, chunk, next) {
if (stream.write(chunk)) next(); // buffer has room: keep going
else stream.once("drain", next); // buffer full: wait for "drain"
}Readers ask
What happens without backpressure?
Data piles up in memory between the fast and slow parts of the system. Latency climbs as the backlog grows, and eventually the process runs out of memory, crashes, or starts dropping data.
Is backpressure the same as rate limiting?
No. Rate limiting applies a fixed, predefined cap on requests, while backpressure is a dynamic signal based on how much the consumer can handle right now. Systems often use both.
How do message queues handle backpressure?
Consumers pull messages at their own pace, often with a limit on how many they take at once, so the queue absorbs bursts. A growing backlog, called consumer lag, is the signal to add consumers or slow producers, and bounded queues can reject new messages when full.
See also
- Message QueueBackend & APIs, p. 29A message queue is a component that stores messages from one service until another is ready to process them, so parts of a system can work asynchronously.
- Event StreamingBackend & APIs, p. 14Event streaming is the practice of recording events as a continuous, ordered and durable log that many applications can read, replay and process in real time.
- Rate LimitingBackend & APIs, p. 37Rate limiting is a technique that caps how many requests a client can make to a server or API within a time window, protecting it from abuse and overload.
- TCPNetworking, p. 30TCP is a core internet protocol that delivers data between two programs reliably and in order, by opening a connection and resending anything that gets lost.
- Node.jsBackend & APIs, p. 32Node.js is an open-source JavaScript runtime that runs JavaScript outside the browser, most often to build web servers, APIs, and command-line tools.
- LatencyNetworking, p. 14Latency is the delay between sending a request and the start of a response, usually measured in milliseconds, and it shapes how responsive an app feels.
Spotted a mistake or something missing on this page?Suggest an edit