Skip to main content

Side by side

LatencyvsThroughput

What is the difference between latency and throughput?

Updated 2 min read6 differences

In short

Latency is how long one request or piece of data takes to arrive, while throughput is how much data or how many requests a system handles per unit of time.

Latency

Latency is the delay between sending a request and the start of a response, usually measured in milliseconds, and it shapes how responsive an app feels.

Read the page on Latency

Throughput

Throughput is the amount of data or work a system actually handles per unit of time, like megabits per second on a network or requests per second on a server.

Read the page on Throughput

Latency and Throughput compared

AspectLatencyThroughput
MeasuresDelay for one request or packetWork or data handled per unit of time
UnitsMillisecondsRequests per second, Mbps
AnalogyHow long one car takes for the tripHow many cars pass per hour
Improved byShorter distance, fewer round trips, cachingParallelism, batching, more capacity
Users notice it asResponsivenessCapacity during peaks
Reported asPercentiles such as p95 and p99Averages and peaks over time

The difference, explained

Latency measures delay: the milliseconds between clicking a button and seeing the response, or a packet traveling from Istanbul to Frankfurt. Throughput measures volume: how many requests per second a server completes, or how many megabits per second a link actually carries. A highway analogy helps: latency is how long one car takes to make the trip, throughput is how many cars pass per hour.

The two are related but independent. A satellite link can have high throughput but high latency; a lightly loaded server can answer quickly but handle few requests at once. Under load they interact: as a system approaches its maximum throughput, queues form and latency rises sharply, which load tests reveal as the point where response times curve upward.

They are improved in different ways. Latency falls with shorter distances, such as CDNs and edge servers, fewer round trips, faster code paths and caching. Throughput rises with more parallelism, batching, more instances, better hardware and removing bottlenecks such as locks or slow queries. Sometimes they trade off: batching increases throughput but makes individual items wait.

A common misconception is that bandwidth, throughput and latency are one thing called speed. Bandwidth is the maximum capacity of a link, throughput is what is actually achieved, and latency is the delay. A fast connection for downloading large files can still feel slow for gaming or video calls if its latency is high.

Which one should you use?

Choose Latency when…

  • You build interactive apps, games or video calls.
  • Users wait on each individual response.
  • You care about tail latency, the slowest requests.

Choose Throughput when…

  • You process large batches, streams or file transfers.
  • You need to handle many requests at peak times.
  • Total work done matters more than each item's delay.

Readers ask

Can a system have high throughput and high latency?

Yes. A satellite link or a batch processing job can move a lot of data while each piece still takes a long time to arrive.

Why does latency increase under load?

As a system nears its capacity, requests start waiting in queues for CPU, connections or locks, so each one takes longer even though throughput stays high.

What is the difference between bandwidth and throughput?

Bandwidth is the maximum rate a link could carry. Throughput is the rate actually achieved, usually lower because of overhead, congestion and the devices at each end.

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings