Load Testing
- In Turkish
- Yük Testi
In short
Load testing is a type of performance testing that simulates many users or requests at once to measure how a system behaves under expected traffic.
What is load testing?
Load testing checks how an application performs when many people use it at the same time. A load testing tool generates traffic, such as thousands of simulated users sending HTTP requests, and measures response times, throughput (the number of requests handled per second), and error rates. The goal is to confirm that the system meets its performance targets before real users find its limits.
A typical load test ramps up virtual users gradually, holds a steady level of traffic for a while, and then ramps down. Results focus on percentiles rather than averages: a p95 latency of 300 ms means 95% of requests finished within 300 ms, which reveals slow outliers that an average hides. While the test runs, teams watch server metrics such as CPU, memory, database connections, and queue lengths to find the bottleneck, the resource that runs out first.
An analogy is testing a new bridge by driving heavy trucks across it before opening it to the public. Teams run load tests before big launches or sales events, after major architecture changes, and when tuning autoscaling rules. Tests should run against an environment that closely matches production, because results from a small laptop setup rarely predict real behavior.
Load testing is often confused with stress testing. A load test checks behavior at the traffic levels you expect, while a stress test keeps increasing traffic beyond them to find the breaking point and see how the system fails and recovers. Related variants include spike tests, which apply a sudden burst of traffic, and soak tests, which hold a normal load for hours to reveal slow problems such as memory leaks.
Key takeaways
- Load testing simulates many concurrent users or requests against a system.
- The key measurements are response time percentiles, throughput, and error rate.
- Percentiles such as p95 and p99 reveal slow outliers that averages hide.
- A load test checks expected traffic; a stress test pushes past it to find the breaking point.
- Test in a production-like environment, and only against systems you are allowed to load.
Example
import time
import urllib.request
from concurrent.futures import ThreadPoolExecutor
URL = "http://localhost:8000/" # only load-test systems you own
def timed_request(_):
start = time.perf_counter()
urllib.request.urlopen(URL).read()
return time.perf_counter() - start
# 50 concurrent virtual users send 1,000 requests in total
with ThreadPoolExecutor(max_workers=50) as pool:
durations = sorted(pool.map(timed_request, range(1000)))
print(f"p95 latency: {durations[949] * 1000:.0f} ms") # the 950th fastest of 1,000Readers ask
What is the difference between load testing and stress testing?
Load testing measures how a system performs under the traffic you expect, such as a normal busy day. Stress testing deliberately pushes traffic beyond that level to find the breaking point and check that the system fails gracefully and recovers.
What is p95 latency?
p95 latency is the response time that 95% of requests were faster than. It is more useful than an average because it shows the experience of the slowest users, which an average tends to hide.
Can I run a load test against production?
Sometimes, but it is risky because test traffic competes with real users and can cause an outage. Most teams test a production-like staging environment, and a load test should only ever target systems you own or have permission to test.
See also
- LatencyNetworking, p. 14Latency is the delay between sending a request and the start of a response, usually measured in milliseconds, and it shapes how responsive an app feels.
- ScalabilitySoftware Architecture, p. 36Scalability is a system's ability to handle growing amounts of work, such as more users or data, by adding resources without a drop in performance.
- AutoscalingDevOps & Cloud, p. 2Autoscaling is the automatic adding or removing of computing resources, such as servers or containers, based on demand to keep performance steady and costs low.
- Load BalancerDevOps & Cloud, p. 34A load balancer is a server or service that spreads incoming traffic across several backend servers so no single one is overloaded and the app stays available.
- BandwidthNetworking, p. 2Bandwidth is the maximum amount of data a network connection can carry per second, usually measured in megabits or gigabits per second (Mbps or Gbps).
- Rate LimitingBackend & APIs, p. 37Rate limiting is a technique that caps how many requests a client can make to a server or API within a time window, protecting it from abuse and overload.
- Performance TestingTesting & Quality, p. 18Performance testing measures how fast, stable and scalable a system is under expected and extreme load, from response times to its breaking point.
Spotted a mistake or something missing on this page?Suggest an edit