Scalability
- In Turkish
- Ölçeklenebilirlik
In short
Scalability is a system's ability to handle growing amounts of work, such as more users or data, by adding resources without a drop in performance.
What is scalability?
Scalability describes how well a system copes as demand grows. A scalable application can go from a hundred users to a million by adding resources while keeping response times and error rates acceptable. It applies to every layer of a system, including web servers, databases, and queues, and even to the team and codebase that maintain them.
There are two main ways to scale. Vertical scaling, or scaling up, means giving one machine more power, such as more CPU cores, memory, or faster disks; it is simple but has a hard ceiling and leaves a single point of failure. Horizontal scaling, or scaling out, means adding more machines and spreading the work among them with a load balancer; it can grow much further and improves resilience, but the application must be designed for it, for example by keeping servers stateless and storing sessions in a shared cache or database.
Think of a busy restaurant. Scaling vertically is hiring a faster chef, which only helps up to a point, while scaling horizontally is opening more kitchens, which needs coordination but can serve far more diners. Common scaling techniques include caching, database replication for read-heavy traffic, sharding to split large datasets, message queues to absorb traffic spikes, and cloud autoscaling to add or remove servers automatically.
Scalability is often confused with performance. Performance is how fast a system handles a request under its current load, while scalability is how well it keeps that performance as load increases; a fast app can still collapse under heavy traffic. Scalability is also different from elasticity, which is the ability to scale up and down automatically as demand changes.
Key takeaways
- Scalability is the ability to handle more load by adding resources.
- Vertical scaling (scaling up) adds power to one machine and has a hard limit.
- Horizontal scaling (scaling out) adds more machines and usually requires stateless services.
- Caching, replication, sharding, and queues are common scaling techniques.
- Performance is speed under current load; scalability is keeping that speed as load grows.
Example
# Vertical scaling: give one container more CPU and memory
docker run --cpus=4 --memory=8g my-app
# Horizontal scaling: run three identical copies of the web service
docker compose up --scale web=3
# In Kubernetes, change the number of copies (replicas) of a deployment
kubectl scale deployment/web-app --replicas=5Readers ask
What is the difference between horizontal and vertical scaling?
Vertical scaling makes one server bigger by adding CPU, memory, or storage, while horizontal scaling adds more servers and splits the work between them. Vertical scaling is simpler, but horizontal scaling can grow much further and survives the failure of a single machine.
What makes an application hard to scale?
Common obstacles are servers that keep user sessions in local memory, a single database that every request depends on, and slow tasks that run inside the request. Making services stateless, adding caches and read replicas, and moving heavy work to background queues all help.
Do microservices make a system scalable?
They can, because each service can be scaled independently, but they are not required. A well-designed monolith can also scale horizontally, and microservices add network and operational complexity that small teams may not need.
See also
- Load BalancerDevOps & Cloud, p. 34A load balancer is a server or service that spreads incoming traffic across several backend servers so no single one is overloaded and the app stays available.
- CacheBackend & APIs, p. 8A cache is a fast, temporary storage layer that keeps copies of frequently used data so later requests can be served quickly without repeating slow work.
- ShardingDatabases, p. 39Sharding is a way of scaling a database by splitting its data across several servers, called shards, so each one stores and handles only part of the total.
- Database ReplicationDatabases, p. 10Database replication is the continuous copying of data from one database server to others, so several servers hold the same data for reliability and scale.
- MicroservicesSoftware Architecture, p. 27Microservices are an architectural style where an application is split into small, independently deployable services that communicate over a network.
- Cloud ComputingDevOps & Cloud, p. 10Cloud computing is the on-demand delivery of computing resources, such as servers, storage, and databases, over the internet with pay-as-you-go pricing.
- Horizontal ScalingSoftware Architecture, p. 23Horizontal scaling (scaling out) increases a system's capacity by adding machines and spreading the work across them, rather than making one machine bigger.
- Vertical ScalingSoftware Architecture, p. 47Vertical scaling (scaling up) increases a system's capacity by giving a single machine more CPU, memory or faster storage, instead of adding more machines.
Spotted a mistake or something missing on this page?Suggest an edit