Autoscaling
- In Turkish
- Otomatik Ölçekleme
In short
Autoscaling is the automatic adding or removing of computing resources, such as servers or containers, based on demand to keep performance steady and costs low.
What is autoscaling?
Autoscaling is a cloud and container feature that automatically changes how much computing capacity an application has. When traffic rises, it adds more servers, virtual machines, or containers; when traffic falls, it removes them. The goal is to keep the app responsive during peaks without paying for idle machines the rest of the time.
An autoscaler watches metrics such as CPU usage, memory, request rate, or the length of a message queue, and compares them with a target, for example 70% average CPU. If the metric stays above the target, it starts new instances, usually behind a load balancer; if it stays below, it shuts some down, always within a configured minimum and maximum. Some systems also use scheduled scaling for predictable peaks, or predictive scaling that forecasts demand from past patterns.
Think of a supermarket that opens more checkout lanes when the lines grow long and closes them when the store is quiet. Autoscaling is used by web applications with daily traffic cycles, by background workers that process queues, and by Kubernetes clusters, which can scale both the number of pods and the number of nodes they run on.
Autoscaling is usually horizontal, meaning it adds more copies of the application, which is different from vertical scaling, where a single machine gets more CPU or memory. Horizontal autoscaling works best when the application is stateless, so any instance can handle any request. It is also not instant: new instances take time to start, so sudden spikes may still need spare capacity or rate limiting.
Key takeaways
- Autoscaling adds or removes capacity automatically based on demand.
- It is driven by metrics such as CPU, memory, request rate, or queue length.
- Minimum and maximum limits keep scaling safe and costs predictable.
- Horizontal scaling adds instances, while vertical scaling makes one instance bigger.
- Stateless applications are the easiest to autoscale.
Example
# Keep average CPU near 70% by running between 2 and 10 pods
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web-app
spec:
scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: web-app }
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target: { type: Utilization, averageUtilization: 70 }Readers ask
What is the difference between horizontal and vertical scaling?
Horizontal scaling adds more machines or containers that share the work, while vertical scaling gives a single machine more CPU, memory, or storage. Autoscaling usually means horizontal scaling because it can happen without downtime.
Does autoscaling save money?
It can, because you pay for extra capacity only while you need it instead of sizing for the peak all day. Poorly tuned limits or scaling on the wrong metric can still waste money or leave the app short on capacity.
What is scale to zero?
Scale to zero means removing every instance when there is no traffic, so an idle app costs almost nothing. It is common on serverless platforms, with the trade-off of a cold start delay when the next request arrives.
See also
- KubernetesDevOps & Cloud, p. 32Kubernetes is an open-source system that automates deploying, scaling, and managing containerized applications across a cluster of machines.
- Load BalancerDevOps & Cloud, p. 34A load balancer is a server or service that spreads incoming traffic across several backend servers so no single one is overloaded and the app stays available.
- Cloud ComputingDevOps & Cloud, p. 10Cloud computing is the on-demand delivery of computing resources, such as servers, storage, and databases, over the internet with pay-as-you-go pricing.
- ServerlessDevOps & Cloud, p. 47Serverless is a cloud model in which the provider runs your code on demand, manages all the servers, scales automatically, and bills only for actual use.
- ScalabilitySoftware Architecture, p. 36Scalability is a system's ability to handle growing amounts of work, such as more users or data, by adding resources without a drop in performance.
- ContainerDevOps & Cloud, p. 12A container is a lightweight, isolated package that bundles an application with its dependencies and runs it on the host's shared operating system kernel.
Spotted a mistake or something missing on this page?Suggest an edit