Skip to main content

Autoscaling

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/autoscaling

In short

Autoscaling is the automatic adding or removing of computing resources, such as servers or containers, based on demand to keep performance steady and costs low.

What is autoscaling?

Autoscaling is a cloud and container feature that automatically changes how much computing capacity an application has. When traffic rises, it adds more servers, virtual machines, or containers; when traffic falls, it removes them. The goal is to keep the app responsive during peaks without paying for idle machines the rest of the time.

An autoscaler watches metrics such as CPU usage, memory, request rate, or the length of a message queue, and compares them with a target, for example 70% average CPU. If the metric stays above the target, it starts new instances, usually behind a load balancer; if it stays below, it shuts some down, always within a configured minimum and maximum. Some systems also use scheduled scaling for predictable peaks, or predictive scaling that forecasts demand from past patterns.

Think of a supermarket that opens more checkout lanes when the lines grow long and closes them when the store is quiet. Autoscaling is used by web applications with daily traffic cycles, by background workers that process queues, and by Kubernetes clusters, which can scale both the number of pods and the number of nodes they run on.

Autoscaling is usually horizontal, meaning it adds more copies of the application, which is different from vertical scaling, where a single machine gets more CPU or memory. Horizontal autoscaling works best when the application is stateless, so any instance can handle any request. It is also not instant: new instances take time to start, so sudden spikes may still need spare capacity or rate limiting.

Key takeaways

  • Autoscaling adds or removes capacity automatically based on demand.
  • It is driven by metrics such as CPU, memory, request rate, or queue length.
  • Minimum and maximum limits keep scaling safe and costs predictable.
  • Horizontal scaling adds instances, while vertical scaling makes one instance bigger.
  • Stateless applications are the easiest to autoscale.

Example

A Kubernetes Horizontal Pod Autoscaleryaml
# Keep average CPU near 70% by running between 2 and 10 pods
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: web-app
spec:
  scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: web-app }
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target: { type: Utilization, averageUtilization: 70 }

Readers ask

What is the difference between horizontal and vertical scaling?

Horizontal scaling adds more machines or containers that share the work, while vertical scaling gives a single machine more CPU, memory, or storage. Autoscaling usually means horizontal scaling because it can happen without downtime.

Does autoscaling save money?

It can, because you pay for extra capacity only while you need it instead of sizing for the peak all day. Poorly tuned limits or scaling on the wrong metric can still waste money or leave the app short on capacity.

What is scale to zero?

Scale to zero means removing every instance when there is no traffic, so an idle app costs almost nothing. It is common on serverless platforms, with the trade-off of a cold start delay when the next request arrives.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings