Chaos Engineering
- Pronunciation
- KAY-os en-jih-NEER-ing
In short
Chaos engineering is the practice of deliberately injecting failures into a system, such as crashing servers, to confirm that it keeps working as expected.
What is chaos engineering?
Chaos engineering is the discipline of running controlled experiments that break parts of a system on purpose to find weaknesses before they cause real outages. Typical experiments shut down a server, kill containers, add network latency, fill up a disk, or make a dependency return errors. The practice became well known in the early 2010s, when Netflix began randomly terminating its own production servers to make sure its streaming service could survive the loss.
Each experiment follows a scientific method. First define the steady state, meaning measurable normal behavior such as the checkout success rate; then form a hypothesis, for example that the site stays within its SLO if one database replica fails; then inject the failure and compare the results. Experiments start small, in a test environment or with a limited blast radius, and they have an abort switch that stops the experiment immediately if users are affected.
It is like a fire drill: instead of waiting for a real fire to discover that an exit is blocked, you practice under controlled conditions and fix what goes wrong. Chaos engineering is most useful for distributed systems, such as microservices on Kubernetes, where failures happen constantly and interactions are too complex to reason about fully. Many teams also run game days, scheduled sessions where engineers inject failures together and practice their incident response.
Chaos engineering is often confused with load testing and stress testing. Load testing checks how a system performs under expected traffic, and stress testing pushes it past its limits, while chaos engineering keeps traffic normal and breaks components to test resilience: fault tolerance, failover, retries, timeouts, and alerting. The aim is not chaos for its own sake but confidence that the system handles failure gracefully.
Key takeaways
- Chaos engineering injects real failures on purpose to reveal hidden weaknesses.
- Each experiment starts from a measurable steady state and a clear hypothesis.
- Experiments keep the blast radius small and can be aborted at any time.
- It tests resilience features such as failover, retries, timeouts, and alerts.
- Load testing adds traffic; chaos engineering breaks components under normal traffic.
Example
# Hypothesis: the web service stays healthy if one of its pods dies
kubectl get pods -l app=web
# Inject the failure: delete one pod at random
kubectl delete "$(kubectl get pods -l app=web -o name | shuf -n 1)"
# Add 200 ms of network latency on a test server (Linux)
sudo tc qdisc add dev eth0 root netem delay 200ms
# Watch error rates and latency, then remove the injected delay
sudo tc qdisc del dev eth0 root netemReaders ask
Is chaos engineering done in production?
Mature teams do run experiments in production, because only production has real traffic and real configurations. They start in test environments, limit the blast radius to a small share of users or servers, and stop automatically if key metrics degrade.
What is the difference between chaos engineering and testing?
A test checks a known condition and either passes or fails. Chaos engineering is an experiment that explores how a complex system behaves when something breaks, and it often reveals problems nobody thought to write a test for.
What is a game day?
A game day is a planned session where a team deliberately causes failures, such as shutting down a whole availability zone, and practices detecting and recovering from them. It tests both the system and the team's incident response.
See also
- Fault ToleranceSoftware Architecture, p. 20Fault tolerance is the ability of a system to keep working correctly, perhaps at reduced capacity, when some of its hardware or software components fail.
- Site Reliability EngineeringDevOps & Cloud, p. 49Site reliability engineering is a discipline that applies software engineering to operations, keeping services reliable with automation and measurable targets.
- Stress TestingTesting & Quality, p. 27Stress testing pushes a system beyond its expected workload on purpose to find its breaking point and to check that it fails gracefully and recovers afterward.
- Load TestingTesting & Quality, p. 15Load testing is a type of performance testing that simulates many users or requests at once to measure how a system behaves under expected traffic.
- MicroservicesSoftware Architecture, p. 27Microservices are an architectural style where an application is split into small, independently deployable services that communicate over a network.
- High AvailabilitySoftware Architecture, p. 22High availability is the ability of a system to stay operational nearly all the time, mainly by removing single points of failure through redundancy.
Spotted a mistake or something missing on this page?Suggest an edit