Health Check
In short
A health check is a small automated test, usually an HTTP endpoint, that reports whether a service is up and able to handle requests, so failures show fast.
What is a health check?
A health check lets other systems ask a running service, 'are you OK?'. Most often it is a lightweight endpoint such as /healthz or /health that returns 200 OK when the service is healthy and an error status, such as 503 Service Unavailable, when it isn't. Load balancers, container orchestrators, and monitoring tools call it every few seconds and act on the answer automatically.
Health checks answer different questions. A liveness check asks whether the process is alive or stuck, for example in a deadlock, and a failure means 'restart me'. A readiness check asks whether the service can accept traffic right now, for example after it has loaded its configuration and connected to its database, and a failure means 'don't send me requests yet' without a restart. Kubernetes uses exactly these as liveness, readiness, and startup probes, and load balancers use health checks to take failing servers out of rotation until they recover.
A health check is like a nurse checking a patient's pulse at regular intervals: a quick, routine measurement that raises the alarm early, rather than a full medical exam. Good health checks are fast and cheap, since they run constantly, and they are usually left out of request logs and rate limits so they don't create noise.
A classic mistake is making the liveness check depend on the database or other services. If the database goes down briefly, every instance fails its liveness check and is restarted at the same moment, turning a small outage into a big one, so dependency checks belong in the readiness check, if anywhere. A health check is also much narrower than observability: it gives a simple yes-or-no signal for automation, while metrics, logs, and traces explain how well the system is performing and why.
Key takeaways
- A health check is a quick endpoint or command that reports whether a service is healthy.
- Load balancers and orchestrators call it regularly and reroute traffic or restart automatically.
- Liveness checks trigger restarts; readiness checks control whether traffic is sent.
- Keep liveness checks free of external dependencies to avoid mass restarts.
- Health checks must be fast and cheap, because they run constantly.
Example
containers:
- name: api
image: registry.example.com/api:1.4.2
livenessProbe: # failing -> the container is restarted
httpGet: { path: /livez, port: 8080 }
periodSeconds: 10
failureThreshold: 3
readinessProbe: # failing -> no traffic, but no restart
httpGet: { path: /readyz, port: 8080 }
periodSeconds: 5
startupProbe: # gives slow-starting apps time to boot
httpGet: { path: /livez, port: 8080 }
failureThreshold: 30Readers ask
What is the difference between a liveness check and a readiness check?
A liveness check tells the platform whether the process is stuck and should be restarted. A readiness check tells it whether the process can take traffic right now; failing it removes the instance from the load balancer without restarting it.
What should a health check endpoint return?
Return 200 when healthy and 503 when not, optionally with a small JSON body describing the status of each dependency. Keep detailed internal information off public endpoints, since it can help attackers.
Why is the endpoint often called /healthz?
The trailing z is a naming convention that came from Google's internal systems and spread through Kubernetes, meant to avoid clashing with real application routes. Kubernetes' own components now expose /livez and /readyz instead.
See also
- Load BalancerDevOps & Cloud, p. 34A load balancer is a server or service that spreads incoming traffic across several backend servers so no single one is overloaded and the app stays available.
- KubernetesDevOps & Cloud, p. 32Kubernetes is an open-source system that automates deploying, scaling, and managing containerized applications across a cluster of machines.
- PodDevOps & Cloud, p. 41A pod is the smallest deployable unit in Kubernetes: one or more containers that share a network address and storage and are scheduled together on one node.
- ObservabilityDevOps & Cloud, p. 38Observability is the ability to understand what is happening inside a running software system by collecting and analyzing its logs, metrics, and traces.
- High AvailabilitySoftware Architecture, p. 22High availability is the ability of a system to stay operational nearly all the time, mainly by removing single points of failure through redundancy.
- EndpointBackend & APIs, p. 12An endpoint is a specific URL, combined with an HTTP method, where an API receives requests and returns responses for one particular resource or action.
Spotted a mistake or something missing on this page?Suggest an edit