Metrics
- In Turkish
- Metrikler
In short
Metrics are numeric measurements of a system collected over time, such as request rate, error rate and CPU usage, used for dashboards, alerts and planning.
What are metrics in software monitoring?
In software operations, metrics are numbers that describe the state or behavior of a system, recorded at regular intervals. Examples include requests per second, the percentage of failed requests, response time, memory usage, and the number of jobs waiting in a queue. Because each data point is just a name, a value, a timestamp, and a few labels, metrics are cheap to store and fast to query, even over months of history.
Most metrics fall into a few types: a counter only goes up, such as total requests served; a gauge goes up and down, such as current memory use; and a histogram sorts measurements into buckets, such as how many requests took under 100, 250, or 500 milliseconds, which lets you calculate percentiles like p95 and p99. Applications expose metrics through a library, and a monitoring system collects them every few seconds, often by scraping an HTTP endpoint, and stores them in a time-series database. Labels such as route or status let you slice a metric, but each unique combination creates a new series, so labels with unbounded values like user IDs, known as high cardinality, must be avoided.
Metrics are like the gauges on a car's dashboard: speed, fuel, and engine temperature tell you at a glance whether things are normal, without describing every event. They power dashboards, alerts, autoscaling decisions, and SLOs. A popular starting set is the four golden signals from site reliability engineering: latency, traffic, errors, and saturation, meaning how full a resource is.
Metrics are often confused with logs. A log line records one event with rich detail, while a metric aggregates many events into a number, so a metric tells you that errors jumped at 14:02 and logs or traces tell you why. Averages can also mislead: a mean response time of 200 milliseconds can hide that 1% of users wait five seconds, which is why teams track percentiles.
Key takeaways
- Metrics are numeric measurements recorded over time and stored as time series.
- The common types are counters, gauges, and histograms.
- Percentiles such as p95 and p99 reveal slow requests that averages hide.
- Labels add dimensions, but high-cardinality labels like user IDs cause cost problems.
- Metrics show that something changed; logs and traces help explain why.
Example
from prometheus_client import Counter, Histogram, start_http_server
# A counter only goes up; labels let you slice it by route and status
REQUESTS = Counter("http_requests_total", "Requests served", ["route", "status"])
# A histogram sorts durations into buckets so percentiles can be calculated
LATENCY = Histogram("http_request_duration_seconds", "Request duration", ["route"])
start_http_server(9100) # serves the metrics for the monitoring system to scrape
def handle_checkout(request):
with LATENCY.labels(route="/checkout").time():
response = process(request)
REQUESTS.labels(route="/checkout", status=response.status).inc()
return responseReaders ask
What is the difference between a counter and a gauge?
A counter is a value that only increases, such as the total number of requests, and is usually viewed as a rate per second. A gauge is a value that can go up or down, such as current memory usage or the number of active connections.
What are the four golden signals?
They are latency, traffic, errors, and saturation, a set of metrics recommended in site reliability engineering for monitoring any user-facing service. Together they show how fast the service is, how much it is used, how often it fails, and how close it is to its limits.
What does p99 latency mean?
p99 latency is the response time that 99% of requests are faster than, so only the slowest 1% take longer. It shows what users in the long tail experience, which an average hides.
See also
- ObservabilityDevOps & Cloud, p. 38Observability is the ability to understand what is happening inside a running software system by collecting and analyzing its logs, metrics, and traces.
- LoggingDevOps & Cloud, p. 35Logging is the practice of recording timestamped messages about events in a running program, such as errors and requests, so people can investigate them later.
- Distributed TracingDevOps & Cloud, p. 15Distributed tracing is a technique that follows a single request as it travels through many services, recording how long each step took and where it failed.
- SLODevOps & Cloud, p. 51An SLO is a measurable reliability target for a service, such as 99.9% of requests succeeding over 30 days, that tells a team how reliable is reliable enough.
- Time-Series DatabaseDatabases, p. 46A time-series database is a database optimized for storing and querying timestamped measurements, such as sensor readings or server metrics, in time order.
- AutoscalingDevOps & Cloud, p. 2Autoscaling is the automatic adding or removing of computing resources, such as servers or containers, based on demand to keep performance steady and costs low.
- PrometheusDevOps & Cloud, p. 43Prometheus is an open-source monitoring system that collects metrics from apps and servers, stores them as time series and alerts when values cross a limit.
- OpenTelemetryDevOps & Cloud, p. 39OpenTelemetry is an open standard and set of tools for collecting traces, metrics and logs from software and sending them to any monitoring backend.
Spotted a mistake or something missing on this page?Suggest an edit