Prometheus
- Pronunciation
- pruh-MEE-thee-us
In short
Prometheus is an open-source monitoring system that collects metrics from apps and servers, stores them as time series and alerts when values cross a limit.
What is Prometheus?
Prometheus was created at SoundCloud in 2012 and became the second project, after Kubernetes, to graduate from the Cloud Native Computing Foundation. It answers questions such as how many requests per second a service handles, how long they take and how much memory it uses, by keeping a history of numbers over time.
It works by pulling: every few seconds Prometheus scrapes an HTTP endpoint, usually /metrics, on each target and stores the values it finds. Applications expose these metrics with client libraries, and ready-made exporters do it for systems such as Linux machines, databases and Nginx. Each series is identified by a name and labels, such as http_requests_total{method="GET", status="500"}.
Data is queried with PromQL, a language for calculations such as request rates, error percentages or the 99th percentile of response times. Alerting rules evaluate PromQL expressions, and the separate Alertmanager groups the resulting alerts and sends them to email, Slack or on-call tools. Grafana is the usual choice for drawing the data as dashboards.
A common misconception is that Prometheus stores logs or traces. It is built for numeric metrics; logs and traces need other tools, such as Loki and Jaeger. A single Prometheus server is also not meant for long-term, global storage, so large setups add systems such as Thanos or Mimir for that.
Key takeaways
- Prometheus is an open-source monitoring system and time-series database.
- It pulls metrics by scraping HTTP endpoints, usually /metrics.
- Each series has a name and labels; PromQL queries them.
- Alerting rules plus Alertmanager notify people when something is wrong.
- It handles metrics, not logs or traces.
Example
# prometheus.yml: scrape the app every 15 seconds
scrape_configs:
- job_name: "shop-api"
scrape_interval: 15s
static_configs:
- targets: ["api:8080"]
# PromQL: share of requests that failed over the last 5 minutes
# sum(rate(http_requests_total{status=~"5.."}[5m]))
# / sum(rate(http_requests_total[5m]))Readers ask
What is PromQL?
The Prometheus Query Language, used to select and calculate over time series, for example the per-second rate of requests or the 95th percentile of response times.
What is the difference between Prometheus and Grafana?
Prometheus collects and stores metrics and evaluates alerts. Grafana draws dashboards from data sources such as Prometheus. They are usually used together.
Does Prometheus push or pull metrics?
It pulls: Prometheus scrapes each target's metrics endpoint on a schedule. For short-lived jobs that can't be scraped, a Pushgateway lets them push their results instead.
See also
- MetricsDevOps & Cloud, p. 36Metrics are numeric measurements of a system collected over time, such as request rate, error rate and CPU usage, used for dashboards, alerts and planning.
- ObservabilityDevOps & Cloud, p. 38Observability is the ability to understand what is happening inside a running software system by collecting and analyzing its logs, metrics, and traces.
- Time-Series DatabaseDatabases, p. 46A time-series database is a database optimized for storing and querying timestamped measurements, such as sensor readings or server metrics, in time order.
- GrafanaDevOps & Cloud, p. 25Grafana is an open-source tool for building dashboards that turn metrics, logs and traces from many data sources into live charts and alerts in one place.
- SLODevOps & Cloud, p. 51An SLO is a measurable reliability target for a service, such as 99.9% of requests succeeding over 30 days, that tells a team how reliable is reliable enough.
- KubernetesDevOps & Cloud, p. 32Kubernetes is an open-source system that automates deploying, scaling, and managing containerized applications across a cluster of machines.
- OpenTelemetryDevOps & Cloud, p. 39OpenTelemetry is an open standard and set of tools for collecting traces, metrics and logs from software and sending them to any monitoring backend.
Spotted a mistake or something missing on this page?Suggest an edit