OpenTelemetry
- Pronunciation
- OH-pen tuh-LEM-uh-tree
In short
OpenTelemetry is an open standard and set of tools for collecting traces, metrics and logs from software and sending them to any monitoring backend.
What is OpenTelemetry?
To understand a running system you need telemetry: traces that follow a request across services, metrics such as request counts and latency, and logs. OpenTelemetry, often shortened to OTel, gives each language the same API and SDK for producing that data, plus a common protocol, OTLP, for sending it. It is a Cloud Native Computing Foundation project, formed in 2019 by merging two earlier efforts, OpenTracing and OpenCensus.
Its main promise is independence. You instrument your code once, and the data can go to Jaeger, Prometheus, Grafana or a commercial service by changing configuration rather than code. Many libraries and frameworks can be instrumented automatically, so HTTP calls and database queries show up as spans without extra work. The OpenTelemetry Collector, a separate program, can receive, filter, enrich and forward the data on its way.
Tracing works by passing context along. When one service calls another, it adds a traceparent header, defined by the W3C Trace Context standard, so the next service attaches its spans to the same trace. The result is a timeline of one request across every service it touched, showing where the time went or where an error began. OpenTelemetry covers producing and moving telemetry; storing it and drawing dashboards is left to the backend.
Key takeaways
- OpenTelemetry is a vendor-neutral standard for traces, metrics and logs.
- Code is instrumented once, and the data can go to any backend.
- OTLP is its protocol; the Collector receives, processes and forwards data.
- Trace context travels between services in the
traceparentheader.
Example
from opentelemetry import trace
tracer = trace.get_tracer("checkout")
def place_order(order):
with tracer.start_as_current_span("place_order") as span:
span.set_attribute("order.items", len(order.items))
charge(order) # spans inside become children of this one
send_confirmation(order)Readers ask
Is OpenTelemetry a monitoring tool?
Not by itself. It produces and moves telemetry but doesn't store it or draw dashboards; for that the data goes to a backend, open source or commercial. That split is the point: you can change the backend without touching the code.
What is the difference between OpenTelemetry and Prometheus?
Prometheus is a monitoring system: it collects metrics, stores them and lets you query them and alert on them. OpenTelemetry is a standard for producing traces, metrics and logs. They work together, and OpenTelemetry metrics can be sent to Prometheus.
See also
- ObservabilityDevOps & Cloud, p. 38Observability is the ability to understand what is happening inside a running software system by collecting and analyzing its logs, metrics, and traces.
- Distributed TracingDevOps & Cloud, p. 15Distributed tracing is a technique that follows a single request as it travels through many services, recording how long each step took and where it failed.
- MetricsDevOps & Cloud, p. 36Metrics are numeric measurements of a system collected over time, such as request rate, error rate and CPU usage, used for dashboards, alerts and planning.
- LoggingDevOps & Cloud, p. 35Logging is the practice of recording timestamped messages about events in a running program, such as errors and requests, so people can investigate them later.
- PrometheusDevOps & Cloud, p. 43Prometheus is an open-source monitoring system that collects metrics from apps and servers, stores them as time series and alerts when values cross a limit.
- MicroservicesSoftware Architecture, p. 27Microservices are an architectural style where an application is split into small, independently deployable services that communicate over a network.
Sources
Spotted a mistake or something missing on this page?Suggest an edit