Distributed Tracing
- In Turkish
- Dağıtık İzleme
In short
Distributed tracing is a technique that follows a single request as it travels through many services, recording how long each step took and where it failed.
What is distributed tracing?
In a system made of many services, one user action, such as placing an order, can trigger calls to a gateway, an order service, a payment service, a database, and a message queue. Distributed tracing records that whole journey as a single trace, so you can see every step in order, how long each one took, and which one failed. It answers questions that are very hard to answer from logs alone, such as why this one request was slow.
A trace is made of spans. Each span represents one unit of work, such as an HTTP call or a database query, and records its start time, duration, status, and attributes; spans point to a parent span, forming a tree. When a service calls another, it passes the trace ID and its current span ID along with the request, usually in the standard W3C traceparent HTTP header, so the next service can attach its spans to the same trace. Tracing libraries, most commonly OpenTelemetry, create spans automatically for popular frameworks and export them to a tracing backend, which draws each trace as a timeline.
It works like a parcel tracking number: every depot scans the parcel, and afterward you can see exactly where it went and where it sat for three days. Tracing is essential in microservices and serverless systems, where no single service's logs show the full story. Because recording every request is expensive at high traffic, systems usually use sampling, for example keeping 1% of normal traces but every trace that contains an error.
Distributed tracing is often confused with logging. A log is an isolated record written by one service, while a trace connects work across services through a shared trace ID and shows timing and cause and effect. The two work best together: adding the trace ID to every log line lets you jump from a slow span straight to the matching log entries.
Key takeaways
- A trace follows one request end to end across services.
- Traces are made of spans, each measuring one operation, linked in parent-child order.
- A trace ID is passed between services, usually in the
traceparentheader. - OpenTelemetry is the common open standard for creating and exporting traces.
- Sampling keeps tracing affordable at high traffic volumes.
Example
// W3C trace context header: version-traceId-parentSpanId-flags
// traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
import { context, propagation } from "@opentelemetry/api";
async function callPaymentService(order) {
const headers = { "content-type": "application/json" };
// Copy the current trace ID and span ID into the outgoing headers
propagation.inject(context.active(), headers);
return fetch("https://payments.internal/charge", {
method: "POST",
headers,
body: JSON.stringify(order),
});
}Readers ask
What is a span in distributed tracing?
A span is one timed operation within a trace, such as an incoming HTTP request, a database query, or a call to another service. It records a start time, a duration, a status, and attributes, and it points to its parent span so the whole request can be shown as a tree.
What is the difference between tracing and logging?
Logging records individual events inside one service. Tracing links the work done by many services for a single request through a shared trace ID, showing the order, timing, and cause of each step.
Does distributed tracing slow down an application?
Creating spans adds a small overhead, usually negligible compared with network calls. The bigger cost is sending and storing trace data, which is why most systems sample only a portion of requests.
See also
- ObservabilityDevOps & Cloud, p. 38Observability is the ability to understand what is happening inside a running software system by collecting and analyzing its logs, metrics, and traces.
- LoggingDevOps & Cloud, p. 35Logging is the practice of recording timestamped messages about events in a running program, such as errors and requests, so people can investigate them later.
- MetricsDevOps & Cloud, p. 36Metrics are numeric measurements of a system collected over time, such as request rate, error rate and CPU usage, used for dashboards, alerts and planning.
- MicroservicesSoftware Architecture, p. 27Microservices are an architectural style where an application is split into small, independently deployable services that communicate over a network.
- LatencyNetworking, p. 14Latency is the delay between sending a request and the start of a response, usually measured in milliseconds, and it shapes how responsive an app feels.
- Service MeshDevOps & Cloud, p. 48A service mesh is an infrastructure layer that manages traffic between microservices, adding encryption, retries, routing, and monitoring without code changes.
- OpenTelemetryDevOps & Cloud, p. 39OpenTelemetry is an open standard and set of tools for collecting traces, metrics and logs from software and sending them to any monitoring backend.
Spotted a mistake or something missing on this page?Suggest an edit