Skip to main content

Observability

Updated 2 min read

Share this page

Send the link, quote the definition with a link back, or show it as a card on your own site.

https://softwaredictionary.org/terms/observability

In short

Observability is the ability to understand what is happening inside a running software system by collecting and analyzing its logs, metrics, and traces.

What is observability?

Observability describes how well you can understand the internal state of a system from the data it produces. In a well-observed system, when something goes wrong, such as a slow checkout page or a sudden spike in errors, engineers can find out why by querying that data, without shipping new code just to investigate. The term comes from control theory and has become central to running distributed systems such as microservices.

Observability is usually built on three kinds of telemetry, often called the three pillars. Logs are timestamped records of individual events, metrics are numbers measured over time, such as requests per second, error rate, or memory usage, and traces follow a single request as it travels through many services, showing where the time was spent. OpenTelemetry is the widely adopted open standard for producing and collecting this data, which can then be sent to many different storage and dashboard tools.

A car makes a helpful comparison: a warning light on the dashboard tells you that something is wrong, while a mechanic's diagnostic tool that reads every sensor helps you find out exactly what and why. Teams typically turn telemetry into dashboards, alerts, and service level objectives (SLOs), which define how reliable a service needs to be for its users.

Observability is often used interchangeably with monitoring, but they are different. Monitoring watches for known problems using predefined checks and dashboards, answering questions you thought of in advance, while observability lets you investigate new, unexpected problems you did not predict. Monitoring is one part of observability, not a replacement for it.

At a glance

A checkout service sends out three kinds of telemetry: metrics show its slowest responses jumping to 900 ms, a trace of one slow request shows most of the time spent in the payment call, and the log line for it says the payment provider timed out.CheckoutserviceMetrics · p95 response time900 msTrace · one slow requestcheckoutpaymentdatabaseLogs12:04:31 ERROR payment provider timed out (750 ms)
Metrics say that something is wrong, traces say where, and logs say why. Together they answer questions nobody planned a dashboard for.

Key takeaways

  • Observability means understanding a system's internal state from the data it emits.
  • Logs, metrics, and traces are the three main types of telemetry.
  • Distributed tracing follows one request across many services.
  • Monitoring answers known questions; observability helps investigate unknown ones.
  • OpenTelemetry is the common open standard for collecting telemetry.

Example

Adding a trace span with OpenTelemetry (Node.js)javascript
import { trace } from "@opentelemetry/api";

const tracer = trace.getTracer("checkout-service");

export async function checkout(order) {
  // A span measures this step and links it to the rest of the request's trace
  return tracer.startActiveSpan("checkout", async (span) => {
    span.setAttribute("order.item_count", order.items.length);
    try {
      return await chargeCard(order);
    } finally {
      span.end(); // the finished span is exported to your tracing backend
    }
  });
}

Readers ask

What is the difference between observability and monitoring?

Monitoring tracks predefined metrics and alerts you when known problems occur, such as high CPU usage or a failing health check. Observability is the broader ability to explore detailed telemetry and understand problems you did not anticipate, and monitoring is one part of it.

What are the three pillars of observability?

The three pillars are logs, metrics, and traces. Logs record individual events, metrics summarize measurements over time, and traces show the path and timing of a single request through a system; many teams now add continuous profiling as a fourth signal.

What is OpenTelemetry?

OpenTelemetry, often shortened to OTel, is an open-source, vendor-neutral standard and set of SDKs for producing and collecting logs, metrics, and traces. Because it is a standard, you can switch storage and dashboard tools without rewriting your instrumentation.

See also

Spotted a mistake or something missing on this page?Suggest an edit

Read a random page
Open today's review
Switch to the dark theme
Read this page in Türkçe

More

Settings