DevOps
Lesson 3 of 8About 3 min readSuggest an edit

Observability

Monitoring tells you that something is wrong. Observability is being able to ask why from the outside, without shipping new code to find out. It rests on three kinds of telemetry.

Logs, metrics and traces

Logs are timestamped records of events. Write them as structured data rather than prose, so you can filter and aggregate:

{"level":"error","msg":"payment failed","order_id":9001,"provider":"stripe","duration_ms":812,"trace_id":"4bf92f35"}

Never log passwords, tokens or full card numbers, and be careful with personal data.

Metrics are numbers aggregated over time: request rate, error count, latency percentiles, queue depth, memory. They are cheap to store and fast to query, which makes them ideal for dashboards and alerts. Watch the labels: a label with unbounded values, such as user id, creates a separate time series for every value and can overwhelm a metrics system.

Traces follow one request across services. Each step is a span with a start time, a duration and a parent. A trace shows that a slow checkout spent 40 ms in your API and 1.9 s waiting on the inventory service. Put the trace id into your logs to jump from a slow trace straight to its log lines.

OpenTelemetry is the vendor-neutral standard for producing all three.

The four golden signals

For any user-facing service, start by measuring:

  1. Latency: how long requests take. Track percentiles (p50, p95, p99), not averages: an average hides the slowest one percent of users. Measure failed requests separately, since fast errors make latency look better than it is.
  2. Traffic: how much demand the service is getting, such as requests per second.
  3. Errors: the rate of failed requests, including “successful” responses with wrong content.
  4. Saturation: how full the service is: CPU, memory, connection pools, queue length.

SLIs, SLOs and error budgets

  • An SLI (service level indicator) measures user experience, for example “the share of requests that succeed in under 300 ms”.
  • An SLO (objective) is the target for it: “99.9% over 30 days”.
  • The error budget is what’s left: 0.1% of requests, roughly 43 minutes of full downtime a month.

The budget turns reliability into a decision rule. While budget remains, ship features. When it is spent, prioritise reliability work. No SLO should be 100%: users can’t tell 99.99% from 100%, but chasing it costs a great deal.

Alert on symptoms, not causes

Page a human only when users are affected or about to be, and only when a human needs to act:

  • Good: “Checkout error rate above 2% for 5 minutes” or “error budget burning 10x faster than sustainable”.
  • Noisy: “CPU at 85% on one host”. High CPU that hurts nobody is a dashboard item or a ticket, not a 3 a.m. page.

Every alert should link to a runbook that says what to check first. Alerts that fire often and get ignored train people to ignore all alerts; delete or fix them.

Dashboards that answer questions

Build one overview per service with the golden signals, then drill-down views. Put deploy markers on the graphs: the first question in most incidents is “what changed?”, and it’s usually the last deploy.

Next: Infrastructure as code

Declarative infrastructure, the plan and apply workflow, state, modules, drift and secrets.