Skip to content

Observability Is About Understanding, Not Monitoring

Dashboards that count errors are not observability. You need to ask new questions about unknown failures without shipping new code, which requires intentional signals and ownership of meaning.

Production Notes #32 · Part of Binary and Beyond. LinkedIn newsletter edition follows.

Most teams own monitoring.

Fewer own understanding.

The difference shows up at 2 a.m., when the dashboard is green-ish, the error rate is "within budget," and customers are still stuck in a state the product cannot explain. Someone opens the same six charts. Someone else asks for a new log line. The incident becomes a feature request for telemetry.

That is monitoring doing its job and still failing the business.

The dashboard that counted the wrong thing

Picture a mid-market payments platform with a proud observability stack: APM, log search, synthetic checks, on-call rotations, and a wall of latency percentiles.

One quarter, a partner integration started failing intermittently. Retries masked the shape. Success rate stayed high enough that paging never fired. Support tickets rose about "pending forever" payments. Engineering saw elevated 5xx on one dependency, then a quiet recovery. No war room. No root cause. The next week it happened again with a different partner and the same story.

The team added another panel: partner error count by code. The panel filled. Meaning did not. Status 92 meant timeout to one engineer, means declined to another, and meant "try later" in a vendor PDF nobody had read since onboarding. The dashboard answered "how many." It could not answer "what did we promise the merchant, and is that promise still true?" (Why Retries Make Bugs Worse, Why Race Conditions Are Business Problems).

They were monitoring infrastructure health. They were not observing business truth under partial failure.

The expensive part was not the missing panel. It was the missing ability to reconstruct a single payment's story across queue, partner, and ledger without convening three teams and guessing which retry won. Understanding failed before tooling did.

What monitoring optimises for (and where it stops)

Monitoring is excellent at known unknowns you already decided matter:

  • Is the process up?
  • Is latency past a threshold?
  • Is error rate above a page line?
  • Did the synthetic login succeed?

Those questions are necessary. They are also closed questions. You write the question when you ship the alert. When production invents a new failure mode, the closed set goes quiet while the business bleeds through an uninstrumented seam.

Observability, used carefully, is the ability to ask a new question about a live system without shipping new code first. That ability is not a product SKU. It is a property of the signals you chose to emit, the context you attached, and the ownership of what those signals mean.

If every novel incident requires a deploy to "add logging," you have monitoring plus hope.

Why understanding needs intentional signals

You cannot query what you never recorded. You also cannot trust what you recorded without semantics.

Intentional signals look like:

Correlation that survives hops. A request, workflow, or payment id that travels through queues, partners, and retries so you can reconstruct a story instead of a pile of timestamps.

Business state, not only HTTP status. Pending, authorised, settled, compensating, failed-with-reason. The states operators already argue about in Slack (Building Systems That Recover Instead of Restart).

Cardinality with discipline. High-cardinality fields that answer "which merchant, which partner, which region" without turning every log into a cost incident.

Owned meanings. A status code glossary that lives next to the code path that emits it, with an owner, not a wiki that drifted two years ago.

Failure shape, not only failure count. Timeouts versus declines versus duplicate accepts. Retries that succeeded late versus retries that doubled a side effect.

Without those, dashboards become aquariums: soothing motion, little explanatory power. Teams then overfit to the metrics they have. They tune for green panels while customers experience a stuck workflow the panels never named.

There is a second failure mode: drowning in signals without owners. High-cardinality traces with empty attributes, unstructured logs that contradict each other, and three competing definitions of "failed" in the same index. Volume is not understanding. Noise is a tax on the on-call brain. Intentional signals are fewer, richer, and contested less often because someone owns what they claim.

A mental model: questions you can ask cold

Observability is not "more graphs." It is whether a stranger on call can form and answer questions like:

  • For this customer id, what was the last committed business state across our system and the partner?
  • Which retries fired, with which keys, and what did each attempt believe was true?
  • Did we compensate, or did we leave a half-promise?
  • Which dependency degraded first, and which symptoms were downstream echoes?

If answering requires a new deploy, a tribal expert, or reading three Slack channels from last Tuesday, the system is not observable in the sense that matters for production. It is instrumented for the failures you predicted.

Monitoring asks: "Did we breach a threshold we defined?"

Understanding asks: "What is true right now for this workflow, and how did we get here?"

Both are needed. Confusing them produces expensive tools and thin confidence.

Implications for teams and roadmaps

Treat signal design as product work. When you ship a new workflow, ship the questions you will need on day thirty, not only the happy path metrics. Budget for correlation ids, state transitions, and reason codes the way you budget for the feature itself.

Separate SLOs that protect user journeys from SLIs that protect hosts. A healthy CPU with a stuck settlement queue is not a healthy product. Page on journey breakage when you can detect it. Keep host metrics as diagnostic depth, not as the definition of success.

Make meaning ownership explicit. If three teams emit status=failed, someone must own the vocabulary. Unowned vocabulary is how dashboards lie politely.

On legacy modernisation programmes, we treat observability as part of the cutover contract: can operators ask new questions about the hybrid estate without inventing temporary log packs for every incident? Dual-run without shared correlation is dual blindness.

Refuse vanity coverage. One hundred percent trace sampling with empty attributes is theatre. Ten carefully chosen events with durable context will out-diagnose a forest of empty spans.

Practice the cold question in rehearsal, not only in incidents. Once a month, pick a random workflow id from production and ask an engineer who did not build it to narrate the business state from signals alone. Where they stall is your observability backlog. That exercise is cheaper than discovering the same gap during a customer-visible outage.

Also separate detective work from product truth. Logs that help an engineer debug a null pointer are useful. They are not the same as events that tell support whether a merchant is owed a retry, a refund, or an apology. If support still opens tickets to learn what the system already "knows," your signals serve developers and abandon operators.

An observability readiness checklist

  1. Can on-call answer "what is the business state of this id?" from signals alone, without a code change?
  2. Do retries, queues, and partner calls carry a stable correlation key end to end?
  3. Are status and reason codes owned, documented next to emitters, and stable enough to query historically?
  4. Which alerts fire on journey failure versus host discomfort, and do pages match customer pain?
  5. For the last three novel incidents, did understanding require a deploy for more logs?
  6. Who owns the meaning of the top twenty signals that drive decisions?
  7. What is the cheapest signal you could add that would have shortened the last war room by half?

If most answers fail, buy fewer panels and invent better questions. Tooling amplifies signal quality. It does not create it.

Fingerprint of a team that understands production

You will notice it in the first fifteen minutes of an incident. Someone pulls a single id and walks the story: attempt, partner response, local commit, retry, compensation. The room argues about the business state, not about which dashboard tab to open. Afterward, the follow-up is a sharper signal or a clearer contract, not another red/green tile.

Monitoring keeps the lights checked. Observability is the practice of remaining curious under failure without waiting for the next release train. That curiosity is engineered: intentional signals, owned meanings, and the humility to admit that the next failure will not match last quarter's alert rules.

Count errors if you must. Just do not confuse the count with understanding.

Related reading

Production Notes #32 · Part of Binary and Beyond. LinkedIn newsletter edition follows. Need signals that explain stuck workflows, not just green dashboards? Start a conversation.

Working through a problem like this on a live system? Start a conversation