FREE Systems LESSON · Systems

Operate from evidence

Capacity, observability, and incident learning

Observability is the ability to answer causal questions from system evidence.

Metrics summarize known dimensions, logs preserve discrete events, and traces connect work across boundaries. None is useful merely because it exists. Begin with the user outcome and a hypothesis, then ask which signal distinguishes competing causes. Capacity models connect arrival rate, service time, concurrency, queue depth, and saturation so an alarm can point toward a mechanism rather than announce pain.

Collect evidence that changes a decision; cardinality and volume are costs, not signs of maturity.

An alert can be accurate and still be useless.

Thresholds on every internal metric create noise, fragmented ownership, and alert fatigue. Page on urgent user-impacting conditions with an actionable response; ticket slower risks; retain diagnostic signals for investigation. During incidents, preserve a timeline of observations, decisions, and state changes so the review can improve the system rather than reward hindsight.

The purpose of an alert is a timely human or automated decision.
Practise this lesson free →