The dashboard nobody looks at
Most dashboards are built during an incident and never opened again. A few are worth keeping, and they look different.
tech, developers, and the code underneath
23 pieces tagged reliability,
from March 10, 2025 to September 28, 2026.
Most dashboards are built during an incident and never opened again. A few are worth keeping, and they look different.
Your service dependency graph exists whether or not anyone has written it down. Writing it down takes an afternoon.
A request crosses a dozen components with a timeout each, and almost nobody has ever added them up.
Every fixed-size resource is a queue. Most of them are unmonitored, and that is where latency hides.
The chat log is where the incident is actually run. Four conventions make it useful rather than a wall of noise.
Error rates get tracked as a percentage. Users experience them as a specific person having a specific bad day.
When a fast producer meets a slow consumer, something has to give. Deciding what, in advance, is the whole discipline.
The most common way a small incident becomes a large one is a retry policy written without thinking about aggregate behavior.
It has no type system, no tests, no review, and it causes a disproportionate share of outages.
The bug that happens once a week and never in staging. A systematic approach to the class of problem everyone handles badly.
Most tracing deployments produce beautiful waterfalls nobody opens. The difference is three implementation details.
A month of synchronized global demand across three countries, sixteen cities, and every streaming platform at once.