<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — reliability</title>
<link>https://readme.news/tags/reliability/</link>
<atom:link href="https://readme.news/tags/reliability/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged reliability.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>The dashboard nobody looks at</title><link>https://readme.news/the-dashboard-nobody-looks-at/</link><guid isPermaLink="true">https://readme.news/the-dashboard-nobody-looks-at/</guid><pubDate>Mon, 28 Sep 2026 09:00:00 +0000</pubDate><description>Most dashboards are built during an incident and never opened again. A few are worth keeping, and they look different.</description><content:encoded><![CDATA[<p>Open your monitoring tool and count the dashboards. Now check when each was last viewed — most tools will tell you.</p>
<p>The distribution is always the same: three dashboards carry all the traffic and forty have not been opened in a year.</p>
<h2 id="why-they-accumulate">why they accumulate<a class="anchor" href="#why-they-accumulate" aria-label="link to this section">#</a></h2>
<p>Dashboards get created during incidents. Someone needs to see a specific correlation right now, builds a view, the incident ends, and the dashboard stays — specific to a problem that has been fixed, named something like "orders debug 2," owned by nobody.</p>
<p>Nothing removes them, so they accumulate until finding the useful one requires knowing its name in advance.</p>
<h2 id="the-three-that-earn-their-place">the three that earn their place<a class="anchor" href="#the-three-that-earn-their-place" aria-label="link to this section">#</a></h2>
<p><strong>The service dashboard.</strong> One per service, same layout for every service, so that anyone can read any service's dashboard without orientation. Rate, errors, duration, saturation. Six to eight panels, above the fold, no scrolling.</p>
<p>The consistency is the point. When every service looks the same, an engineer responding to an unfamiliar service already knows where to look.</p>
<p><strong>The journey dashboard.</strong> One per critical user journey — checkout, signup, search. Not per service: per outcome, end to end, across every service involved.</p>
<p>This is the one most organisations lack, and it is the one that answers the question that matters during an incident: <em>are users able to do the thing?</em> Every service being green while checkout is broken is a common and confusing state, and only a journey view shows it.</p>
<p><strong>The capacity dashboard.</strong> Resource headroom over weeks, not minutes. Reviewed on a schedule, not during incidents. This is the only one meant for planning rather than response.</p>
<h2 id="what-makes-a-panel-useful">what makes a panel useful<a class="anchor" href="#what-makes-a-panel-useful" aria-label="link to this section">#</a></h2>
<p><strong>A threshold drawn on it.</strong> A line with no reference is a shape. A line with "target: 200ms" drawn across it is information. Every latency and error panel should show what "fine" is.</p>
<p><strong>Comparison to last week.</strong> Most values are meaningless in isolation. 4,000 requests per minute — is that high? The same panel with last week's line overlaid answers instantly.</p>
<p><strong>A unit and a direction.</strong> Label it. State whether up is good. This sounds trivial and half of all panels fail it.</p>
<p><strong>Percentiles, not averages,</strong> for anything latency-shaped. And show p50 next to p99: the gap between them is often the real signal.</p>
<p><strong>Fewer panels.</strong> A dashboard with forty panels is a data dump. During an incident nobody scrolls. Eight panels that fit on a screen beat forty that do not.</p>
<h2 id="the-annotations-that-make-it-work">the annotations that make it work<a class="anchor" href="#the-annotations-that-make-it-work" aria-label="link to this section">#</a></h2>
<p>The highest-value dashboard feature is the least used: <strong>deploy markers</strong>.</p>
<p>Vertical lines on every time series showing when a deploy happened. Most tools support it and most teams have not wired it up.</p>
<p>With them, "did this start after the deploy?" is answered by looking. Without them it is answered by opening another tab, finding the deploy log, and correlating timestamps by hand — during an incident, under pressure.</p>
<p>Add config changes and <a class="xref" href="/the-graph-nobody-drew/" title="The graph nobody drew">feature flag</a> flips to the same annotation stream and you have covered the causes of most self-inflicted incidents.</p>
<h2 id="the-maintenance-rule">the maintenance rule<a class="anchor" href="#the-maintenance-rule" aria-label="link to this section">#</a></h2>
<p><strong>Delete dashboards nobody opens.</strong> Quarterly, using the view counts your tool already records. If someone misses one, they can rebuild it, and the rebuild will be better because it will reflect the current system.</p>
<p><strong>Give every dashboard an owner and a purpose in its description.</strong> "Used by <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> to triage checkout failures" tells the next person whether to trust it. An unlabelled dashboard is one nobody dares delete and nobody quite believes.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>During your next incident, notice which dashboard you actually opened first.</p>
<p>That is your real service dashboard. Whether it is the one you designated as such, and how much of what you needed was on it, is the most honest review of your monitoring you will ever get — and it costs nothing but paying attention for five minutes while you are already there.</p>]]></content:encoded></item><item><title>The graph nobody drew</title><link>https://readme.news/the-graph-nobody-drew/</link><guid isPermaLink="true">https://readme.news/the-graph-nobody-drew/</guid><pubDate>Mon, 14 Sep 2026 09:00:00 +0000</pubDate><description>Your service dependency graph exists whether or not anyone has written it down. Writing it down takes an afternoon.</description><content:encoded><![CDATA[<p>Ask an engineering team to list what their service depends on to serve a request. You will get the database and two or three obvious APIs.</p>
<p>The real list is usually fifteen things, and the difference between the two lists is where the surprising outages come from.</p>
<h2 id="the-full-inventory">the full inventory<a class="anchor" href="#the-full-inventory" aria-label="link to this section">#</a></h2>
<p>For one service, one request path:</p>
<ul><li><strong>Datastores.</strong> Primary, replicas, cache, search index, object storage.</li><li><strong>Internal services</strong> it calls synchronously.</li><li><strong>Third-party APIs</strong> — payments, email, geocoding, <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">feature flags</a>, auth.</li><li><strong>Infrastructure control planes.</strong> Service discovery, DNS, the secret manager, the config service. These are the ones nobody lists and they are the ones that take everything down.</li><li><strong>Identity.</strong> The token issuer, the JWKS endpoint it fetches keys from.</li><li><strong>Observability.</strong> If your telemetry agent blocks when the collector is down — and some do — it is in your request path.</li><li><strong>The container registry</strong>, at start-up. Fine until you need to scale during an incident and cannot pull an image.</li><li><strong>Certificate infrastructure</strong>, including OCSP and the ACME endpoint.</li><li><strong>NTP.</strong> Rarely, spectacularly.</li></ul>
<p>The pattern: the dependencies people list are the ones they <em>call</em>. The ones that hurt are the ones they <em>need in order to start, authenticate, or scale.</em></p>
<h2 id="the-two-questions-per-dependency">the two questions per dependency<a class="anchor" href="#the-two-questions-per-dependency" aria-label="link to this section">#</a></h2>
<p>For each item, two answers, written down:</p>
<p><strong>What happens when it is slow?</strong> Slow is worse than down and much more common. Down fails fast; slow holds your connections, exhausts your pool, and turns one dependency's bad day into your outage.</p>
<p><strong>What happens when it is unavailable?</strong> Three possible answers, and only one is a decision:</p>
<ul><li><em>We fail.</em> Fine, if the dependency is genuinely essential — a database for a write path.</li><li><em>We degrade.</em> Serve stale cache, hide the feature, use a default. This is the answer for most non-essential dependencies and it requires code that usually does not exist yet.</li><li><em>We do not know.</em> This is the real answer for most dependencies, and it is the finding.</li></ul>
<h2 id="the-classification-that-matters">the classification that matters<a class="anchor" href="#the-classification-that-matters" aria-label="link to this section">#</a></h2>
<p>Sort every dependency into two buckets:</p>
<p><strong>Hard</strong> — the request genuinely cannot be served without it. Should be a very short list. Each one caps your maximum availability at its own.</p>
<p><strong>Soft</strong> — the request can be served in degraded form. Every soft dependency needs a timeout, a fallback, and a circuit breaker, or it is a hard dependency that nobody has admitted to.</p>
<p>The arithmetic is the reason this matters: five hard dependencies at 99.9% each gives you at best 99.5%, before any failure of your own. If your availability target is higher than that, some of those dependencies have to become soft, and that is an engineering project rather than a target you can declare.</p>
<h2 id="the-feature-flag-trap">the feature flag trap<a class="anchor" href="#the-feature-flag-trap" aria-label="link to this section">#</a></h2>
<p>Worth naming specifically because it catches good teams. Feature flag services are adopted as a safety mechanism — a way to turn things off when they go wrong.</p>
<p>Then the flag SDK becomes a hard dependency on the request path, and when the flag service has an outage your service does too. The tool you adopted to make incidents smaller has made one bigger.</p>
<p>The fix is standard for this whole category and rarely applied: cache flag values locally, evaluate from the cache, refresh in the background, and ship a default in the binary. Then the flag service being down means flags are stale, which is survivable.</p>
<h2 id="how-to-draw-it">how to draw it<a class="anchor" href="#how-to-draw-it" aria-label="link to this section">#</a></h2>
<p>Do not buy anything. Open a file:</p>
<div class="code"><pre><code>orders-api
  HARD  postgres-primary        write path; no fallback
  HARD  auth-jwks               cached 1h, so 1h of grace
  SOFT  redis                   cache; miss → primary, 3x latency
  SOFT  inventory-svc           timeout 300ms → show "check availability"
  SOFT  pricing-svc             timeout 200ms → list price, no promo
  SOFT  flags-svc               local cache + baked defaults
  BOOT  vault                   secrets at start; running pods unaffected
  BOOT  registry                image pull; blocks scale-up only</code></pre></div>
<p>Twenty minutes per service. The <code>BOOT</code> category — needed to start but not to serve — is the one that catches people out during incidents, because those dependencies are invisible right up until you need to replace a pod.</p>
<p>Then, for every <code>SOFT</code> line, check that the fallback described actually exists in code. That check is where the real findings are, and it is usually not the timeout that is missing but the fallback behind it.</p>]]></content:encoded></item><item><title>Timeouts: every one of them</title><link>https://readme.news/timeouts-every-one-of-them/</link><guid isPermaLink="true">https://readme.news/timeouts-every-one-of-them/</guid><pubDate>Thu, 10 Sep 2026 09:00:00 +0000</pubDate><description>A request crosses a dozen components with a timeout each, and almost nobody has ever added them up.</description><content:encoded><![CDATA[<p>Every layer of your stack has a timeout. Most of them are defaults. Almost nobody has written them down in one place and checked that they make sense together.</p>
<p>The result is systems where the client gives up at 10 seconds, the server keeps working for 60, and the database holds a lock for 300.</p>
<h2 id="the-inventory">the inventory<a class="anchor" href="#the-inventory" aria-label="link to this section">#</a></h2>
<p>For a single HTTP request, in rough order:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">layer</th><th style="text-align:left">typical default</th></tr></thead><tbody><tr><td style="text-align:left">browser / client library</td><td style="text-align:left">30s or none</td></tr><tr><td style="text-align:left">DNS resolution</td><td style="text-align:left">5s per attempt</td></tr><tr><td style="text-align:left">TCP connect</td><td style="text-align:left">20–75s (OS)</td></tr><tr><td style="text-align:left">TLS handshake</td><td style="text-align:left">inherits connect</td></tr><tr><td style="text-align:left">load balancer idle</td><td style="text-align:left">60s</td></tr><tr><td style="text-align:left">reverse proxy read</td><td style="text-align:left">60s</td></tr><tr><td style="text-align:left">application server request</td><td style="text-align:left">often none</td></tr><tr><td style="text-align:left">HTTP client to downstream</td><td style="text-align:left">often none</td></tr><tr><td style="text-align:left">connection pool acquire</td><td style="text-align:left">30s</td></tr><tr><td style="text-align:left">database statement</td><td style="text-align:left">often none</td></tr><tr><td style="text-align:left">database lock wait</td><td style="text-align:left">often none</td></tr></tbody></table></div>
<p>Two things stand out. <strong>"Often none" appears five times</strong> — most application-level timeouts are unset by default. And <strong>the values are not coordinated</strong> with each other in any way.</p>
<h2 id="the-rule-that-fixes-most-of-it">the rule that fixes most of it<a class="anchor" href="#the-rule-that-fixes-most-of-it" aria-label="link to this section">#</a></h2>
<p><strong>Timeouts must decrease as you go deeper.</strong></p>
<div class="code"><pre><code>client 10s
  └─ load balancer 9s
       └─ application 8s
            └─ downstream call 3s (with 1 retry → 6s worst case)
                 └─ connection acquire 1s
                      └─ database statement 2s</code></pre></div>
<p>Each layer must be shorter than its caller, with room for <a class="xref" href="/retries-a-complete-guide-to-not-making-it-worse/" title="Retries: a complete guide to not making it worse">retries</a>. If an inner layer can outlast its caller, the caller gives up while the work continues — which means you are burning capacity on results nobody will receive. Under load, that is most of your capacity.</p>
<h2 id="deadline-propagation">deadline propagation<a class="anchor" href="#deadline-propagation" aria-label="link to this section">#</a></h2>
<p>The better version of the rule: do not configure each timeout independently. Pass the deadline down.</p>
<div class="code"><span class="code-lang">go</span><pre><code class="lang-go">// caller has 10s; every downstream inherits what remains
ctx, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()

// 3s into the request, this call gets at most 7s, automatically
resp, err := downstream.Fetch(ctx, id)</code></pre></div>
<p>Go's <code>context</code>, gRPC's deadlines and equivalents elsewhere exist for this. Once you propagate deadlines, a slow first step automatically shortens the budget for later ones, and work never continues past the point where anyone is waiting.</p>
<p>This is the single highest-value change available in most service codebases, and it is usually a few days of threading a parameter through.</p>
<h2 id="the-ones-people-forget">the ones people forget<a class="anchor" href="#the-ones-people-forget" aria-label="link to this section">#</a></h2>
<p><strong>Lock wait timeouts.</strong> A transaction waiting on a row lock with no timeout waits forever. Set <code>lock_timeout</code> in Postgres, <code>innodb_lock_wait_timeout</code> in MySQL.</p>
<p><strong>DDL statements.</strong> A migration that cannot get its lock will queue behind a long transaction — and everything else <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> behind it. <code>SET lock_timeout = '3s'</code> before DDL is the line that prevents a large share of migration incidents.</p>
<p><strong>Idle-in-transaction.</strong> A connection that opened a transaction and went away holds locks and blocks vacuum indefinitely. <code>idle_in_transaction_session_timeout</code> is the safety net.</p>
<p><strong>Client-side connect vs. read.</strong> These are different, and most libraries let you set them separately. A short connect timeout with a long read timeout is usually what you want.</p>
<p><strong>DNS.</strong> Rarely configured, occasionally the whole problem.</p>
<h2 id="write-them-down">write them down<a class="anchor" href="#write-them-down" aria-label="link to this section">#</a></h2>
<p>One table, in the repository, listing every timeout in the request path with its current value and where it is configured.</p>
<p>The exercise takes an afternoon and reliably finds at least one inversion — an inner layer waiting longer than the outer one — and at least two places where the value is a framework default nobody chose.</p>
<p>That document then becomes something you can review when latency changes, rather than a set of numbers scattered across six config files and three languages.</p>
<h2 id="the-failure-mode-this-prevents">the failure mode this prevents<a class="anchor" href="#the-failure-mode-this-prevents" aria-label="link to this section">#</a></h2>
<p>Without coordinated timeouts, a slow dependency does not degrade your service — it exhausts it. Requests pile up waiting on something their callers abandoned long ago, connections stay held, the pool empties, and healthy requests start failing for want of a connection.</p>
<p>That is a full outage caused by one slow downstream, and the difference between it and a brief latency blip is entirely whether the numbers were coordinated.</p>]]></content:encoded></item><item><title>The queues you did not know you had</title><link>https://readme.news/the-queues-you-did-not-know-you-had/</link><guid isPermaLink="true">https://readme.news/the-queues-you-did-not-know-you-had/</guid><pubDate>Tue, 01 Sep 2026 09:00:00 +0000</pubDate><description>Every fixed-size resource is a queue. Most of them are unmonitored, and that is where latency hides.</description><content:encoded><![CDATA[<p>You know about the message queue, because you chose it and it has a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>. The queues that hurt are the ones nobody named.</p>
<p>Anywhere a fixed-size resource is shared by more requests than it has capacity, there is a queue. It has a depth, a wait time, and a failure mode, and almost none of them are instrumented.</p>
<h2 id="the-inventory">the inventory<a class="anchor" href="#the-inventory" aria-label="link to this section">#</a></h2>
<p><strong>The connection pool.</strong> Twenty connections, forty concurrent requests: twenty requests are waiting. Pool wait time is the single most under-measured latency component in web applications, and it is invisible in a database query timer because the clock starts after the connection is acquired.</p>
<p><strong>The thread pool or worker pool.</strong> Same shape. Requests queue for a worker, and the time spent waiting is not attributed to any handler.</p>
<p><strong>The TCP accept backlog.</strong> The kernel holds connections your process has not accepted yet. Overflow silently drops them, and the client sees a timeout with no server-side trace at all.</p>
<p><strong>The HTTP client's per-host connection limit.</strong> Most clients cap concurrent connections per destination. Exceed it and your requests queue in the client, before any network activity, invisible to server-side metrics on both ends.</p>
<p><strong>The DNS resolver.</strong> A limited number of in-flight lookups with a cache that can stampede on expiry.</p>
<p><strong>The disk queue.</strong> Storage devices have a queue depth. Exceed it and I/O waits.</p>
<p><strong>The garbage collector.</strong> Not a queue exactly, but the same behaviour: work that accumulates and is paid in a burst.</p>
<p><strong>Rate limiters.</strong> A limiter that delays rather than rejecting is a queue with an enforced service rate.</p>
<h2 id="why-this-matters-more-than-it-sounds">why this matters more than it sounds<a class="anchor" href="#why-this-matters-more-than-it-sounds" aria-label="link to this section">#</a></h2>
<p>Little's Law: <code>L = λW</code>. Items in the system equals arrival rate times time in system. Rearranged, <strong>wait time is queue length over service rate</strong>.</p>
<p>That means your queue depth choice <em>is</em> a latency choice, whether or not you made it deliberately. A connection pool with a 30-second acquisition timeout is a promise that some requests will wait 30 seconds.</p>
<p>And these queues compose. A request waits for a worker, then waits for a connection, then waits for the disk. Each is modest; the sum is the p99 nobody can explain, because each layer's own metrics look fine.</p>
<h2 id="how-to-find-them">how to find them<a class="anchor" href="#how-to-find-them" aria-label="link to this section">#</a></h2>
<p><strong>Measure acquisition, not just use.</strong> Time the wait for the resource separately from the work done with it:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">t0 = time.perf_counter()
with pool.acquire() as conn:
    t1 = time.perf_counter()
    result = conn.execute(query)
span.set_attribute("db.pool_wait_ms", (t1 - t0) * 1000)
span.set_attribute("db.query_ms", (time.perf_counter() - t1) * 1000)</code></pre></div>
<p>Two numbers instead of one. The first one is the one you did not have, and it is frequently the larger.</p>
<p><strong>Add up your span times.</strong> If the parent span is 400 ms and the children total 180 ms, the missing 220 ms is queueing somewhere. That gap is the most useful signal in a trace and almost nobody looks for it.</p>
<p><strong>Check <code>netstat -s</code> for accept queue overflows.</strong> A non-zero and growing "listen queue overflowed" counter means you are dropping connections before your application sees them.</p>
<p><strong>Load test past capacity deliberately.</strong> Push to 3× and watch which metric degrades first. That is your binding queue, and you will not find it at normal load.</p>
<h2 id="the-rule">the rule<a class="anchor" href="#the-rule" aria-label="link to this section">#</a></h2>
<p>Every queue needs three things, and most have none:</p>
<ol><li><strong>A bounded size.</strong> Unbounded is a memory leak with a friendly name.</li><li><strong>A wait metric</strong> — specifically the <em>age of the oldest waiter</em>, which tells you how far behind you are in time rather than in count.</li><li><strong>A decision about overflow.</strong> Reject, drop, or block. Pick one deliberately, because the default is usually "queue forever," and queueing forever means serving requests whose callers gave up ten seconds ago.</li></ol>
<p>Do that for the queues you chose. Then go find the six you did not.</p>]]></content:encoded></item><item><title>What a good incident channel looks like</title><link>https://readme.news/what-a-good-incident-channel-looks-like/</link><guid isPermaLink="true">https://readme.news/what-a-good-incident-channel-looks-like/</guid><pubDate>Sat, 29 Aug 2026 09:00:00 +0000</pubDate><description>The chat log is where the incident is actually run. Four conventions make it useful rather than a wall of noise.</description><content:encoded><![CDATA[<p>Every incident runs in a chat channel. Almost nobody has thought about how that channel should work, so the default is thirty people speculating in parallel while two people try to fix something.</p>
<p>Four conventions fix most of it, and none of them need a tool.</p>
<h2 id="1-one-channel-created-at-declaration">1. one channel, created at declaration<a class="anchor" href="#1-one-channel-created-at-declaration" aria-label="link to this section">#</a></h2>
<p>Not the team channel. Not a thread in the team channel. A dedicated channel per incident, named predictably: <code>inc-2026-08-29-checkout-errors</code>.</p>
<p>Why it matters: the incident becomes searchable as a unit, the timeline is the channel, and people who join late can read from the top instead of asking what happened. It also means the incident does not drown the normal channel, and the normal channel does not drown the incident.</p>
<p>Archive it afterwards rather than deleting it. It is the primary source for the postmortem.</p>
<h2 id="2-named-roles-stated-out-loud">2. named roles, stated out loud<a class="anchor" href="#2-named-roles-stated-out-loud" aria-label="link to this section">#</a></h2>
<p>Three, at minimum, posted in the channel as the first message:</p>
<div class="code"><pre><code>IC: @priya       — decides, delegates, does not debug
Comms: @marcus   — status page, stakeholders, customer support
Ops: @sam        — hands on keyboard</code></pre></div>
<p>The incident commander not debugging is the rule people resist and the one that matters most. The moment the IC opens a terminal, nobody is tracking the whole picture, and the incident gets longer.</p>
<p>For a small incident one person can hold two roles. They should still say which ones, because the alternative is everyone assuming someone else is doing comms.</p>
<h2 id="3-a-pinned-status-message-edited-in-place">3. a pinned status message, edited in place<a class="anchor" href="#3-a-pinned-status-message-edited-in-place" aria-label="link to this section">#</a></h2>
<p>One message, pinned, rewritten as things change:</p>
<div class="code"><pre><code>STATUS 14:22 — Checkout failing for ~12% of users since 13:58.
Cause: unknown. Suspect the payments deploy at 13:55.
Now: rolling back that deploy (@sam), ETA 5 min.
Impact: card payments only, wallet payments unaffected.
Next update: 14:35</code></pre></div>
<p>Five lines: what is broken, since when, what we think, what we are doing, when the next update comes.</p>
<p>This single convention removes most of the noise, because it answers the question that generates the noise — "what's the current state?" — without anyone having to ask. Everyone joining reads the pin instead of scrolling.</p>
<p><strong>Always include the next-update time.</strong> It is what stops people asking for updates.</p>
<h2 id="4-mark-the-speculation">4. mark the speculation<a class="anchor" href="#4-mark-the-speculation" aria-label="link to this section">#</a></h2>
<p>The most common way incidents go wrong is that a guess gets repeated until it becomes the working theory, and then twenty minutes go into the wrong system.</p>
<p>A one-word convention fixes it:</p>
<div class="code"><pre><code>FACT: error rate went from 0.1% to 12% at 13:58:20
FACT: the payments deploy completed at 13:55:41
GUESS: the deploy caused it
ACTION: rolling back to confirm</code></pre></div>
<p>Facts have evidence attached. Guesses are labelled as guesses. Actions say who is doing them.</p>
<p>It looks pedantic for about four minutes and then it saves the incident, because somebody reading the channel can tell the difference between what is known and what somebody said.</p>
<h2 id="what-to-keep-out">what to keep out<a class="anchor" href="#what-to-keep-out" aria-label="link to this section">#</a></h2>
<p><strong>Speculation from people who are not investigating.</strong> Well-meant and it fills the channel that responders are trying to read. If you are not on a role, watch.</p>
<p><strong>"Is it fixed yet?"</strong> The pinned status has the next update time.</p>
<p><strong>Root-cause analysis during the incident.</strong> Restore first. The good question at 14:22 is "what makes this stop", not "why did this happen". Why is a question for Thursday.</p>
<p><strong>Blame, in any form, including jokes.</strong> It is in a permanent record that a person will read afterwards.</p>
<h2 id="the-handoff">the handoff<a class="anchor" href="#the-handoff" aria-label="link to this section">#</a></h2>
<p>Long incidents cross shift boundaries and the handoff is where context dies. It needs to be explicit and in the channel:</p>
<div class="code"><pre><code>HANDOFF 22:00 — @priya → @dan (IC)
Ruled out: deploy (rolled back, no change), CDN (unaffected regions also failing)
Current theory: connection pool exhaustion on the orders DB
In flight: @sam is capturing pg_stat_activity every 30s → thread above
Not yet tried: failover to replica
Customers: status page updated 21:40, support has the template</code></pre></div>
<p>Ruled out, current theory, in flight, not tried, customer state. Five lines, and the incoming IC starts with the accumulated knowledge instead of rediscovering it.</p>
<h2 id="why-bother">why bother<a class="anchor" href="#why-bother" aria-label="link to this section">#</a></h2>
<p>The channel is not a side effect of the incident. During the incident it is the coordination mechanism, and afterwards it is the only complete record of what was known and when.</p>
<p>A channel that follows these four conventions produces a postmortem that half writes itself, and — more importantly — an incident that ends sooner, because the people fixing it spent their attention on the system rather than on the chat.</p>]]></content:encoded></item><item><title>The cost of a 500</title><link>https://readme.news/the-cost-of-a-500/</link><guid isPermaLink="true">https://readme.news/the-cost-of-a-500/</guid><pubDate>Thu, 20 Aug 2026 09:00:00 +0000</pubDate><description>Error rates get tracked as a percentage. Users experience them as a specific person having a specific bad day.</description><content:encoded><![CDATA[<p>A 0.1% error rate sounds like nothing. It is a rounding error on a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>, it is well inside almost any SLO, and it is the kind of number that gets a green tick in a weekly review.</p>
<p>Run the arithmetic on it and the picture is different.</p>
<h2 id="the-arithmetic">the arithmetic<a class="anchor" href="#the-arithmetic" aria-label="link to this section">#</a></h2>
<p>Ten million requests a day at 0.1% is <strong>ten thousand failed requests</strong>. If a user session averages twenty requests, that is roughly five hundred sessions with at least one failure — every day.</p>
<p>But the failures are not evenly spread, and that is the part that matters. Errors cluster: by endpoint, by tenant, by client version, by region, by input shape. A 0.1% global rate is frequently one customer at 40% and everyone else at 0.001%.</p>
<p>For that customer, your service is broken. Your dashboard says 99.9%.</p>
<h2 id="the-slices-that-reveal-it">the slices that reveal it<a class="anchor" href="#the-slices-that-reveal-it" aria-label="link to this section">#</a></h2>
<p>Aggregate error rate is nearly useless on its own. The same number means completely different things depending on the distribution beneath it, and you cannot tell which without slicing:</p>
<ul><li><strong>By customer or tenant.</strong> The single highest-value slice in any multi-tenant system. One broken integration hides perfectly in a global average.</li><li><strong>By endpoint.</strong> A 0.1% average across forty endpoints can be 4% on the one that matters.</li><li><strong>By client version.</strong> Errors concentrated in one mobile build means a bad release, not a server problem.</li><li><strong>By region.</strong> Often a networking or dependency issue, not application logic.</li><li><strong>By authenticated vs anonymous.</strong> Frequently a different code path entirely.</li></ul>
<p>If your telemetry cannot answer "which customers saw errors in the last hour," that is the gap to close before adding any more dashboards.</p>
<h2 id="the-errors-that-are-not-500s">the errors that are not 500s<a class="anchor" href="#the-errors-that-are-not-500s" aria-label="link to this section">#</a></h2>
<p>Counting HTTP status codes undercounts real failure, sometimes badly:</p>
<ul><li><strong>200 with an empty result</strong> because a downstream timed out and the code swallowed it. The most dangerous category — invisible to every status-based metric.</li><li><strong>200 with a partial result.</strong> The page rendered, three of the eight widgets are missing.</li><li><strong>Client-side failures</strong> that never reach your server at all: a bad bundle, a CSP violation, a network drop mid-request.</li><li><strong>Slow enough to be a failure.</strong> A request that succeeds after 45 seconds has failed as far as the user is concerned, and it is counted as a success.</li><li><strong>The retry that worked.</strong> The user saw a spinner for eight seconds. Your success rate is 100%.</li></ul>
<p><strong>The fix is measuring outcomes rather than responses.</strong> Did the order get placed? Did the message send? Those are answerable from your own data and they count the failures that status codes miss.</p>
<h2 id="what-a-single-failure-actually-costs">what a single failure actually costs<a class="anchor" href="#what-a-single-failure-actually-costs" aria-label="link to this section">#</a></h2>
<p>Worth being concrete, because "0.1%" abstracts it away:</p>
<ul><li>The user <a class="xref" href="/retries-a-complete-guide-to-not-making-it-worse/" title="Retries: a complete guide to not making it worse">retries</a>. If it fails again, some fraction leave.</li><li>Some fraction contact support. Each of those is real money and an engineer's attention.</li><li>Some fraction never come back, and you will not attribute that to today.</li><li>If it is a payment or a submission, someone may not know whether it went through — which generates duplicates, which generates a different class of problem.</li></ul>
<p>None of that appears on the error-rate graph, which is why the graph being green is not the same as things being fine.</p>
<h2 id="the-practical-program">the practical program<a class="anchor" href="#the-practical-program" aria-label="link to this section">#</a></h2>
<p><strong>Alert on the slice, not the aggregate.</strong> "Any single tenant above 5% for ten minutes" catches things the global rate never will.</p>
<p><strong>Sample real failures and read them.</strong> Not the counts — the actual requests, with their inputs and their traces. An hour a week reading real failed requests teaches you more about your system than any dashboard.</p>
<p><strong>Put a request ID on the error page.</strong> So a user's complaint becomes a trace lookup rather than an investigation.</p>
<p><strong>Track outcomes for your critical journeys.</strong> Checkout completed, message delivered, deploy finished. That number is the one to put in the weekly review, not the status-code ratio.</p>
<h2 id="the-framing-that-changes-behaviour">the framing that changes behaviour<a class="anchor" href="#the-framing-that-changes-behaviour" aria-label="link to this section">#</a></h2>
<p>Stop saying "we're at 99.9%."</p>
<p>Start saying "<strong>about five hundred people had a broken session yesterday.</strong>"</p>
<p>Same data. Completely different meeting.</p>]]></content:encoded></item><item><title>Backpressure is the concept your system is missing</title><link>https://readme.news/backpressure-is-the-concept-your-system-is-missing/</link><guid isPermaLink="true">https://readme.news/backpressure-is-the-concept-your-system-is-missing/</guid><pubDate>Fri, 03 Jul 2026 09:00:00 +0000</pubDate><description>When a fast producer meets a slow consumer, something has to give. Deciding what, in advance, is the whole discipline.</description><content:encoded><![CDATA[<p>Every system with a producer and a consumer eventually has a moment where the producer is faster. What happens next is either a design decision you made or an emergent behavior you discover during an incident.</p>
<h2 id="the-four-options">the four options<a class="anchor" href="#the-four-options" aria-label="link to this section">#</a></h2>
<p>There are only four. Every system picks one, explicitly or by accident.</p>
<p><strong>1. Buffer.</strong> Queue the excess.</p>
<p>Works for bursts. Fails for sustained overload, because a buffer is a delay, and an unbounded buffer is a memory leak with a friendly name. The failure mode is that memory grows until the process dies, taking the buffer with it — so you lose everything, at the worst possible moment.</p>
<p><strong>2. Drop.</strong> Discard the excess.</p>
<p>Correct more often than people are comfortable with. Metrics, logs, telemetry, non-critical events — dropping 5% of samples under load is fine and dying is not.</p>
<p>The requirement: <strong>know that you dropped, and how much.</strong> Silent drops are how you get a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> that looks healthy while data is missing.</p>
<p><strong>3. Block.</strong> Make the producer wait.</p>
<p>This is real backpressure. The consumer's slowness propagates upstream, the producer slows down, and the system reaches equilibrium at the consumer's rate.</p>
<p>Correct for internal pipelines where the producer can wait. Dangerous when the producer is a user-facing request handler, because now user requests are blocked on a background process.</p>
<p><strong>4. Reject.</strong> Tell the producer no.</p>
<p>The right answer at a service boundary. A 429 or 503 returned in one millisecond is much better than a request that waits thirty seconds and then times out, because the caller can make a decision — retry later, degrade, or tell the user.</p>
<h2 id="the-wrong-default">the wrong default<a class="anchor" href="#the-wrong-default" aria-label="link to this section">#</a></h2>
<p>Most systems buffer by default, unboundedly, without anyone deciding.</p>
<ul><li>An in-memory list that grows.</li><li>A queue with no maximum length.</li><li>A connection pool that <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> waiters forever.</li><li>A channel with a very large capacity, which is unbounded in practice.</li></ul>
<p>The failure mode is always the same: latency grows, memory grows, and then the process dies. And the requests you were holding were abandoned by their callers ten seconds earlier, so all that work was for nothing.</p>
<p><strong>A queue that is always full is not a buffer. It is a delay you cannot see.</strong></p>
<h2 id="the-practical-rules">the practical rules<a class="anchor" href="#the-practical-rules" aria-label="link to this section">#</a></h2>
<p><strong>Every queue has a maximum size.</strong> Every one. Choose the number by asking: how long should a request be willing to wait? Multiply by the consumer's rate. That is your queue depth.</p>
<p>If you cannot answer that question, the queue is not designed.</p>
<p><strong>Prefer rejecting to queueing at the edge.</strong> When a request arrives and the system is saturated, reject fast. The caller has a timeout; use it as your budget.</p>
<p><strong>Propagate deadlines.</strong> If the caller has 200 ms left, every downstream operation should know that. Work performed after the caller has given up is pure waste, and under overload it is the majority of the work being done.</p>
<div class="code"><span class="code-lang">go</span><pre><code class="lang-go">ctx, cancel := context.WithTimeout(ctx, remaining)
defer cancel()</code></pre></div>
<p><strong>Shed by priority, not randomly.</strong> Under load, serve health checks, serve authenticated users, serve the critical path. Shed background work, analytics, and prefetches. Random shedding means your health checks fail and your orchestrator kills healthy instances, which is a self-inflicted outage.</p>
<p>**Measure queue <em>age</em>, not depth.** Depth tells you how many. Age tells you how far behind. "The oldest item has been waiting four minutes" is actionable in a way that "there are 30,000 items" is not.</p>
<h2 id="the-pattern-that-ties-it-together">the pattern that ties it together<a class="anchor" href="#the-pattern-that-ties-it-together" aria-label="link to this section">#</a></h2>
<p>Little's Law: <code>L = λW</code>. Items in the system equals arrival rate times time in system.</p>
<p>Rearranged: <strong>wait time equals queue length divided by service rate.</strong></p>
<p>If your queue holds 10,000 items and you process 100 per second, the newest item waits 100 seconds. That is not a hypothetical — it is arithmetic, and it means your queue length choice <em>is</em> your latency choice, whether or not you framed it that way.</p>
<p>Pick the latency you can accept, multiply by the service rate, and that is your maximum queue length. Everything past it gets rejected.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>Point a load generator at your service at three times its capacity. Watch:</p>
<ul><li>Does memory grow without bound?</li><li>Does latency grow without bound?</li><li>Does anything get rejected, or does everything just get slower?</li><li>After you stop the load, how long until it recovers?</li></ul>
<p>That last one is the important one. A system that takes twenty minutes to recover from a two-minute overload has a backpressure problem, and it will turn a small incident into a large one on a day you did not choose.</p>]]></content:encoded></item><item><title>Retries: a complete guide to not making it worse</title><link>https://readme.news/retries-a-complete-guide-to-not-making-it-worse/</link><guid isPermaLink="true">https://readme.news/retries-a-complete-guide-to-not-making-it-worse/</guid><pubDate>Mon, 29 Jun 2026 09:00:00 +0000</pubDate><description>The most common way a small incident becomes a large one is a retry policy written without thinking about aggregate behavior.</description><content:encoded><![CDATA[<p>Retries are the most commonly implemented and most commonly wrong piece of resilience engineering. A retry policy that seems obviously correct in isolation is frequently the mechanism that turns a brief degradation into a full outage.</p>
<h2 id="the-failure-mode">the failure mode<a class="anchor" href="#the-failure-mode" aria-label="link to this section">#</a></h2>
<p>A downstream service slows down. Every caller times out. Every caller retries.</p>
<p>The downstream now receives double its normal traffic while already struggling. More requests time out. More retries. The load multiplies with every round.</p>
<p>The original problem might have been a five-second blip. The retry storm keeps the service down for twenty minutes, and it stays down after the original cause is resolved because the queued retries are still arriving.</p>
<p>This is metastable failure and retries are its most common cause.</p>
<h2 id="the-rules">the rules<a class="anchor" href="#the-rules" aria-label="link to this section">#</a></h2>
<p><strong>1. Only retry idempotent operations.</strong></p>
<p>A retried non-idempotent operation charges the card twice. If you need to retry a write, make it idempotent first with an <a class="xref" href="/idempotency-is-the-only-distributed-systems-concept-you-need/" title="Idempotency is the only distributed systems concept you need">idempotency</a> key.</p>
<p><strong>2. Only retry retryable errors.</strong></p>
<p>A 400 will be a 400 next time. Retrying it wastes a request and delays the error the caller needs to see.</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">retry</th><th style="text-align:left">do not retry</th></tr></thead><tbody><tr><td style="text-align:left">connection refused / reset</td><td style="text-align:left">400, 401, 403, 404</td></tr><tr><td style="text-align:left">timeout</td><td style="text-align:left">422</td></tr><tr><td style="text-align:left">502, 503, 504</td><td style="text-align:left">any deterministic validation failure</td></tr><tr><td style="text-align:left">429 (respect <code>Retry-After</code>)</td><td style="text-align:left">501</td></tr></tbody></table></div>
<p>The one people get wrong: <strong>500 is ambiguous.</strong> It might be transient, it might be a deterministic bug. Retrying it once is usually reasonable; retrying it five times is usually pointless.</p>
<p><strong>3. Exponential backoff with jitter. Always jitter.</strong></p>
<p>Without jitter, all your clients retry at the same moments, and you have built a synchronized load generator.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def delay(attempt, base=0.1, cap=30):
    return random.uniform(0, min(cap, base * (2 ** attempt)))</code></pre></div>
<p>That is full jitter, and it is the recommended default. It spreads retries across the whole interval, which is what you want. Half jitter — <code>d/2 + random(0, d/2)</code> — is a reasonable alternative when you want a guaranteed minimum delay.</p>
<p>The version without jitter is the one everyone writes first and it is the one that causes the storm.</p>
<p><strong>4. Cap the attempts and cap the total time.</strong></p>
<p>Three attempts, usually. And a total deadline — if the caller is going to give up after two seconds, retrying at three seconds is pure waste and it is load on a struggling service.</p>
<p>Propagate the deadline. If your caller has 500 ms left, your retry budget is 500 ms, not your configured default.</p>
<p><strong>5. Do not retry at every layer.</strong></p>
<p>This is the one that produces the shocking numbers. Three attempts at the HTTP client, three in the service wrapper, three in the caller, three at the gateway: 3⁴ = 81 requests for one logical call.</p>
<p><strong>Retry at exactly one layer.</strong> Usually the outermost one that has the context to decide. Every other layer fails fast and propagates.</p>
<p>Audit this. Most systems that have grown organically retry at three or four layers and nobody knows.</p>
<p><strong>6. Use a retry budget.</strong></p>
<p>The refinement that actually prevents storms: cap retries as a <em>fraction of total traffic</em>, not per request.</p>
<div class="code"><pre><code>if retries_in_window / requests_in_window &gt; 0.1:
    do_not_retry()</code></pre></div>
<p>When things are healthy, occasional retries are well under the budget and everything works. When things are broken, the budget is exhausted immediately and retries stop entirely — exactly when they would do the most harm.</p>
<p>This single mechanism converts retries from a failure amplifier into a bounded safety net, and it is not widely implemented.</p>
<p><strong>7. Circuit break.</strong></p>
<p>When a downstream is clearly failing, stop calling it. Fail fast, return a cached or degraded response, and probe occasionally to see if it has recovered.</p>
<p>Three states: closed (normal), open (failing fast), half-open (probing). The half-open state must allow only a trickle — if you send full traffic at a recovering service you will knock it over again.</p>
<h2 id="the-server-side">the server side<a class="anchor" href="#the-server-side" aria-label="link to this section">#</a></h2>
<p>The other half, which is usually forgotten.</p>
<p><strong>Send <code>Retry-After</code> on 429 and 503.</strong> Then clients that respect it retry at a time you chose rather than a time they chose.</p>
<p><strong>Shed load rather than queueing it.</strong> A request that will time out anyway should be rejected immediately, not queued. Queueing under overload increases latency for everything without increasing throughput, and the requests you eventually serve have often already been abandoned.</p>
<p><strong>Prioritize.</strong> Under load, serve health checks and critical paths, shed the rest. An unprioritized overload sheds randomly, which means your health checks fail and your orchestrator kills healthy instances.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>Take your service. Make a downstream dependency return 503 for everything. Watch the request rate at the downstream.</p>
<p>If it goes up by more than a small factor, your retry configuration will cause an outage. You have just not had the trigger yet.</p>]]></content:encoded></item><item><title>Configuration is the most under-designed part of your system</title><link>https://readme.news/configuration-is-the-most-under-designed-part-of-your-system/</link><guid isPermaLink="true">https://readme.news/configuration-is-the-most-under-designed-part-of-your-system/</guid><pubDate>Wed, 24 Jun 2026 09:00:00 +0000</pubDate><description>It has no type system, no tests, no review, and it causes a disproportionate share of outages.</description><content:encoded><![CDATA[<p>Look at the last ten significant outages you can remember reading about. A disproportionate number were caused by configuration, not code.</p>
<p>A config file with a wrong value. A <a class="xref" href="/the-graph-nobody-drew/" title="The graph nobody drew">feature flag</a> flipped. A generated file that doubled in size. A DNS record. A permission change.</p>
<p>Configuration receives a fraction of the engineering rigor that code does, and it has comparable power to break things.</p>
<h2 id="why-it-goes-wrong">why it goes wrong<a class="anchor" href="#why-it-goes-wrong" aria-label="link to this section">#</a></h2>
<p><strong>No <a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">type system</a>.</strong> A YAML file will happily contain <code>timeout: 30s</code> where the code expects an integer, or <code>enabled: "false"</code> which is a truthy string.</p>
<p><strong>No tests.</strong> Nobody writes a test for their config.</p>
<p><strong>No review, or perfunctory review.</strong> A config change is "just a value" and gets approved in ten seconds.</p>
<p><strong>Deployed differently from code.</strong> Frequently faster, frequently without staging, frequently without a canary. The safety mechanisms built for code deployment routinely do not cover config.</p>
<p><strong>Environment drift.</strong> Staging and production differ in ways nobody has enumerated, so staging validates a configuration that is not production's.</p>
<p><strong>No rollback story.</strong> "What was this value before" is often unanswerable.</p>
<h2 id="the-fixes">the fixes<a class="anchor" href="#the-fixes" aria-label="link to this section">#</a></h2>
<p><strong>1. Parse and validate at startup, not at use.</strong></p>
<p>The worst failure mode is a config error that manifests three hours later when a rarely-used code path reads a malformed value.</p>
<p>Validate everything at boot. Fail loudly and immediately if anything is wrong.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">class Settings(BaseModel):
    database_url: PostgresDsn
    timeout_ms: int = Field(gt=0, le=60_000)
    max_connections: int = Field(ge=1, le=1000)
    feature_new_checkout: bool = False

settings = Settings(**load_config())   # raises at startup, with a clear message</code></pre></div>
<p>An application that will not start is far better than one that starts and behaves wrongly.</p>
<p><strong>2. Types, with a schema.</strong></p>
<p>Whatever your language, there is a library that turns untyped config into a validated typed object. Use it. This eliminates the entire category of string-that-should-be-a-number bugs.</p>
<p><strong>3. Config changes go through the same pipeline as code.</strong></p>
<p>Version control. Review. Staging. Canary. Rollback.</p>
<p>The argument against is that config changes need to be fast, especially for incident response. That is a real need and the answer is a small explicitly-defined set of emergency levers — kill switches, <a class="xref" href="/rate-limits-are-a-product-decision-not-an-infrastructure-one/" title="Rate limits are a product decision, not an infrastructure one">rate limits</a> — with a fast path, and everything else on the normal pipeline.</p>
<p>Not "all config is fast" because that is how config takes down your service.</p>
<p><strong>4. Validate generated configuration.</strong></p>
<p>If a config file is produced by a program, that program can be wrong. A size check against the previous version, a schema check, a sanity check on record count.</p>
<p>This is the specific failure that has caused several high-profile outages: a generated file changed unexpectedly and propagated globally before anyone looked at it.</p>
<p><strong>5. Make the environment explicit and diff it.</strong></p>
<p>You should be able to answer "how does staging differ from production" with a command. If you cannot, staging is not validating production.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">config-diff staging production</code></pre></div>
<p><strong>6. Log the effective configuration at startup.</strong></p>
<p>Not the file — the resolved values after defaults, overrides, and environment variables are applied. Redact secrets. This is the single most useful thing for debugging "it works on my machine," because the effective config is frequently not what anyone thinks it is.</p>
<p><strong>7. <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">Feature flags</a> need the same discipline as code.</strong></p>
<p>Owner. Expiry. Test both branches. Log the flag state on every event. A flag flip is a production change and should be treated as one.</p>
<h2 id="the-hierarchy-that-works">the hierarchy that works<a class="anchor" href="#the-hierarchy-that-works" aria-label="link to this section">#</a></h2>
<p>Most systems end up with layered configuration and the layering should be explicit and simple:</p>
<div class="code"><pre><code>defaults in code
  ← config file
    ← environment variables
      ← command line flags</code></pre></div>
<p>Later overrides earlier. Log which layer each effective value came from when debugging.</p>
<p><strong>Keep the layers few.</strong> Systems with six overlapping sources of configuration — defaults, file, environment, a service, a database table, a flag system — produce values nobody can trace. Each layer you add makes "why is this value what it is" harder to answer.</p>
<h2 id="secrets">secrets<a class="anchor" href="#secrets" aria-label="link to this section">#</a></h2>
<p>Not the same thing as configuration and should not be in the same place.</p>
<ul><li>Never in the config file. Never in version control. Never in the image.</li><li>A secret manager, injected at runtime.</li><li>Rotated on a schedule, and the rotation must be tested.</li><li>Never logged. Redact by default at the logger, not at each call site — because somebody will forget.</li></ul>
<h2 id="the-general-point">the general point<a class="anchor" href="#the-general-point" aria-label="link to this section">#</a></h2>
<p>Configuration is the input to your program that is most likely to be wrong and least likely to be checked.</p>
<p>Every rigor you apply to code — types, validation, tests, review, staged rollout, rollback — applies to configuration, and applying it costs a day of setup.</p>
<p>The reason it does not happen is that config does not feel like code. It is data, and data feels safe.</p>
<p>It is not data. It is the arguments to your program, and passing wrong arguments to a program is how programs go wrong.</p>]]></content:encoded></item><item><title>Randomness, and the bugs you cannot reproduce</title><link>https://readme.news/randomness-and-the-bugs-you-cannot-reproduce/</link><guid isPermaLink="true">https://readme.news/randomness-and-the-bugs-you-cannot-reproduce/</guid><pubDate>Mon, 22 Jun 2026 09:00:00 +0000</pubDate><description>The bug that happens once a week and never in staging. A systematic approach to the class of problem everyone handles badly.</description><content:encoded><![CDATA[<p>The worst bugs are the ones that happen sometimes. They resist the standard debugging loop entirely, because the loop requires reproduction and reproduction is exactly what you do not have.</p>
<p>Most engineers approach these by staring at code and hoping. There is a better method.</p>
<h2 id="the-sources-of-nondeterminism">the sources of nondeterminism<a class="anchor" href="#the-sources-of-nondeterminism" aria-label="link to this section">#</a></h2>
<p>There is a finite list. Work through it.</p>
<p><strong>Concurrency.</strong> Two things running at once with insufficient ordering. The largest category by far. Includes: unsynchronized shared state, check-then-act races, lost updates, and the classic where two requests both check "does this exist" and both create it.</p>
<p><strong>Time.</strong> Anything that depends on wall clock: <a class="xref" href="/timeouts-every-one-of-them/" title="Timeouts: every one of them">timeouts</a>, expiry, scheduled work, date boundaries. These fail at midnight, at month end, on the DST transition, on leap day, and when NTP adjusts the clock backward.</p>
<p><strong>Ordering.</strong> Hash map iteration order, filesystem directory order, unordered message delivery, parallel test execution. Code that accidentally depends on an order that is not guaranteed works until it does not.</p>
<p><strong>External state.</strong> A cache that is sometimes warm. A connection that is sometimes pooled. A DNS response that sometimes returns a different IP. A downstream that is sometimes slow.</p>
<p><strong>Resource exhaustion.</strong> Memory pressure changes GC timing, which changes interleaving. Connection pool exhaustion changes behavior under load. File descriptor limits. These make other latent bugs appear.</p>
<p><strong>Actual randomness.</strong> UUIDs, load balancing, sampling, <a class="xref" href="/retries-a-complete-guide-to-not-making-it-worse/" title="Retries: a complete guide to not making it worse">retries</a> with jitter, partitioning by hash.</p>
<p><strong>Uninitialized memory or undefined behavior</strong>, in languages that permit it.</p>
<h2 id="the-method">the method<a class="anchor" href="#the-method" aria-label="link to this section">#</a></h2>
<p><strong>1. Instrument before you theorize.</strong></p>
<p>You cannot reproduce it, so you must capture it. Add logging around the suspicious area — not "entering function," but the actual values, the timing, the thread or task identity, the state.</p>
<p>The instinct is to avoid adding logging to production. Add it. A bug you cannot reproduce is a bug you must observe in the wild, and that requires observation.</p>
<p><strong>2. Find the correlation.</strong></p>
<p>You have occurrences. What do they have in common?</p>
<ul><li>Time of day. (Points at scheduled work, or peak load.)</li><li>Specific tenant or user. (Points at data-dependent behavior.)</li><li>Specific host or region.</li><li>Specific client version.</li><li>Load level.</li><li>Whether some other event happened first.</li></ul>
<p>This is why wide structured events matter. If every request logs its context, the correlation is a query. Without it, you are guessing.</p>
<p><strong>3. Make it more likely.</strong></p>
<p>Once you have a hypothesis, try to increase the failure rate:</p>
<ul><li><strong>Suspect a race?</strong> Add a sleep in the window you think is unprotected. If the failure rate goes from 0.1% to 90%, you found it.</li><li><strong>Suspect ordering?</strong> Randomize the order deliberately. Run tests with randomized seeds and shuffled execution.</li><li><strong>Suspect load-related?</strong> Load test with the specific pattern.</li><li><strong>Suspect time?</strong> Set the clock. Run at 23:59:59. Run on 29 February.</li></ul>
<p>Making a rare bug common is the single most effective technique in this whole category, and it is underused because it feels like making things worse.</p>
<p><strong>4. Add an assertion.</strong></p>
<p>If you believe an invariant holds, assert it. In production, with an alert.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">assert order.total == sum(i.price * i.qty for i in order.items), \
    f"order {order.id} total mismatch: {order.total}"</code></pre></div>
<p>You will find out that it does not hold, and you will find out with the context attached rather than downstream where the symptom appears.</p>
<p>The window between "the invariant broke" and "someone noticed the symptom" is where the information lives, and assertions collapse it to zero.</p>
<p><strong>5. Bisect it.</strong></p>
<p>If it started recently, <code>git bisect</code> still works on statistical failures — you just need a test script that runs the operation many times and reports failure if the rate exceeds a threshold. Slower than a deterministic bisect and far better than reading diffs.</p>
<h2 id="the-prevention">the prevention<a class="anchor" href="#the-prevention" aria-label="link to this section">#</a></h2>
<p><strong>Make things deterministic where you can.</strong> Inject the clock rather than calling it. Inject the random source with a seed. Sort collections before iterating where order matters. Deterministic systems have bugs you can reproduce, which means bugs you can fix.</p>
<p><strong>Run tests in randomized order, with a printed seed.</strong> If a test only passes in a specific order, it has a hidden dependency, and that dependency is a bug in your production code more often than people assume.</p>
<p><strong>Use your language's race detector.</strong> Go's <code>-race</code>, thread sanitizer for C and C++, Java's concurrency tooling. These find real races that have never yet manifested. Run them in CI.</p>
<p><strong>Prefer immutability.</strong> A value that cannot change cannot be changed concurrently. This eliminates the largest category of nondeterminism structurally.</p>
<p><strong>Log the identifiers.</strong> Trace ID, request ID, tenant. Nine tenths of the difficulty in these investigations is being unable to correlate events you already recorded.</p>
<h2 id="the-hard-truth">the hard truth<a class="anchor" href="#the-hard-truth" aria-label="link to this section">#</a></h2>
<p>Some of these take weeks. A concurrency bug that occurs once per million requests in a system with a subtle ordering dependency is genuinely difficult, and no method makes it easy.</p>
<p>The method makes it <em>tractable</em>: instead of staring and hoping, you have a list of candidate sources, a way to gather evidence, and a way to raise the failure rate until it is reproducible.</p>
<p>That is the difference between a bug you eventually fix and one that sits in the backlog for two years labeled "cannot reproduce."</p>]]></content:encoded></item><item><title>Distributed tracing that people actually use</title><link>https://readme.news/distributed-tracing-that-people-actually-use/</link><guid isPermaLink="true">https://readme.news/distributed-tracing-that-people-actually-use/</guid><pubDate>Mon, 15 Jun 2026 09:00:00 +0000</pubDate><description>Most tracing deployments produce beautiful waterfalls nobody opens. The difference is three implementation details.</description><content:encoded><![CDATA[<p>Distributed tracing is the right answer to "what happened to this request across nine services." Most implementations produce a system that is technically correct and that nobody opens during an incident.</p>
<p>Three details separate the two outcomes.</p>
<h2 id="detail-one-the-trace-id-must-be-everywhere">detail one: the trace ID must be everywhere<a class="anchor" href="#detail-one-the-trace-id-must-be-everywhere" aria-label="link to this section">#</a></h2>
<p>A trace is only useful if you can find it. Which means the trace ID must appear:</p>
<ul><li><strong>In every log line</strong>, so you can pivot from a log to the trace.</li><li><strong>In the response headers</strong>, so a client can report it.</li><li><strong>On the error page</strong>, so a user can paste it into a support ticket.</li><li><strong>In your error tracking</strong>, so an exception links to its trace.</li><li><strong>In the support tool</strong>, so a support engineer can hand it to an engineer.</li></ul>
<p>The workflow that makes tracing valuable: customer reports a problem → support gets the trace ID → engineer opens exactly that request and sees everything.</p>
<p>Without the ID in those places, the workflow is: customer reports a problem → engineer tries to guess which of eight million traces it was.</p>
<p>That single difference determines whether the investment pays off, and it is a plumbing problem rather than a tracing problem.</p>
<h2 id="detail-two-spans-need-attributes-not-just-timing">detail two: spans need attributes, not just timing<a class="anchor" href="#detail-two-spans-need-attributes-not-just-timing" aria-label="link to this section">#</a></h2>
<p>A span that says "database query, 45 ms" is nearly useless. Which query? Which table? How many rows? Was it cached?</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">with tracer.start_as_current_span("db.query") as span:
    span.set_attribute("db.system", "postgresql")
    span.set_attribute("db.operation", "select")
    span.set_attribute("db.table", "orders")
    span.set_attribute("db.rows_returned", len(rows))
    span.set_attribute("app.tenant_id", tenant)
    span.set_attribute("app.cache_hit", cached)</code></pre></div>
<p>The attributes are what let you ask questions the waterfall cannot answer: is this slow for one tenant, is it slow when the cache misses, is it slow when the result set is large.</p>
<p><strong>Put your business identifiers on the root span.</strong> Tenant, user, plan tier, <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">feature flags</a>, client version. Then you can filter traces by them, which is how you find the pattern rather than the instance.</p>
<p>High cardinality is correct here. This is not a metrics system.</p>
<h2 id="detail-three-sample-intelligently-keep-the-interesting-ones">detail three: sample intelligently, keep the interesting ones<a class="anchor" href="#detail-three-sample-intelligently-keep-the-interesting-ones" aria-label="link to this section">#</a></h2>
<p>Full sampling is expensive and mostly wasteful — the successful, fast, boring requests are identical to each other.</p>
<p>Tail-based sampling makes the decision after the trace completes, when you know whether it was interesting:</p>
<ul><li><strong>100% of traces with an error.</strong></li><li><strong>100% of traces above a latency threshold.</strong></li><li><strong>A small percentage of everything else</strong>, for baseline comparison.</li><li><strong>100% of traces for a specific customer</strong>, when you are debugging that customer.</li></ul>
<p>That last one is worth building explicitly. A flag that says "capture everything for this tenant for the next hour" turns an unreproducible customer report into a solvable problem, and it is the single most useful debugging feature you can add to a multi-tenant system.</p>
<h2 id="the-things-that-waste-effort">the things that waste effort<a class="anchor" href="#the-things-that-waste-effort" aria-label="link to this section">#</a></h2>
<p><strong>Instrumenting everything.</strong> Auto-instrumentation gives you spans for every framework operation, most of which are noise. A trace with four hundred spans is not more informative than one with twenty; it is less, because the signal is buried.</p>
<p>Instrument the boundaries — service calls, database queries, external APIs, queue operations — and add manual spans only for genuinely expensive internal operations.</p>
<p><strong>Perfect propagation.</strong> You will have services that drop the context: a message queue without header support, a third-party integration, a legacy component. Do not block the rollout on 100% coverage. A trace with a gap is still much better than no trace.</p>
<p><strong>Building dashboards from traces.</strong> Traces answer "what happened to this request." Metrics answer "what is happening to all requests." Using traces for aggregate views is expensive and slow. Use both, for what each is good at.</p>
<h2 id="the-question-tracing-answers-that-nothing-else-does">the question tracing answers that nothing else does<a class="anchor" href="#the-question-tracing-answers-that-nothing-else-does" aria-label="link to this section">#</a></h2>
<p>Not "is the system slow" — metrics tell you that, cheaper.</p>
<p>Not "what error happened" — logs tell you that.</p>
<p>**"Why was <em>this specific request</em> slow, and what was different about it?"**</p>
<p>That is the question. It is the question you have during an incident, when one customer is affected and the aggregate metrics look fine. And it is unanswerable without tracing.</p>
<p>If your tracing setup does not make that question easy to answer in under a minute, the setup is the problem, not the concept.</p>
<h2 id="the-minimum-viable-version">the minimum viable version<a class="anchor" href="#the-minimum-viable-version" aria-label="link to this section">#</a></h2>
<p>If you are starting from nothing:</p>
<ol><li>Adopt OpenTelemetry. It is the standard, the instrumentation libraries are broad, and it keeps you portable across backends.</li><li>Propagate context across every service boundary.</li><li>Put the trace ID in every log line and every response header.</li><li>Add business identifiers to the root span.</li><li>Tail-sample: all errors, all slow, 1% of the rest.</li></ol>
<p>That is a week of work and it covers most of the value. Everything beyond it is refinement.</p>]]></content:encoded></item><item><title>The World Cup is the largest load test ever run</title><link>https://readme.news/the-world-cup-is-the-largest-load-test-ever-run/</link><guid isPermaLink="true">https://readme.news/the-world-cup-is-the-largest-load-test-ever-run/</guid><pubDate>Wed, 10 Jun 2026 09:00:00 +0000</pubDate><description>A month of synchronized global demand across three countries, sixteen cities, and every streaming platform at once.</description><content:encoded><![CDATA[<p>The tournament starts tomorrow: forty-eight teams, three host countries, sixteen venues, and a match schedule designed so that a very large fraction of the planet is watching the same thing at the same moment, repeatedly, for a month.</p>
<p>From an engineering perspective this is the most demanding recurring event in consumer computing, and almost nothing about how it works gets written up.</p>
<h2 id="the-shape-of-the-demand">the shape of the demand<a class="anchor" href="#the-shape-of-the-demand" aria-label="link to this section">#</a></h2>
<p><strong>Synchronized, not smooth.</strong> Streaming traffic for a scheduled match is a step function. Millions of concurrent sessions establish within a two-minute window around kickoff, and the ones that fail to establish are a product failure with no recovery — the user missed the start.</p>
<p><strong>Multi-peak within a session.</strong> A goal produces a spike in social traffic, in betting platforms, in messaging, in news sites, and in the streams of people switching from another match. These arrive within seconds of each other and are correlated across completely unrelated companies.</p>
<p><strong>Correlated across the industry.</strong> This is the part that is genuinely unusual. Every CDN, every mobile network, every payment processor, and every messaging platform experiences the peak simultaneously. There is no borrowing capacity from a quiet neighbor, because there is no quiet neighbor.</p>
<p><strong>Multi-region with different profiles.</strong> Matches in three countries across several <a class="xref" href="/time-zones-and-why-your-calendar-code-is-wrong/" title="Time zones, and why your calendar code is wrong">time zones</a> means the traffic profile shifts across the tournament in ways that capacity planning must anticipate.</p>
<h2 id="what-breaks-historically">what breaks, historically<a class="anchor" href="#what-breaks-historically" aria-label="link to this section">#</a></h2>
<p><strong>Payment processing at scale.</strong> Betting and merchandise platforms see enormous transaction spikes at specific moments. Payment processors are a shared dependency and their capacity is a hard constraint nobody downstream controls.</p>
<p><strong>Mobile networks in venue areas.</strong> Sixty thousand people in one place, all trying to upload video. This is a well-understood problem with an expensive solution (temporary cell capacity) and it still degrades.</p>
<p><strong>Authentication systems.</strong> Everyone logs in at once. Login is frequently the least scaled part of a streaming stack because it is not on the hot path during normal operation.</p>
<p><strong>The thundering herd on recovery.</strong> Something fails, comes back, and every client reconnects simultaneously — causing a second failure. Retry jitter is the single most important line of code in a system like this and it is routinely absent.</p>
<p><strong>Ad insertion.</strong> Server-side ad insertion personalizes streams per viewer while keeping segments cacheable. It is a hard constraint and it is a disproportionate source of live-stream incidents.</p>
<h2 id="what-the-good-operators-do">what the good operators do<a class="anchor" href="#what-the-good-operators-do" aria-label="link to this section">#</a></h2>
<p><strong>Pre-scale, do not autoscale.</strong> Autoscaling responds to load after it arrives. Instance startup is measured in tens of seconds; the demand arrives in one. For a known event at a known time, capacity is provisioned in advance and the autoscaler is a safety net, not the mechanism.</p>
<p><strong>Load shed by feature, not by user.</strong> When capacity is short, disable the recommendations, the comments, the statistics overlay — keep the video. A degraded experience for everyone beats a perfect experience for 80% and nothing for the rest.</p>
<p><strong>Multi-CDN with active steering.</strong> Not for capacity alone — for the fact that any single CDN will have a bad region on a given day. A client-side or DNS-level steering layer that measures real performance and shifts traffic is the difference between a degraded minute and an outage.</p>
<p><strong>Rehearse.</strong> Group stage matches are the rehearsal for the knockout rounds. The teams that treat early matches as production load tests, with instrumentation and a retrospective after each one, are the ones that survive the final.</p>
<p><strong>A war room with authority.</strong> Not a monitoring <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> — a room with the people who can make decisions, including the decision to turn features off, without an approval chain.</p>
<h2 id="the-generalizable-lesson">the generalizable lesson<a class="anchor" href="#the-generalizable-lesson" aria-label="link to this section">#</a></h2>
<p>The interesting property here is <strong>demand you cannot smooth, shed, or refuse</strong>.</p>
<p>Most systems get to spread load over time, queue it, or degrade gracefully by making users wait. None of those work when the product is a live event — a queue means the user misses the goal.</p>
<p>What is left is: provision for peak, make every layer redundant, degrade by feature rather than by user, and rehearse.</p>
<p>That is expensive, unglamorous, and the only thing that works. It is worth remembering the next time someone proposes autoscaling as the answer to a spike that arrives faster than a machine can boot.</p>
<p>Some capacity you have to already own.</p>]]></content:encoded></item><item><title>The queue is the architecture</title><link>https://readme.news/the-queue-is-the-architecture/</link><guid isPermaLink="true">https://readme.news/the-queue-is-the-architecture/</guid><pubDate>Fri, 15 May 2026 09:00:00 +0000</pubDate><description>Most scaling problems are solved by making something asynchronous. Most reliability problems are caused by doing it badly.</description><content:encoded><![CDATA[<p>The single most effective architectural move available to most systems is: take the slow thing out of the request path and put it in a queue.</p>
<p>It is also the move that introduces the most subtle <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a>, and the gap between "we added a queue" and "we added a queue correctly" is large.</p>
<h2 id="what-it-buys">what it buys<a class="anchor" href="#what-it-buys" aria-label="link to this section">#</a></h2>
<p><strong>Latency.</strong> The user gets a response when the work is accepted, not when it is done. A checkout that returns in 80 ms and sends the confirmation email asynchronously is a much better product than one that returns in 900 ms.</p>
<p><strong>Absorbing spikes.</strong> A queue is a buffer. Traffic that would overwhelm a synchronous system accumulates and drains. This is the difference between a slow period and an outage.</p>
<p><strong>Isolation.</strong> If the email provider is down, checkout still works. The messages accumulate and send later.</p>
<p><strong>Retry for free.</strong> A failed message goes back on the queue. A failed synchronous call is a user-visible error.</p>
<h2 id="what-it-costs">what it costs<a class="anchor" href="#what-it-costs" aria-label="link to this section">#</a></h2>
<p><strong>Eventual consistency, everywhere.</strong> The user completed checkout and the confirmation has not arrived. The record exists and the search index does not have it. Every asynchronous boundary introduces a window where the system is inconsistent, and your UI has to be honest about it.</p>
<p><strong>Debugging across the boundary.</strong> A synchronous stack trace tells you the whole story. An asynchronous failure requires correlating a producer, a broker, and a consumer, possibly hours apart.</p>
<p><strong>Ordering.</strong> Most <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> do not guarantee it, or guarantee it only within a partition. If your consumer must process events in order, that is a design constraint that reaches back into how you partition.</p>
<p><strong>Duplicate delivery.</strong> Almost all queues are at-least-once. Your consumer <em>will</em> receive the same message twice. If that is not safe, you have a bug that appears under load, weeks after launch.</p>
<h2 id="the-rules">the rules<a class="anchor" href="#the-rules" aria-label="link to this section">#</a></h2>
<p><strong>1. Consumers must be idempotent. Non-negotiable.</strong></p>
<p>At-least-once delivery means duplicates. The consumer must produce the same result whether it processes a message once or five times.</p>
<p>The usual implementation: a natural <a class="xref" href="/idempotency-is-the-only-distributed-systems-concept-you-need/" title="Idempotency is the only distributed systems concept you need">idempotency</a> key, and a record of processed keys.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def handle(msg):
    key = msg.idempotency_key
    with tx():
        if already_processed(key):
            return
        do_the_work(msg)
        mark_processed(key)</code></pre></div>
<p>The <code>mark_processed</code> must be in the same transaction as the work, or you have moved the race rather than eliminated it.</p>
<p><strong>2. Every queue needs a dead letter queue, and someone must watch it.</strong></p>
<p>A message that fails repeatedly must go somewhere. A DLQ nobody monitors is a place where data goes to be silently lost, which is worse than an error, because errors are visible.</p>
<p>Alert on DLQ depth. Not on it being non-zero — on it growing.</p>
<p><strong>3. Retry with backoff and jitter, and cap the attempts.</strong></p>
<p>Immediate retry on a failing downstream is a denial of service you are performing against yourself. Exponential backoff with jitter, a maximum attempt count, then the DLQ.</p>
<p><strong>4. Monitor queue depth and age, not just throughput.</strong></p>
<p>Throughput looks healthy right up until it does not. The metrics that tell you something is wrong:</p>
<ul><li><strong>Depth</strong> — how many messages are waiting.</li><li><strong>Oldest message age</strong> — the most useful single metric. If it is growing, your consumers cannot keep up, and you know how far behind you are in time rather than in count.</li></ul>
<p><strong>5. Decide what happens when the queue is full.</strong></p>
<p>It will be. Reject the producer, drop messages, or block? Each is right in different cases and the default is usually wrong for you. An unbounded queue is not a solution; it is a memory leak with extra steps.</p>
<p><strong>6. Keep the payload small and the reference stable.</strong></p>
<p>Put an ID in the message, not the whole object. The consumer fetches current state. This avoids stale data in the message and keeps the broker fast.</p>
<p>The exception: if you need the state <em>as it was</em> when the event occurred, put it in the message deliberately, and say so.</p>
<p><strong>7. Version your message schema from day one.</strong></p>
<p>Producers and consumers deploy independently. Old consumers will see new messages. Include a version field. Make additive changes only, or handle both shapes.</p>
<h2 id="the-thing-that-surprises-people">the thing that surprises people<a class="anchor" href="#the-thing-that-surprises-people" aria-label="link to this section">#</a></h2>
<p><strong>Queues do not reduce load. They defer it.</strong></p>
<p>If your consumer processes 100 messages per second and you produce 150, you do not have a working system with a buffer. You have a system that is failing slowly, and the queue depth graph is a countdown.</p>
<p>A queue absorbs <em>bursts</em>. It does not fix a sustained capacity deficit, and the failure mode when you use it that way is a queue that grows for six hours and then an incident where you are simultaneously behind and unable to catch up.</p>
<p>Alert on the age, watch the trend, and size the consumers for the sustained rate.</p>]]></content:encoded></item><item><title>Run the incident before the incident</title><link>https://readme.news/run-the-incident-before-the-incident/</link><guid isPermaLink="true">https://readme.news/run-the-incident-before-the-incident/</guid><pubDate>Wed, 13 May 2026 09:00:00 +0000</pubDate><description>Game days, failure injection, and the specific reason your untested runbook is wrong.</description><content:encoded><![CDATA[<p>Every runbook that has never been executed is wrong. Not might be — is. The commands have changed, the <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> moved, the person who wrote it left, and the system it describes has been modified fourteen times.</p>
<p>The only way to find out is to run it, and the only good time to run it is when nothing is actually broken.</p>
<h2 id="what-a-game-day-is">what a game day is<a class="anchor" href="#what-a-game-day-is" aria-label="link to this section">#</a></h2>
<p>A scheduled exercise where you deliberately break something in a controlled way and have the <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> rotation respond as if it were real.</p>
<p>Two hours. A stated scenario. The people who would actually respond. Real tools, real runbooks, real dashboards. Someone taking notes on everything that did not work.</p>
<p>That is the whole practice. It is not chaos engineering in the automated-random sense, though that is a good adjacent practice. It is a rehearsal.</p>
<h2 id="what-you-find-reliably">what you find, reliably<a class="anchor" href="#what-you-find-reliably" aria-label="link to this section">#</a></h2>
<p>I have never run one of these that did not find at least four of the following:</p>
<p><strong>The runbook references something that does not exist.</strong> A dashboard that was renamed, a script in a repository that was archived, an alias nobody has.</p>
<p><strong>Nobody has the access.</strong> The runbook says to restart the service. The on-call engineer does not have permission and does not know who does. This is the single most common finding.</p>
<p><strong>The dashboard does not show the thing.</strong> You built the alert. You never built the view that tells you what to do about it.</p>
<p><strong>The escalation path is a person, not a rotation.</strong> "Ask Sarah." Sarah is on vacation. Sarah left last year.</p>
<p><strong>Nobody knows the customer impact.</strong> The service is degraded. Which customers? Which features? Nobody can answer, so nobody can decide how urgent it is.</p>
<p><strong>Recovery has an undocumented step.</strong> The service restarts and does not work because a cache must be cleared first, which is knowledge that lives in one person's head.</p>
<p><strong>The communication path is unclear.</strong> Who tells the customers? Who updates the status page? Who has the credentials for the status page?</p>
<p>Every one of those is cheap to fix on a Tuesday afternoon and extremely expensive to discover at 3 a.m.</p>
<h2 id="the-scenarios-worth-running">the scenarios worth running<a class="anchor" href="#the-scenarios-worth-running" aria-label="link to this section">#</a></h2>
<p>Start with the boring ones. Exotic <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> are fun and the common ones are what actually happens.</p>
<p><strong>The database primary fails over.</strong> Does the application reconnect? How long? Does anything need a manual restart?</p>
<p><strong>A dependency returns 500s.</strong> Not down — erroring. Does your circuit breaker work? Does your retry policy make it worse?</p>
<p><strong>A dependency gets slow.</strong> Harder than down and much more common. Does your timeout fire? Do connections exhaust? Does the slowness propagate to callers?</p>
<p><strong>Disk fills up.</strong> On the database host, on the log host, on the application host. This is the single most common self-inflicted outage.</p>
<p><strong>A bad deploy.</strong> Deploy something broken to staging and time the rollback. Not the theoretical rollback — the actual one, executed by the actual on-call person.</p>
<p><strong>Certificate expiration.</strong> Set one to expire in staging. Watch what happens. Most teams discover their monitoring does not cover this.</p>
<p><strong>The person who knows is unavailable.</strong> Run a game day where the subject matter expert is explicitly not allowed to help. This finds the knowledge concentration problems that nothing else does.</p>
<h2 id="how-to-run-one-that-works">how to run one that works<a class="anchor" href="#how-to-run-one-that-works" aria-label="link to this section">#</a></h2>
<p><strong>Announce it.</strong> Do not surprise people. Surprise exercises generate resentment and teach people to distrust the process. Everyone should know it is a drill.</p>
<p><strong>Staging first, production eventually.</strong> Staging finds most of the runbook problems. Production finds the ones that only exist because staging is not production, which are real and are the ones that matter most.</p>
<p><strong>Have a stop condition.</strong> A named person who can call it off, and a defined way to revert whatever you broke.</p>
<p><strong>Write down every friction point</strong>, including small ones. "It took four minutes to find the right dashboard" is a real finding.</p>
<p><strong>Fix things within a week.</strong> A game day that produces a list nobody acts on is theater, and the second one will have lower attendance.</p>
<p><strong>Do it quarterly.</strong> Systems change. A runbook validated a year ago is a runbook that has not been validated.</p>
<h2 id="the-cultural-part">the cultural part<a class="anchor" href="#the-cultural-part" aria-label="link to this section">#</a></h2>
<p>The purpose is to find gaps in the system, not gaps in people.</p>
<p>If someone cannot resolve the scenario, that is a finding about documentation, tooling, or access — not about them. Say this explicitly before you start, and mean it, or people will optimize for looking competent rather than for surfacing problems.</p>
<p>The best outcome of a game day is a long list of things that went wrong, discovered by people who were not stressed, on a schedule, with time to fix them.</p>]]></content:encoded></item><item><title>Observability costs more than the thing it observes</title><link>https://readme.news/observability-costs-more-than-the-thing-it-observes/</link><guid isPermaLink="true">https://readme.news/observability-costs-more-than-the-thing-it-observes/</guid><pubDate>Fri, 10 Apr 2026 09:00:00 +0000</pubDate><description>At some point your telemetry bill exceeded your compute bill and nobody noticed. Here&#x27;s how to fix it without going blind.</description><content:encoded><![CDATA[<p>A meaningful number of organizations now spend more on observability tooling than on the compute being observed. That is not automatically wrong — visibility has real value — but it is almost never a decision anyone made deliberately.</p>
<h2 id="how-it-happens">how it happens<a class="anchor" href="#how-it-happens" aria-label="link to this section">#</a></h2>
<p><strong>Logs at debug level in production</strong>, because someone turned it on during an incident in 2023 and nobody turned it off.</p>
<p><strong>Every metric at every dimension.</strong> A counter with five labels, each with a hundred values, is ten billion time series. Cardinality multiplies and the pricing follows.</p>
<p><strong>100% trace sampling</strong> because sampling felt like giving something up.</p>
<p><strong>Retention set to the maximum</strong> because nobody knew what to pick and the default was generous.</p>
<p><strong>Duplicate pipelines.</strong> Logs going to two vendors during a migration that never completed.</p>
<p>Each decision was locally reasonable. The bill is the sum.</p>
<h2 id="the-reduction-in-order-of-return">the reduction, in order of return<a class="anchor" href="#the-reduction-in-order-of-return" aria-label="link to this section">#</a></h2>
<p><strong>1. Sample traces intelligently.</strong> You do not need every trace. You need:</p>
<ul><li>All traces that errored.</li><li>All traces above a latency threshold.</li><li>A small percentage of the rest, for baseline.</li></ul>
<p>Tail-based sampling — decide after the trace completes, when you know whether it is interesting — typically cuts volume by 90%+ while keeping essentially all diagnostic value. This is the single largest win available.</p>
<p><strong>2. Fix your cardinality.</strong> Find the metrics with the most series. There will be one or two that dominate.</p>
<p>The usual culprits: user ID as a metric label, URL path with IDs in it (<code>/user/12345/profile</code> instead of <code>/user/:id/profile</code>), and <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a> as labels.</p>
<p><strong>Metrics are for aggregates. Events are for specifics.</strong> A user ID belongs in a structured event where it costs one field, not in a metric label where it multiplies the series count.</p>
<p><strong>3. Drop the logs you never query.</strong> Audit what you actually search. Most organizations find that a large fraction of log volume is from a handful of noisy sources nobody has ever looked at — <a class="xref" href="/your-monitoring-is-measuring-the-wrong-nines/" title="Your monitoring is measuring the wrong nines">health check</a> logs, framework debug output, successful-request logs that duplicate a metric.</p>
<p>Route those to cheap object storage rather than an indexed store, or drop them.</p>
<p><strong>4. Tier your retention.</strong> You do not need 90 days of everything.</p>
<ul><li>Metrics: long retention, they are small when aggregated.</li><li>Traces: short, days. You investigate recent things.</li><li>Logs: short and hot for search, long and cold in object storage for compliance.</li></ul>
<p><strong>5. Aggregate at the source.</strong> Emit a histogram, not a thousand individual timing events. Most agents can pre-aggregate and it moves cost from the vendor to your own process, where it is much cheaper.</p>
<h2 id="what-not-to-cut">what not to cut<a class="anchor" href="#what-not-to-cut" aria-label="link to this section">#</a></h2>
<p>Be careful here, because the failure mode of aggressive cost reduction is finding out during an incident that you deleted the thing you needed.</p>
<p><strong>Keep all errors.</strong> Never sample errors. They are rare and they are the whole point.</p>
<p><strong>Keep the request-level events for your critical journeys.</strong> One wide structured event per request, with all the context, is the highest value telemetry you have per byte. It is cheaper than the interior logging it replaces.</p>
<p><strong>Keep enough cardinality to slice by customer and region.</strong> Aggregate-only metrics hide the failures that matter — "one tenant is completely broken" looks identical to "everyone is slightly slower" in a global average.</p>
<h2 id="the-structural-fix">the structural fix<a class="anchor" href="#the-structural-fix" aria-label="link to this section">#</a></h2>
<p><strong>Move from logs to events.</strong> The expensive pattern is many log lines per request, each a string, indexed for full-text search.</p>
<p>The cheap pattern is one structured event per unit of work, with everything you know attached as fields, queried by field rather than by text.</p>
<p>Fewer records, more information per record, dramatically cheaper to store and query. This is the change that actually fixes the cost problem rather than trimming it.</p>
<p><strong>Own your pipeline.</strong> An open telemetry collector between your services and your vendor lets you filter, sample, aggregate, and route without changing application code, and without being locked into one destination's pricing model.</p>
<p>That is worth setting up before your bill is a problem, because doing it under cost pressure means making decisions in a hurry.</p>
<h2 id="the-question-to-ask-quarterly">the question to ask quarterly<a class="anchor" href="#the-question-to-ask-quarterly" aria-label="link to this section">#</a></h2>
<p>For each telemetry stream: <strong>when did we last use this to answer a question?</strong></p>
<p>Anything nobody has queried in a quarter is a candidate for deletion or cold storage. Run the audit. The answer is usually uncomfortable and the savings are usually large.</p>]]></content:encoded></item><item><title>Rate limiting: the four algorithms and when each is wrong</title><link>https://readme.news/rate-limiting-the-four-algorithms-and-when-each-is-wrong/</link><guid isPermaLink="true">https://readme.news/rate-limiting-the-four-algorithms-and-when-each-is-wrong/</guid><pubDate>Mon, 06 Apr 2026 09:00:00 +0000</pubDate><description>Token bucket, leaky bucket, fixed window, sliding window. They fail differently and the differences matter.</description><content:encoded><![CDATA[<p>Rate limiting looks like a solved problem until you have to pick an algorithm, at which point the choice determines your failure mode.</p>
<h2 id="fixed-window">fixed window<a class="anchor" href="#fixed-window" aria-label="link to this section">#</a></h2>
<p>Count requests per fixed interval. Reset the counter at the boundary.</p>
<div class="code"><pre><code>key: user:1234:minute:2026-04-06T14:23
incr, expire 60s, reject if &gt; limit</code></pre></div>
<p><strong>Pros:</strong> trivial to implement, one counter, minimal memory.</p>
<p><strong>Cons:</strong> the boundary problem. A client can send the full limit at 14:23:59 and the full limit again at 14:24:00 — double the intended rate, in one second, legally.</p>
<p><strong>Use when:</strong> the limit is generous relative to the burst you can absorb, and simplicity matters more than precision. Most internal services.</p>
<h2 id="sliding-window-log">sliding window log<a class="anchor" href="#sliding-window-log" aria-label="link to this section">#</a></h2>
<p>Store a timestamp per request. Count the ones inside the window.</p>
<p><strong>Pros:</strong> exactly correct. No boundary artifacts.</p>
<p><strong>Cons:</strong> memory proportional to the limit times the number of clients. A limit of 10,000 per hour per user across a million users is a lot of timestamps.</p>
<p><strong>Use when:</strong> limits are small, precision matters, and client count is bounded. API keys with strict quotas.</p>
<h2 id="sliding-window-counter">sliding window counter<a class="anchor" href="#sliding-window-counter" aria-label="link to this section">#</a></h2>
<p>The practical compromise. Keep counters for the current and previous window, interpolate based on how far into the current window you are.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def allowed(now, limit, window):
    cur_start = now - (now % window)
    elapsed = (now - cur_start) / window
    estimate = prev_count * (1 - elapsed) + cur_count
    return estimate &lt; limit</code></pre></div>
<p><strong>Pros:</strong> two counters per client, no boundary problem in practice, cheap.</p>
<p><strong>Cons:</strong> an approximation. Can be slightly wrong at the edges under a very uneven request distribution.</p>
<p><strong>Use when:</strong> this is the default. It is what most production rate limiters actually do and it is almost always the right choice.</p>
<h2 id="token-bucket">token bucket<a class="anchor" href="#token-bucket" aria-label="link to this section">#</a></h2>
<p>A bucket holds tokens, refilled at a constant rate up to a maximum. Each request consumes one. Empty bucket, request rejected.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def allowed(bucket, now, rate, capacity):
    bucket.tokens = min(capacity, bucket.tokens + (now - bucket.last) * rate)
    bucket.last = now
    if bucket.tokens &gt;= 1:
        bucket.tokens -= 1
        return True
    return False</code></pre></div>
<p><strong>Pros:</strong> allows bursts up to the bucket capacity while enforcing a long-run average. This matches how real clients behave — idle, then a burst of activity — much better than a strict window.</p>
<p><strong>Cons:</strong> two parameters instead of one, and people set them without thinking about what the burst allowance means.</p>
<p><strong>Use when:</strong> you want to allow legitimate bursts. Most user-facing APIs. This is the second-best default after sliding window counter and is better when burstiness is expected.</p>
<h2 id="leaky-bucket">leaky bucket<a class="anchor" href="#leaky-bucket" aria-label="link to this section">#</a></h2>
<p>A queue drained at a constant rate. Requests enter the queue; overflow is rejected.</p>
<p><strong>Pros:</strong> output rate is perfectly smooth, which is what you want when you are protecting a downstream system that cannot handle bursts at all.</p>
<p><strong>Cons:</strong> adds latency — requests wait in the queue. Not appropriate for interactive traffic.</p>
<p><strong>Use when:</strong> you are shaping traffic to a fixed-capacity downstream, like a third-party API with a hard rate limit, or a legacy system that falls over above a threshold.</p>
<h2 id="the-parts-everyone-gets-wrong">the parts everyone gets wrong<a class="anchor" href="#the-parts-everyone-gets-wrong" aria-label="link to this section">#</a></h2>
<p><strong>Rate limiting the wrong key.</strong> Limiting by IP breaks for users behind NAT and does nothing against a distributed attacker. Limit by authenticated identity where you have one, and by IP only as a coarse pre-auth defense.</p>
<p><strong>Not telling the client anything useful.</strong> A 429 with no information is hostile. Send the standard headers:</p>
<div class="code"><pre><code>RateLimit-Limit: 100
RateLimit-Remaining: 0
RateLimit-Reset: 42
Retry-After: 42</code></pre></div>
<p>A client that knows when to retry will retry then. A client that does not will hammer you.</p>
<p><strong>No jitter on the client side.</strong> If a thousand clients are told to retry in 42 seconds, they all retry at exactly the same moment. Always jitter, and say so in your API documentation.</p>
<p><strong>Rate limiting at the wrong layer.</strong> Limiting at the application means the request already consumed a connection, a thread, and possibly a database query. For abuse protection, limit at the edge. For fairness between legitimate users, limit in the application where you know who they are.</p>
<p><strong>Ignoring cost variance.</strong> Not all requests are equal. A search that scans a million rows and a <a class="xref" href="/your-monitoring-is-measuring-the-wrong-nines/" title="Your monitoring is measuring the wrong nines">health check</a> are both "one request." Weight by cost, or rate-limit expensive endpoints separately, or you will limit the wrong thing.</p>
<p><strong>Global counters in a distributed system.</strong> Exact global rate limiting requires coordination on every request, which is a latency and availability problem. The usual answer is per-node limits with the total divided across nodes, accepting some imprecision, or a shared store with local <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> and periodic reconciliation.</p>
<p>Decide which imprecision you can live with, deliberately, rather than discovering it during an incident.</p>
<h2 id="the-one-line-recommendation">the one-line recommendation<a class="anchor" href="#the-one-line-recommendation" aria-label="link to this section">#</a></h2>
<p>Sliding window counter for fairness, token bucket where bursts are legitimate, limit by authenticated identity, always send <code>Retry-After</code>, always jitter.</p>]]></content:encoded></item><item><title>On-call is a design problem</title><link>https://readme.news/on-call-is-a-design-problem/</link><guid isPermaLink="true">https://readme.news/on-call-is-a-design-problem/</guid><pubDate>Wed, 25 Feb 2026 09:00:00 +0000</pubDate><description>If your rotation is painful, that&#x27;s information about your architecture, not about your people&#x27;s resilience.</description><content:encoded><![CDATA[<p>Bad on-call is treated as a fact of life, a personal endurance test, or a staffing problem. It is none of those. It is a measurement of your system's design, delivered directly to your team's sleep.</p>
<h2 id="what-the-pages-are-telling-you">what the pages are telling you<a class="anchor" href="#what-the-pages-are-telling-you" aria-label="link to this section">#</a></h2>
<p>Every page is a statement about the system.</p>
<p><strong>A page that requires a human to restart something</strong> says the system cannot recover from a condition it will encounter again. That is automatable and the automation is usually a supervisor and a <a class="xref" href="/your-monitoring-is-measuring-the-wrong-nines/" title="Your monitoring is measuring the wrong nines">health check</a>.</p>
<p><strong>A page for a transient issue that resolved itself</strong> says your alert threshold is wrong, or your alert lacks a duration condition, or the underlying flakiness needs a retry with backoff.</p>
<p><strong>A page where the runbook is "look at the <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> and see if it is bad"</strong> is not a page. That is a dashboard, and it should be checked during business hours.</p>
<p><strong>A page that only one person can resolve</strong> is a knowledge distribution failure and a single point of human failure. The fix is documentation and rotation of the work, not heroism.</p>
<p><strong>A page at 3 a.m. for something that could have waited until 9 a.m.</strong> says nobody has classified alerts by urgency. Not everything that is wrong is urgent.</p>
<h2 id="the-audit">the audit<a class="anchor" href="#the-audit" aria-label="link to this section">#</a></h2>
<p>Pull the last ninety days of pages. For each one:</p>
<ul><li>What was the user impact? (Often: none.)</li><li>Could it have been automated?</li><li>Could it have waited?</li><li>Did the runbook exist and was it correct?</li><li>What was the actual fix?</li></ul>
<p>Then categorize:</p>
<p><strong>Should not have paged</strong> — no user impact or no urgency. Delete the alert or downgrade it. This is usually a large fraction and deleting alerts is the highest value hour available to most teams.</p>
<p><strong>Should have been automatic</strong> — the fix was mechanical. Automate it. If the runbook says "run this command," a computer can run that command.</p>
<p><strong>Genuine incidents</strong> — real user impact requiring judgment. These are the ones on-call exists for, and there should not be many.</p>
<p>A healthy rotation is a small number of genuine incidents. If you are getting paged more than a couple of times per week, the problem is not that your systems are complex.</p>
<h2 id="the-design-changes-that-reduce-pages">the design changes that reduce pages<a class="anchor" href="#the-design-changes-that-reduce-pages" aria-label="link to this section">#</a></h2>
<p><strong>Make things self-healing.</strong> Restart on crash. Retry with exponential backoff and jitter. Circuit-break to a degraded mode. Shed load rather than falling over. Each of these converts a page into a metric.</p>
<p><strong>Make degradation graceful and explicit.</strong> Decide in advance which features can be disabled. <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">Feature flags</a> that turn off the expensive path let a page become "turn off recommendations, investigate Monday."</p>
<p><strong>Make everything reversible fast.</strong> A large fraction of incidents are caused by a deploy. If rollback is one command and takes ninety seconds, the incident is ninety seconds long. If it requires reversing a migration, it is four hours.</p>
<p><strong>Add capacity headroom.</strong> Running at 85% utilization to save money means every spike is a page. Headroom is cheaper than the human cost, and much cheaper than the turnover.</p>
<h2 id="the-human-side">the human side<a class="anchor" href="#the-human-side" aria-label="link to this section">#</a></h2>
<p><strong>Pay for it.</strong> On-call is work performed outside working hours with real cost to the person's life. Compensate it explicitly — money or time off, and enough that the cost is visible to whoever decides whether to fix the alert noise.</p>
<p>Unpaid on-call means the cost is borne entirely by the person and is invisible to the organization, which guarantees it never improves.</p>
<p><strong>Time off after a bad night.</strong> Non-negotiable, automatic, not something the person has to ask for.</p>
<p><strong>Never one person.</strong> A primary and a secondary, always. The primary needs to be able to escalate without feeling like a failure.</p>
<p><strong>Rotate wide.</strong> If only three people can be on-call, they will burn out and leave, and the knowledge leaves with them.</p>
<p><strong>The person who was paged decides what gets fixed.</strong> Give the on-call engineer authority to prioritize the follow-up work. They have the information and the motivation, and nobody else has both.</p>
<h2 id="the-metric-to-track">the metric to track<a class="anchor" href="#the-metric-to-track" aria-label="link to this section">#</a></h2>
<p>Pages per on-call shift, trended over quarters.</p>
<p>If it is flat or rising, nobody is fixing the causes. That is a management failure, not a resilience failure, and the outcome of ignoring it is attrition — which is more expensive than every fix combined.</p>]]></content:encoded></item><item><title>Your monitoring is measuring the wrong nines</title><link>https://readme.news/your-monitoring-is-measuring-the-wrong-nines/</link><guid isPermaLink="true">https://readme.news/your-monitoring-is-measuring-the-wrong-nines/</guid><pubDate>Tue, 27 Jan 2026 09:00:00 +0000</pubDate><description>Availability percentages are a compliance artifact. Here&#x27;s what to measure if you want to know whether users are having a bad time.</description><content:encoded><![CDATA[<p>Your <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> says 99.95% availability. Your support queue says otherwise. Both are correct, and the dashboard is measuring something that does not matter.</p>
<h2 id="what-available-usually-means">what "available" usually means<a class="anchor" href="#what-available-usually-means" aria-label="link to this section">#</a></h2>
<p>For most teams it means: a health check endpoint returned 200 within a timeout, from a monitoring location, on a schedule.</p>
<p>That measures whether a process is running. It does not measure whether anyone can do anything.</p>
<p>Real user-visible failures that a health check will not catch:</p>
<ul><li>The API returns 200 with an empty result set because a downstream cache is cold.</li><li>Requests succeed at p50 and take 30 seconds at p95, and the p95 is entirely one large customer.</li><li>Login works, but only for users whose sessions were created before the deploy.</li><li>Everything works except in one region, which is 8% of traffic and 40% of revenue.</li><li>Writes succeed and are silently not persisted because a queue consumer is wedged.</li></ul>
<p>Every one of those is a full outage for the affected users and none of them move a health-check-based availability number.</p>
<h2 id="what-to-measure-instead">what to measure instead<a class="anchor" href="#what-to-measure-instead" aria-label="link to this section">#</a></h2>
<p><strong>Measure user journeys, not endpoints.</strong> Pick the three to five things users come to your product to do. Checkout. Send a message. Run a query. Deploy. Instrument those end to end, and define success as "the journey completed correctly," not "each service returned 200."</p>
<p><strong>Measure from the client where you can.</strong> Server-side metrics cannot see DNS failures, TLS problems, CDN issues, or the user's terrible network. A meaningful fraction of real user-visible failure is invisible from inside your datacenter.</p>
<p><strong>Use success rate, not uptime.</strong> The fraction of requests that succeeded, over a window. This handles partial failure correctly, which uptime does not. A service serving 60% of requests successfully is not "up," and a binary metric says it is.</p>
<p><strong>Slice by everything.</strong> Aggregate metrics hide the failures that matter. The same success rate can mean "everyone has a slightly bad time" or "one tenant is completely broken," and those require completely different responses.</p>
<p>Slice by: customer, region, client version, endpoint, <a class="xref" href="/the-graph-nobody-drew/" title="The graph nobody drew">feature flag</a>. If you cannot slice by these, that is your observability gap.</p>
<p><strong>Set a latency threshold and count violations.</strong> "p99 latency" as a number is hard to alert on and easy to game. "Percentage of requests slower than 2 seconds" is a success rate, comparable across time, and directly meaningful. If more than 1% of checkouts take longer than 2 seconds, something is wrong, and you can say that in a sentence a product manager understands.</p>
<h2 id="the-slo-framing-without-ceremony">the SLO framing, without ceremony<a class="anchor" href="#the-slo-framing-without-ceremony" aria-label="link to this section">#</a></h2>
<p>You do not need a formal SLO program. You need three numbers per critical journey:</p>
<ol><li><strong>What fraction of attempts must succeed.</strong> 99.9%, say.</li><li><strong>What counts as success.</strong> Completed, correct, within a latency bound.</li><li><strong>Over what window.</strong> 28 days, rolling.</li></ol>
<p>The difference between your target and 100% is your error budget. When you are burning it fast, stop shipping and fix reliability. When you have plenty, ship faster and take more risk.</p>
<p>That is the entire value of the framework: it converts "should we prioritize reliability or features" from an argument into an arithmetic question. Everything else in the SLO literature is optional.</p>
<h2 id="alerting-on-the-right-thing">alerting on the right thing<a class="anchor" href="#alerting-on-the-right-thing" aria-label="link to this section">#</a></h2>
<p><strong>Alert on symptoms, not causes.</strong> "Checkout success rate below 99%" pages someone. "CPU above 80%" does not, because high CPU is sometimes fine and low CPU with a broken service is not.</p>
<p><strong>Alert on burn rate, not instantaneous values.</strong> A single failed request should not page anyone. Burning a month's error budget in an hour should page everyone. Multi-window burn-rate alerting is the single best improvement most teams can make to their paging, and it dramatically reduces false pages.</p>
<p><strong>Every page should have an action.</strong> If the runbook is "look at it and see if it recovers," it is not a page. It is a dashboard.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>Take your last five user-reported incidents. For each, ask: did an alert fire before the user reported it?</p>
<p>If the answer is mostly no, your monitoring is measuring the wrong things, and no amount of additional dashboards will fix that. The gap is not visibility. It is that you are watching the system instead of watching the users.</p>]]></content:encoded></item><item><title>The postmortem that changes something</title><link>https://readme.news/the-postmortem-that-changes-something/</link><guid isPermaLink="true">https://readme.news/the-postmortem-that-changes-something/</guid><pubDate>Tue, 16 Dec 2025 09:00:00 +0000</pubDate><description>Most incident reviews produce a document and a ticket that never gets done. Here&#x27;s the difference.</description><content:encoded><![CDATA[<p>Most organizations run blameless postmortems. Most organizations also have the same class of incident repeatedly. Both of those things are true simultaneously and nobody finds it strange.</p>
<h2 id="the-failure-pattern">the failure pattern<a class="anchor" href="#the-failure-pattern" aria-label="link to this section">#</a></h2>
<p>The typical postmortem:</p>
<ol><li>Timeline of what happened. Accurate, detailed, useful.</li><li>Root cause. Usually a single technical fact.</li><li>Action items. Five to twelve of them.</li><li>Filed. Two action items get done. The rest age out.</li></ol>
<p>Six months later, a similar incident, with a similar document.</p>
<p>Three things are wrong here.</p>
<h2 id="problem-one-root-cause-is-singular">problem one: "root cause" is singular<a class="anchor" href="#problem-one-root-cause-is-singular" aria-label="link to this section">#</a></h2>
<p>Complex systems do not fail because of one thing. They fail because several conditions aligned, each of which was individually survivable.</p>
<p>The Cloudflare outage in November: a permissions change, a query that returned duplicates, a fixed buffer size, an error path that panicked instead of degrading, and a config pipeline that propagated globally in minutes. Remove any one and it does not happen, or it is much smaller.</p>
<p>Picking one and calling it "the root cause" means you fix one and leave four.</p>
<p>Better framing: <strong>contributing factors</strong>, plural, each with its own assessment of whether it is worth addressing. Some will not be — that is a legitimate decision if it is made explicitly.</p>
<h2 id="problem-two-action-items-without-owners-and-dates-are-wishes">problem two: action items without owners and dates are wishes<a class="anchor" href="#problem-two-action-items-without-owners-and-dates-are-wishes" aria-label="link to this section">#</a></h2>
<p>An action item that says "improve monitoring for the config pipeline" with no owner, no date, and no definition of done is a sentence, not a plan.</p>
<p>The fix is unglamorous:</p>
<ul><li><strong>Every action item has one named person.</strong> Not a team. A person.</li><li><strong>Every action item has a date.</strong> If nobody will commit to a date, it is not going to happen and you should delete it and say so.</li><li><strong>Every action item has a definition of done</strong> that someone else could verify.</li><li><strong>They go in the same backlog as feature work</strong>, prioritized against it. An action item in a separate "incident follow-up" list that is never sprint planned is a list of things that will not be done.</li></ul>
<p>And the one that actually forces it: <strong>review the open action items at the start of the next postmortem.</strong> Nothing motivates completion like a room full of people looking at your undone item from last quarter's incident while discussing this quarter's similar one.</p>
<h2 id="problem-three-nobody-asks-about-the-near-misses">problem three: nobody asks about the near misses<a class="anchor" href="#problem-three-nobody-asks-about-the-near-misses" aria-label="link to this section">#</a></h2>
<p>The incidents you review are the ones that broke through. For every one, there were several that did not — a bad deploy caught by a canary, a config error someone noticed in review, a query that would have taken down the database if it had run on Monday instead of Sunday.</p>
<p>Those contain the same information at a fraction of the cost, and almost nobody collects them.</p>
<p>Add a lightweight channel for it. "I nearly broke prod today, here is how." No document, no meeting, no blame. Just a note. The pattern that emerges over a quarter is more valuable than any individual postmortem.</p>
<h2 id="the-questions-that-produce-useful-findings">the questions that produce useful findings<a class="anchor" href="#the-questions-that-produce-useful-findings" aria-label="link to this section">#</a></h2>
<p>Replace "what was the root cause" with:</p>
<ul><li><strong>"What made this hard to detect?"</strong> Detection time is usually the largest component of impact and is the most improvable.</li><li><strong>"What made this hard to diagnose?"</strong> Usually missing observability. This produces the highest-value action items.</li><li><strong>"What made recovery slow?"</strong> Often a missing runbook, a missing <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">kill switch</a>, or a rollback that was not actually tested.</li><li><strong>"Who knew something that would have helped, and why did that not reach the responders?"</strong> This is an organizational question and it is frequently the real finding.</li><li><strong>"What did we do that helped?"</strong> Genuinely important. Practices that worked should be named so they get kept.</li></ul>
<h2 id="the-blameless-part-done-correctly">the blameless part, done correctly<a class="anchor" href="#the-blameless-part-done-correctly" aria-label="link to this section">#</a></h2>
<p>Blameless does not mean nobody made a mistake. It means the analysis focuses on why the mistake was possible and easy, rather than on the person.</p>
<p>"Alice deployed without running the tests" is blame and it is also useless. "The deploy path does not require tests to pass, and the shortcut that skips them is the fastest way to deploy" is the same fact stated in a way you can act on.</p>
<p>If the answer to "why did they do that" is "because the system made it easy and the correct path was hard," you have found something.</p>
<p>If the answer is genuinely "they were careless," you still fix the system, because the next person will also be careless eventually. Humans are the constant; the system is the variable.</p>]]></content:encoded></item><item><title>Cloudflare falls over because of a config file</title><link>https://readme.news/cloudflare-falls-over-because-of-a-config-file/</link><guid isPermaLink="true">https://readme.news/cloudflare-falls-over-because-of-a-config-file/</guid><pubDate>Thu, 20 Nov 2025 09:00:00 +0000</pubDate><description>A permissions change doubles the size of a generated feature file, which overflows a fixed-size buffer, which 500s a fifth of the web.</description><content:encoded><![CDATA[<p>Cloudflare had a significant outage on Tuesday, returning 5xx errors across a large portion of its network for several hours. Their published postmortem is detailed and worth reading in full.</p>
<p>The chain of events is a small masterpiece of the genre.</p>
<h2 id="what-happened">what happened<a class="anchor" href="#what-happened" aria-label="link to this section">#</a></h2>
<p>A database permissions change caused a query that generates a Bot Management feature configuration file to return <strong>duplicate rows</strong>. The query had been returning one row per feature; after the permissions change it returned rows from multiple underlying schemas.</p>
<p>The generated file therefore roughly doubled in size.</p>
<p>The proxy that consumes this file preallocates a fixed-size buffer sized against a limit of 200 features — comfortably above the ~60 actually in use. The doubled file exceeded that limit.</p>
<p>The Rust code handling this hit an unrecoverable error path and the proxy panicked rather than degrading. Because the configuration file propagates network-wide every few minutes, the failure propagated network-wide within minutes.</p>
<p>Recovery was complicated by the fact that the bad file kept regenerating and redeploying.</p>
<h2 id="the-lessons-which-are-old">the lessons, which are old<a class="anchor" href="#the-lessons-which-are-old" aria-label="link to this section">#</a></h2>
<p><strong>Configuration is code and needs the same rigor.</strong> This was not a code deploy. It was a data change that propagated to production automatically with no staging, no canary, and no validation. Config deployment pipelines are consistently held to a lower standard than code deployment pipelines, and config causes a large share of major outages.</p>
<p><strong>Fixed-size limits need to fail soft.</strong> The 200-feature limit was reasonable. Panicking when exceeded was not. The correct behavior for a proxy encountering an oversized config is to log loudly, alert, and continue with the previous known-good version.</p>
<p><strong>Validate generated artifacts before propagation.</strong> A size check, a schema check, a sanity check on row count against the previous version — any of these would have caught it. Generated files should be validated as rigorously as user input, because the generator can be wrong.</p>
<p><strong>Blast radius follows deployment speed.</strong> Config that propagates globally in minutes is a feature until it propagates a bad config globally in minutes. Staged rollout applies to configuration too.</p>
<p><strong>Have a <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">kill switch</a> for automated pipelines.</strong> Much of the recovery time went to stopping the thing that kept redeploying the bad file. Every automated deployment path needs a way to stop it that does not require fixing the underlying problem first.</p>
<h2 id="the-rust-note">the Rust note<a class="anchor" href="#the-rust-note" aria-label="link to this section">#</a></h2>
<p>This will be used as an argument about Rust, and it should not be.</p>
<p>The panic was a deliberate choice at that call site — the code used a construct that terminates on error rather than propagating it. That is a design decision about error handling, available in any language. In C the equivalent code would have written past the buffer, which is worse.</p>
<p>The actual lesson is about where you choose to make errors fatal. In a proxy handling live traffic, almost nothing should be fatal. Degrade, alert, continue. "Fail fast" is good advice for a batch job and bad advice for a load balancer.</p>
<h2 id="the-credit-due">the credit due<a class="anchor" href="#the-credit-due" aria-label="link to this section">#</a></h2>
<p>Cloudflare published a detailed technical postmortem within a day, named the specific code path, and did not hide behind "an issue with a third-party provider."</p>
<p>That is how it should be done and it is rarer than it should be. A company that publishes real postmortems earns more trust than one that never has visible incidents, because the second one is not telling you about them.</p>]]></content:encoded></item><item><title>us-east-1 goes down and takes a large chunk of the internet with it</title><link>https://readme.news/us-east-1-goes-down-and-takes-a-large-chunk-of-the-internet-with-it/</link><guid isPermaLink="true">https://readme.news/us-east-1-goes-down-and-takes-a-large-chunk-of-the-internet-with-it/</guid><pubDate>Tue, 21 Oct 2025 09:00:00 +0000</pubDate><description>A DNS race condition in DynamoDB&#x27;s automation cascades across dozens of AWS services. The lesson is about coupling, not DNS.</description><content:encoded><![CDATA[<p>AWS's us-east-1 region suffered a multi-hour outage yesterday that affected a very large number of services and, through them, a very large fraction of consumer internet applications.</p>
<p>AWS's public summary attributes the trigger to a latent race condition in the automation that manages DynamoDB's DNS records, which resulted in an empty record set for a regional endpoint and no automatic recovery path.</p>
<h2 id="the-cascade">the cascade<a class="anchor" href="#the-cascade" aria-label="link to this section">#</a></h2>
<p>The failure sequence is the interesting part.</p>
<p>DynamoDB's endpoint became unresolvable. That alone would be bad. What made it a regional event is that <strong>an enormous number of AWS's own services use DynamoDB internally.</strong> The EC2 instance launch path, IAM's control plane, Lambda's invocation machinery, and dozens of others depend on it.</p>
<p>So the failure propagated: DynamoDB down means new EC2 instances cannot launch, which means autoscaling cannot replace failing capacity, which means load shedding, which means more failures. Network Load Balancer health checks destabilized. The recovery itself was slowed by the backlog of queued work that had accumulated.</p>
<p>This is textbook <strong>metastable failure</strong>: a system that is stable under normal load and stable under no load, but which, once pushed past a threshold, sustains its own failure through retry amplification and queue buildup even after the original trigger is fixed.</p>
<h2 id="why-us-east-1">why us-east-1<a class="anchor" href="#why-us-east-1" aria-label="link to this section">#</a></h2>
<p>It is the oldest region, the largest, and the default in approximately every tutorial ever written. Several global AWS control planes are homed there — IAM, CloudFront configuration, Route 53's control plane, and others. That means a us-east-1 event has global blast radius even for customers with no resources in the region.</p>
<p>That architecture is a historical artifact. It is also extremely hard to change now, which is a lesson about early decisions in systems that grow.</p>
<h2 id="the-honest-customer-takeaway">the honest customer takeaway<a class="anchor" href="#the-honest-customer-takeaway" aria-label="link to this section">#</a></h2>
<p>The reflexive response is "multi-region." Before you spend a year on that, do the arithmetic.</p>
<p><strong>Multi-region active-active is genuinely hard.</strong> Data consistency across regions, failover testing that actually works, doubled infrastructure cost, and a substantially more complex system that fails in new ways. Many organizations that attempt it end up with a system that is <em>less</em> reliable overall because the complexity introduces more <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> than the regional risk it removes.</p>
<p><strong>The dependency you cannot escape.</strong> If your multi-region architecture depends on a global control plane that lives in us-east-1, you did not achieve independence. Check this specifically. A lot of people discovered it yesterday.</p>
<p><strong>What is actually worth doing, in order:</strong></p>
<ol><li><strong>Know your dependencies.</strong> Most teams cannot enumerate what their service requires to start. Write it down. The exercise is revealing.</li><li><strong>Static stability.</strong> Design so existing capacity keeps serving when the control plane is unavailable. If your service needs to call an API to keep running, it will stop when that API stops. Cache aggressively, fail open where safe, and do not require a control plane call on the request path.</li><li><strong>Graceful degradation.</strong> Decide in advance which features can be turned off. A checkout that works without recommendations is much better than a site that is down.</li><li><strong>Exponential backoff with jitter, and circuit breakers.</strong> Retry storms are what turns an incident into an outage. This is the single highest-leverage code change available.</li><li><strong>Multi-region for the tier that genuinely warrants it.</strong> Which is usually not everything.</li></ol>
<h2 id="the-industry-level-observation">the industry-level observation<a class="anchor" href="#the-industry-level-observation" aria-label="link to this section">#</a></h2>
<p>A meaningful fraction of the world's software depends on a small number of regions operated by a small number of companies. That concentration produces excellent reliability most of the time and correlated failure occasionally.</p>
<p>There is no individual fix. Every company independently choosing the most reliable provider produces exactly this concentration. It is a collective action problem, and the only actors who can address it are regulators thinking about systemic risk, who are — belatedly — starting to.</p>]]></content:encoded></item><item><title>Your logs are a product and you are shipping a bad one</title><link>https://readme.news/your-logs-are-a-product-and-you-are-shipping-a-bad-one/</link><guid isPermaLink="true">https://readme.news/your-logs-are-a-product-and-you-are-shipping-a-bad-one/</guid><pubDate>Tue, 03 Jun 2025 09:00:00 +0000</pubDate><description>Structured events, one line per unit of work, and the specific reason your grep-based debugging is slow.</description><content:encoded><![CDATA[<p>Here is a log line from a real production system, lightly anonymized:</p>
<div class="code"><pre><code>2025-06-02 14:23:11 INFO  Processing request</code></pre></div>
<p>Processing which request? For whom? Started when? Finished when? Did it succeed? This line costs money to produce, store, index, and retain, and it conveys approximately zero bits of information.</p>
<p>Multiply by four hundred million a day.</p>
<h2 id="the-mental-model-that-fixes-this">the mental model that fixes this<a class="anchor" href="#the-mental-model-that-fixes-this" aria-label="link to this section">#</a></h2>
<p>Stop thinking of logs as a narrative of what your program did. Think of them as <strong>structured events about units of work</strong>, emitted once, with everything you know attached.</p>
<p>The canonical form: one event per request, per job, per message consumed. Wide. Structured. Emitted at the end when you know how it went.</p>
<div class="code"><span class="code-lang">json</span><pre><code class="lang-json">{
  "event": "http_request",
  "trace_id": "4bf92f3577b34da6",
  "method": "POST",
  "route": "/api/orders",
  "status": 201,
  "duration_ms": 143,
  "user_id": "u_8812",
  "org_id": "org_44",
  "db_queries": 7,
  "db_ms": 89,
  "cache_hits": 3,
  "cache_misses": 1,
  "upstream_ms": 22,
  "region": "us-east-1",
  "version": "2025.06.02-a1b2c3d",
  "feature_flags": ["new_checkout", "fast_path"]
}</code></pre></div>
<p>One line. Now you can answer questions you did not think to ask when you wrote it:</p>
<ul><li>Which org has the slowest p99 on this route?</li><li>Are requests with <code>new_checkout</code> enabled making more database queries?</li><li>Did latency change after the deploy of <code>a1b2c3d</code>?</li><li>Is the cache miss rate correlated with the slow requests, or is that coincidence?</li></ul>
<p>None of those are answerable by grepping "Processing request."</p>
<h2 id="the-rules">the rules<a class="anchor" href="#the-rules" aria-label="link to this section">#</a></h2>
<p><strong>Log at the boundary of a unit of work, not at every step inside it.</strong> Interior logging is what tracing is for. If you need step-level visibility, emit spans, not log lines.</p>
<p><strong>Never log a string you have to parse later.</strong> <code>"user " + id + " failed"</code> means somebody writes a regex. Put <code>id</code> in a field.</p>
<p><strong>Attach identity to everything.</strong> Trace ID, user, org, request ID. The single most common debugging failure is having the information and being unable to correlate it.</p>
<p><strong>High cardinality is the point.</strong> The advice to avoid high-cardinality fields comes from metrics systems, where cardinality multiplies storage. Events are not metrics. <code>user_id</code> in an event is exactly what makes it useful, and any observability tool that cannot handle it is the wrong tool.</p>
<p><strong>Log the decision, not the branch.</strong> Not <code>"entering fast path"</code>. Rather <code>"fast_path": true</code> on the one event.</p>
<p><strong>Errors get context, not just stack traces.</strong> A stack trace tells you where. The fields tell you which input, which user, which version, which flag combination. Where is the easy part.</p>
<h2 id="what-this-costs">what this costs<a class="anchor" href="#what-this-costs" aria-label="link to this section">#</a></h2>
<p>Fewer lines, each larger. In practice, storage usually goes <em>down</em>, because you delete the interior noise. Query performance improves dramatically because you are filtering structured fields rather than doing full-text search.</p>
<p>The real cost is discipline at write time, and one afternoon to set up a logger that makes the structured path the easy path. If emitting a field requires more typing than emitting a string, people will emit strings.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>Pick an incident from the last quarter. Ask: could I have diagnosed that from logs alone, without adding new logging and redeploying?</p>
<p>If the answer is no — and it usually is — that is the gap. The information you needed was in the process's memory at the time and you did not write it down.</p>
<p>Write it down.</p>]]></content:encoded></item><item><title>In defense of the boring deploy</title><link>https://readme.news/in-defense-of-the-boring-deploy/</link><guid isPermaLink="true">https://readme.news/in-defense-of-the-boring-deploy/</guid><pubDate>Mon, 10 Mar 2025 09:00:00 +0000</pubDate><description>Blue-green, canaries, feature flags, and the deeply unfashionable practice of shipping the same way every time.</description><content:encoded><![CDATA[<p>The best deployment I ever worked on took eleven minutes, ran the same way every time, and had a rollback that was a single command anyone on the team could run from their phone. Nobody wrote a conference talk about it.</p>
<p>Deployment is a solved problem that organizations keep re-opening because the solution is boring and boring does not get anyone promoted.</p>
<h2 id="the-properties-that-matter">the properties that matter<a class="anchor" href="#the-properties-that-matter" aria-label="link to this section">#</a></h2>
<p>Strip away the tooling debates and a good deploy has four properties.</p>
<p><strong>It is the same every time.</strong> Not "mostly the same, except for the database migration, and except on Fridays, and except when Sarah does it." The same. If there is a manual step, it is in the pipeline as an approval gate, not in someone's head.</p>
<p><strong>It is reversible in under a minute.</strong> This is the one people get wrong. They build elaborate progressive rollouts and then discover that rolling back requires reverting a migration that dropped a column. Reversibility is a design constraint on your <em>changes</em>, not a feature of your deploy tool.</p>
<p><strong>Failure is detected without a human looking.</strong> If the way you find out about a bad deploy is a customer email, you do not have a deployment process, you have a deployment ritual.</p>
<p><strong>It happens often enough to be unremarkable.</strong> The failure rate of a deploy scales superlinearly with the amount of change in it. Ten small deploys are dramatically safer than one deploy containing ten changes, and this is the single most robust finding in the entire delivery-metrics literature.</p>
<h2 id="the-expand-contract-discipline">the expand-contract discipline<a class="anchor" href="#the-expand-contract-discipline" aria-label="link to this section">#</a></h2>
<p>Most rollback pain is schema pain. The fix is a discipline, not a tool:</p>
<ol><li><strong>Expand.</strong> Add the new column, nullable. Deploy. The old code ignores it.</li><li><strong>Migrate.</strong> Backfill. Dual-write from application code. Deploy. Both shapes work.</li><li><strong>Switch.</strong> Read from the new column. Deploy. Still reversible — the old column is intact.</li><li><strong>Contract.</strong> Drop the old column, days or weeks later, once you are certain.</li></ol>
<p>Four deploys instead of one. Every intermediate state is reversible. It feels slow and it is the reason you do not have a 3 a.m. incident where the rollback made things worse.</p>
<h2 id="flags-are-not-free">flags are not free<a class="anchor" href="#flags-are-not-free" aria-label="link to this section">#</a></h2>
<p><a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">Feature flags</a> decouple deploy from release and that is genuinely valuable. They also multiply your state space: <code>n</code> flags means <code>2^n</code> possible configurations, and you are testing approximately one of them.</p>
<p>Rules that keep this manageable:</p>
<ul><li>Every flag gets an owner and an expiry date at creation.</li><li>A flag that has been at 100% for a month is deleted, not left "just in case."</li><li>Flags never nest. If flag A only matters when flag B is on, you have made something unreasonable.</li><li>Kill-switch flags are a separate category with separate rules and they can live forever.</li></ul>
<p>The failure mode is flag debt: a codebase with two hundred flags where nobody knows which combinations are actually exercised in production. That is worse than no flags, because it looks like safety.</p>
<h2 id="the-actually-unfashionable-opinion">the actually unfashionable opinion<a class="anchor" href="#the-actually-unfashionable-opinion" aria-label="link to this section">#</a></h2>
<p>Most teams do not need Kubernetes to deploy well, and many teams running Kubernetes deploy worse than they did before, because the platform's complexity consumed the attention that used to go into the process.</p>
<p>If your deploys are fast, boring, and reversible, your infrastructure is correct regardless of what it is. If they are not, no amount of platform will fix it, because the problem is that nobody has treated the deploy as a product with a user.</p>
<p>The user is your team at 2 a.m. Design for them.</p>]]></content:encoded></item>
</channel>
</rss>
