<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — architecture</title>
<link>https://readme.news/tags/architecture/</link>
<atom:link href="https://readme.news/tags/architecture/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged architecture.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>The migration that never finished</title><link>https://readme.news/the-migration-that-never-finished/</link><guid isPermaLink="true">https://readme.news/the-migration-that-never-finished/</guid><pubDate>Mon, 21 Sep 2026 09:00:00 +0000</pubDate><description>Half-completed migrations are the most expensive state a system can be in, and the most common one.</description><content:encoded><![CDATA[<p>Every mature codebase has at least one: the migration that got to 70% and stopped. The new system handles most traffic, the old one handles the awkward remainder, and both are maintained forever.</p>
<p>This is worse than either finishing or never starting, and it is the default outcome unless something prevents it.</p>
<h2 id="why-it-costs-more-than-both">why it costs more than both<a class="anchor" href="#why-it-costs-more-than-both" aria-label="link to this section">#</a></h2>
<p><strong>Two systems to maintain.</strong> Every change lands twice, or lands in one and silently diverges in the other.</p>
<p><strong>Two sets of bugs</strong>, plus a third set caused by the interaction.</p>
<p><strong>Nobody knows which is authoritative.</strong> New engineers ask; the answer is "it depends."</p>
<p><strong>The benefit never arrives.</strong> The reason for the migration — delete the old thing, get the performance, simplify the model — is only realised at 100%. At 70% you have paid the full cost and collected none of the return.</p>
<p><strong>It gets harder over time.</strong> The remaining 30% is the hard 30%: the weird integrations, the customer with the bespoke arrangement, the code nobody understands. And it gets harder as the people who understood the original migration leave.</p>
<h2 id="why-it-happens">why it happens<a class="anchor" href="#why-it-happens" aria-label="link to this section">#</a></h2>
<p>Not laziness. The incentives genuinely point this way:</p>
<p><strong>The easy 70% delivers most of the visible benefit.</strong> The graph goes up, the demo works, the announcement is made. The remaining 30% has no visible reward.</p>
<p><strong>The hard cases are hard for a reason.</strong> They were skipped because someone did not know how to handle them, and that has not changed.</p>
<p><strong>Priorities move.</strong> The migration was urgent in Q1. In Q3 there is a launch.</p>
<p><strong>Nobody owns the finish.</strong> The person who drove it moved on, and completion was never anyone's explicit goal.</p>
<h2 id="the-mechanisms-that-actually-work">the mechanisms that actually work<a class="anchor" href="#the-mechanisms-that-actually-work" aria-label="link to this section">#</a></h2>
<p><strong>Name the deletion, not the migration.</strong> The project is not "migrate to the new pricing service." It is <strong>"delete the old pricing service."</strong> The deliverable is the deletion, and the migration is how you get there. This one rewording changes what people track and what counts as done.</p>
<p><strong>Set a date, publicly, at the start.</strong> Not "by end of year" — a date, in the plan, with the deletion as the milestone. Dates without deletions slip silently; a date attached to "and then this code is gone" is checkable.</p>
<p><strong>Make the old path visibly worse.</strong> Log a warning on every use. Add latency — genuinely, deliberately. Put a banner in the internal tool. The old path being comfortable is why nobody leaves it.</p>
<p><strong>Count the stragglers on a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>.</strong> "Requests still on the legacy path" as a number that goes down, reviewed weekly. Anything not measured stalls at whatever level nobody notices.</p>
<p><strong>Do the hard cases first.</strong> Backwards from the usual instinct and correct. The easy 70% will always be doable; the hard 30% is what determines whether the project is possible at all. Finding out in month one that a case cannot be migrated is a much better outcome than finding out in month nine.</p>
<p><strong>Budget the finish before starting.</strong> If you cannot fund the last 30%, do not start. A migration you cannot complete is worse than the system you have.</p>
<h2 id="the-decision-worth-making-explicitly">the decision worth making explicitly<a class="anchor" href="#the-decision-worth-making-explicitly" aria-label="link to this section">#</a></h2>
<p>Sometimes finishing genuinely is not worth it. The remaining cases are rare, the old path works, and the effort is better spent elsewhere.</p>
<p>That is a legitimate call — but it has to be <em>made</em>, written down, and the state made permanent rather than provisional:</p>
<blockquote><p>The legacy importer stays for CSV uploads from the four enterprise customers using it. New integrations use the API. We are not migrating them; the importer is now a supported, frozen component with an owner, not a migration in progress. Reviewed 2026-09-21, revisit 2028.</p></blockquote>
<p>That is a completely different thing from a stalled migration, even though the code looks identical. One is a decision with an owner. The other is a mess nobody has admitted to.</p>
<p>The cost of the second one is not the code. It is that every engineer who touches that area has to reconstruct which way things are supposed to be going, and nobody can tell them.</p>]]></content:encoded></item><item><title>Choosing an ID</title><link>https://readme.news/choosing-an-id/</link><guid isPermaLink="true">https://readme.news/choosing-an-id/</guid><pubDate>Fri, 18 Sep 2026 09:00:00 +0000</pubDate><description>Integers, UUIDs, ULIDs, or something with meaning. The decision is permanent and it leaks into everything.</description><content:encoded><![CDATA[<p>Primary key type is one of the few decisions that is genuinely hard to reverse. It propagates into every foreign key, every URL, every log line, every external integration, and every client that ever stored one.</p>
<p>Worth twenty minutes up front.</p>
<h2 id="the-options">the options<a class="anchor" href="#the-options" aria-label="link to this section">#</a></h2>
<p><strong>Auto-increment integer.</strong> Compact, fast, perfect index locality, human-readable in logs.</p>
<p>The problems are real. It leaks volume — <code>/orders/48213</code> tells a competitor how many orders you have taken. It requires a round trip to the database to learn the ID, which blocks client-side generation and batching. And it makes merging data from two systems a genuine ordeal, because both start at 1.</p>
<p><strong>UUIDv4.</strong> Random, globally unique, generatable anywhere without coordination, leaks nothing.</p>
<p>The cost is index behaviour. Random values scatter inserts across the entire B-tree, which destroys cache locality, inflates the index, and causes page splits on every insert. On a large, write-heavy table this is a genuine and measurable problem, not a theoretical one.</p>
<p><strong>UUIDv7.</strong> Time-ordered UUID: a millisecond timestamp prefix, then randomness. Keeps global uniqueness and client-side generation, restores insert locality because new rows land at the end of the index.</p>
<p>This is the default I would now recommend for most new systems. It fixes the one serious problem with v4 while keeping everything that made v4 attractive. Postgres has <code>uuidv7()</code> built in; most languages have a library.</p>
<p>The trade: it leaks creation time. Usually fine, occasionally not.</p>
<p><strong>ULID / KSUID and friends.</strong> Same idea as v7 — sortable, time-prefixed — with a more compact text encoding. Fine choices. UUIDv7 has the advantage of being a standard with native database support, which matters more than encoding length.</p>
<p><strong>Natural keys.</strong> Email, ISBN, SKU. Almost always a mistake as a primary key, because "naturally unique and never changes" turns out to be false: people change email addresses, and standards get revised. Use them as unique constraints, not as the identifier everything else points at.</p>
<h2 id="the-two-identifier-pattern">the two-identifier pattern<a class="anchor" href="#the-two-identifier-pattern" aria-label="link to this section">#</a></h2>
<p>Frequently the right answer is not to choose:</p>
<ul><li><strong>Internal key</strong>: <code>bigint</code>, auto-increment. Used for foreign keys and joins. Compact, fast, never exposed.</li><li><strong>External ID</strong>: UUIDv7 or a prefixed string. Used in URLs, APIs, logs, support conversations. Unique index on it.</li></ul>
<p>You get join performance and index locality internally, and no information leakage externally. The cost is one extra column and remembering which is which.</p>
<h2 id="prefix-your-external-ids">prefix your external IDs<a class="anchor" href="#prefix-your-external-ids" aria-label="link to this section">#</a></h2>
<p>If you expose identifiers, prefix them by type:</p>
<div class="code"><pre><code>cus_01J9F3K2M4N5P6Q7R8S9T0V1W2
ord_01J9F3K8X1Y2Z3A4B5C6D7E8F9</code></pre></div>
<p>This is a small thing with an outsized payoff:</p>
<ul><li>A support engineer can tell what an ID refers to without asking.</li><li>Passing a customer ID where an order ID belongs becomes detectable, and can be rejected at the API boundary rather than producing a confusing not-found.</li><li>Logs and error reports become self-describing.</li><li>You can rotate the encoding later without ambiguity.</li></ul>
<p>Several well-run APIs do this and it is consistently one of the things developers say they appreciate about them.</p>
<h2 id="the-practical-rules">the practical rules<a class="anchor" href="#the-practical-rules" aria-label="link to this section">#</a></h2>
<p><strong>Never expose auto-increment integers publicly.</strong> Volume leakage plus trivial enumeration.</p>
<p><strong>Do not use UUIDv4 as a clustered primary key</strong> on a table that will get large and write-heavy. If you already have, v7 for new tables and consider whether the old one needs a migration — usually it does not, but measure the index bloat before deciding.</p>
<p><strong>Store UUIDs as a native <code>uuid</code> type</strong>, not as text. Sixteen bytes versus thirty-six, and correct comparison semantics.</p>
<p><strong>Decide the case and format once</strong>, write it down, and validate at the boundary. Half your system accepting hyphenated lowercase and the other half accepting bare uppercase is a bug that surfaces years later in one integration.</p>
<p><strong>Never reuse an ID.</strong> Ever. Deleted means gone; the identifier is retired with it. Reuse turns a stale reference from a clean not-found into silent corruption pointing at the wrong record.</p>]]></content:encoded></item><item><title>The graph nobody drew</title><link>https://readme.news/the-graph-nobody-drew/</link><guid isPermaLink="true">https://readme.news/the-graph-nobody-drew/</guid><pubDate>Mon, 14 Sep 2026 09:00:00 +0000</pubDate><description>Your service dependency graph exists whether or not anyone has written it down. Writing it down takes an afternoon.</description><content:encoded><![CDATA[<p>Ask an engineering team to list what their service depends on to serve a request. You will get the database and two or three obvious APIs.</p>
<p>The real list is usually fifteen things, and the difference between the two lists is where the surprising outages come from.</p>
<h2 id="the-full-inventory">the full inventory<a class="anchor" href="#the-full-inventory" aria-label="link to this section">#</a></h2>
<p>For one service, one request path:</p>
<ul><li><strong>Datastores.</strong> Primary, replicas, cache, search index, object storage.</li><li><strong>Internal services</strong> it calls synchronously.</li><li><strong>Third-party APIs</strong> — payments, email, geocoding, <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">feature flags</a>, auth.</li><li><strong>Infrastructure control planes.</strong> Service discovery, DNS, the secret manager, the config service. These are the ones nobody lists and they are the ones that take everything down.</li><li><strong>Identity.</strong> The token issuer, the JWKS endpoint it fetches keys from.</li><li><strong>Observability.</strong> If your telemetry agent blocks when the collector is down — and some do — it is in your request path.</li><li><strong>The container registry</strong>, at start-up. Fine until you need to scale during an incident and cannot pull an image.</li><li><strong>Certificate infrastructure</strong>, including OCSP and the ACME endpoint.</li><li><strong>NTP.</strong> Rarely, spectacularly.</li></ul>
<p>The pattern: the dependencies people list are the ones they <em>call</em>. The ones that hurt are the ones they <em>need in order to start, authenticate, or scale.</em></p>
<h2 id="the-two-questions-per-dependency">the two questions per dependency<a class="anchor" href="#the-two-questions-per-dependency" aria-label="link to this section">#</a></h2>
<p>For each item, two answers, written down:</p>
<p><strong>What happens when it is slow?</strong> Slow is worse than down and much more common. Down fails fast; slow holds your connections, exhausts your pool, and turns one dependency's bad day into your outage.</p>
<p><strong>What happens when it is unavailable?</strong> Three possible answers, and only one is a decision:</p>
<ul><li><em>We fail.</em> Fine, if the dependency is genuinely essential — a database for a write path.</li><li><em>We degrade.</em> Serve stale cache, hide the feature, use a default. This is the answer for most non-essential dependencies and it requires code that usually does not exist yet.</li><li><em>We do not know.</em> This is the real answer for most dependencies, and it is the finding.</li></ul>
<h2 id="the-classification-that-matters">the classification that matters<a class="anchor" href="#the-classification-that-matters" aria-label="link to this section">#</a></h2>
<p>Sort every dependency into two buckets:</p>
<p><strong>Hard</strong> — the request genuinely cannot be served without it. Should be a very short list. Each one caps your maximum availability at its own.</p>
<p><strong>Soft</strong> — the request can be served in degraded form. Every soft dependency needs a timeout, a fallback, and a circuit breaker, or it is a hard dependency that nobody has admitted to.</p>
<p>The arithmetic is the reason this matters: five hard dependencies at 99.9% each gives you at best 99.5%, before any failure of your own. If your availability target is higher than that, some of those dependencies have to become soft, and that is an engineering project rather than a target you can declare.</p>
<h2 id="the-feature-flag-trap">the feature flag trap<a class="anchor" href="#the-feature-flag-trap" aria-label="link to this section">#</a></h2>
<p>Worth naming specifically because it catches good teams. Feature flag services are adopted as a safety mechanism — a way to turn things off when they go wrong.</p>
<p>Then the flag SDK becomes a hard dependency on the request path, and when the flag service has an outage your service does too. The tool you adopted to make incidents smaller has made one bigger.</p>
<p>The fix is standard for this whole category and rarely applied: cache flag values locally, evaluate from the cache, refresh in the background, and ship a default in the binary. Then the flag service being down means flags are stale, which is survivable.</p>
<h2 id="how-to-draw-it">how to draw it<a class="anchor" href="#how-to-draw-it" aria-label="link to this section">#</a></h2>
<p>Do not buy anything. Open a file:</p>
<div class="code"><pre><code>orders-api
  HARD  postgres-primary        write path; no fallback
  HARD  auth-jwks               cached 1h, so 1h of grace
  SOFT  redis                   cache; miss → primary, 3x latency
  SOFT  inventory-svc           timeout 300ms → show "check availability"
  SOFT  pricing-svc             timeout 200ms → list price, no promo
  SOFT  flags-svc               local cache + baked defaults
  BOOT  vault                   secrets at start; running pods unaffected
  BOOT  registry                image pull; blocks scale-up only</code></pre></div>
<p>Twenty minutes per service. The <code>BOOT</code> category — needed to start but not to serve — is the one that catches people out during incidents, because those dependencies are invisible right up until you need to replace a pod.</p>
<p>Then, for every <code>SOFT</code> line, check that the fallback described actually exists in code. That check is where the real findings are, and it is usually not the timeout that is missing but the fallback behind it.</p>]]></content:encoded></item><item><title>Timeouts: every one of them</title><link>https://readme.news/timeouts-every-one-of-them/</link><guid isPermaLink="true">https://readme.news/timeouts-every-one-of-them/</guid><pubDate>Thu, 10 Sep 2026 09:00:00 +0000</pubDate><description>A request crosses a dozen components with a timeout each, and almost nobody has ever added them up.</description><content:encoded><![CDATA[<p>Every layer of your stack has a timeout. Most of them are defaults. Almost nobody has written them down in one place and checked that they make sense together.</p>
<p>The result is systems where the client gives up at 10 seconds, the server keeps working for 60, and the database holds a lock for 300.</p>
<h2 id="the-inventory">the inventory<a class="anchor" href="#the-inventory" aria-label="link to this section">#</a></h2>
<p>For a single HTTP request, in rough order:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">layer</th><th style="text-align:left">typical default</th></tr></thead><tbody><tr><td style="text-align:left">browser / client library</td><td style="text-align:left">30s or none</td></tr><tr><td style="text-align:left">DNS resolution</td><td style="text-align:left">5s per attempt</td></tr><tr><td style="text-align:left">TCP connect</td><td style="text-align:left">20–75s (OS)</td></tr><tr><td style="text-align:left">TLS handshake</td><td style="text-align:left">inherits connect</td></tr><tr><td style="text-align:left">load balancer idle</td><td style="text-align:left">60s</td></tr><tr><td style="text-align:left">reverse proxy read</td><td style="text-align:left">60s</td></tr><tr><td style="text-align:left">application server request</td><td style="text-align:left">often none</td></tr><tr><td style="text-align:left">HTTP client to downstream</td><td style="text-align:left">often none</td></tr><tr><td style="text-align:left">connection pool acquire</td><td style="text-align:left">30s</td></tr><tr><td style="text-align:left">database statement</td><td style="text-align:left">often none</td></tr><tr><td style="text-align:left">database lock wait</td><td style="text-align:left">often none</td></tr></tbody></table></div>
<p>Two things stand out. <strong>"Often none" appears five times</strong> — most application-level timeouts are unset by default. And <strong>the values are not coordinated</strong> with each other in any way.</p>
<h2 id="the-rule-that-fixes-most-of-it">the rule that fixes most of it<a class="anchor" href="#the-rule-that-fixes-most-of-it" aria-label="link to this section">#</a></h2>
<p><strong>Timeouts must decrease as you go deeper.</strong></p>
<div class="code"><pre><code>client 10s
  └─ load balancer 9s
       └─ application 8s
            └─ downstream call 3s (with 1 retry → 6s worst case)
                 └─ connection acquire 1s
                      └─ database statement 2s</code></pre></div>
<p>Each layer must be shorter than its caller, with room for <a class="xref" href="/retries-a-complete-guide-to-not-making-it-worse/" title="Retries: a complete guide to not making it worse">retries</a>. If an inner layer can outlast its caller, the caller gives up while the work continues — which means you are burning capacity on results nobody will receive. Under load, that is most of your capacity.</p>
<h2 id="deadline-propagation">deadline propagation<a class="anchor" href="#deadline-propagation" aria-label="link to this section">#</a></h2>
<p>The better version of the rule: do not configure each timeout independently. Pass the deadline down.</p>
<div class="code"><span class="code-lang">go</span><pre><code class="lang-go">// caller has 10s; every downstream inherits what remains
ctx, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()

// 3s into the request, this call gets at most 7s, automatically
resp, err := downstream.Fetch(ctx, id)</code></pre></div>
<p>Go's <code>context</code>, gRPC's deadlines and equivalents elsewhere exist for this. Once you propagate deadlines, a slow first step automatically shortens the budget for later ones, and work never continues past the point where anyone is waiting.</p>
<p>This is the single highest-value change available in most service codebases, and it is usually a few days of threading a parameter through.</p>
<h2 id="the-ones-people-forget">the ones people forget<a class="anchor" href="#the-ones-people-forget" aria-label="link to this section">#</a></h2>
<p><strong>Lock wait timeouts.</strong> A transaction waiting on a row lock with no timeout waits forever. Set <code>lock_timeout</code> in Postgres, <code>innodb_lock_wait_timeout</code> in MySQL.</p>
<p><strong>DDL statements.</strong> A migration that cannot get its lock will queue behind a long transaction — and everything else <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> behind it. <code>SET lock_timeout = '3s'</code> before DDL is the line that prevents a large share of migration incidents.</p>
<p><strong>Idle-in-transaction.</strong> A connection that opened a transaction and went away holds locks and blocks vacuum indefinitely. <code>idle_in_transaction_session_timeout</code> is the safety net.</p>
<p><strong>Client-side connect vs. read.</strong> These are different, and most libraries let you set them separately. A short connect timeout with a long read timeout is usually what you want.</p>
<p><strong>DNS.</strong> Rarely configured, occasionally the whole problem.</p>
<h2 id="write-them-down">write them down<a class="anchor" href="#write-them-down" aria-label="link to this section">#</a></h2>
<p>One table, in the repository, listing every timeout in the request path with its current value and where it is configured.</p>
<p>The exercise takes an afternoon and reliably finds at least one inversion — an inner layer waiting longer than the outer one — and at least two places where the value is a framework default nobody chose.</p>
<p>That document then becomes something you can review when latency changes, rather than a set of numbers scattered across six config files and three languages.</p>
<h2 id="the-failure-mode-this-prevents">the failure mode this prevents<a class="anchor" href="#the-failure-mode-this-prevents" aria-label="link to this section">#</a></h2>
<p>Without coordinated timeouts, a slow dependency does not degrade your service — it exhausts it. Requests pile up waiting on something their callers abandoned long ago, connections stay held, the pool empties, and healthy requests start failing for want of a connection.</p>
<p>That is a full outage caused by one slow downstream, and the difference between it and a brief latency blip is entirely whether the numbers were coordinated.</p>]]></content:encoded></item><item><title>The queues you did not know you had</title><link>https://readme.news/the-queues-you-did-not-know-you-had/</link><guid isPermaLink="true">https://readme.news/the-queues-you-did-not-know-you-had/</guid><pubDate>Tue, 01 Sep 2026 09:00:00 +0000</pubDate><description>Every fixed-size resource is a queue. Most of them are unmonitored, and that is where latency hides.</description><content:encoded><![CDATA[<p>You know about the message queue, because you chose it and it has a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>. The queues that hurt are the ones nobody named.</p>
<p>Anywhere a fixed-size resource is shared by more requests than it has capacity, there is a queue. It has a depth, a wait time, and a failure mode, and almost none of them are instrumented.</p>
<h2 id="the-inventory">the inventory<a class="anchor" href="#the-inventory" aria-label="link to this section">#</a></h2>
<p><strong>The connection pool.</strong> Twenty connections, forty concurrent requests: twenty requests are waiting. Pool wait time is the single most under-measured latency component in web applications, and it is invisible in a database query timer because the clock starts after the connection is acquired.</p>
<p><strong>The thread pool or worker pool.</strong> Same shape. Requests queue for a worker, and the time spent waiting is not attributed to any handler.</p>
<p><strong>The TCP accept backlog.</strong> The kernel holds connections your process has not accepted yet. Overflow silently drops them, and the client sees a timeout with no server-side trace at all.</p>
<p><strong>The HTTP client's per-host connection limit.</strong> Most clients cap concurrent connections per destination. Exceed it and your requests queue in the client, before any network activity, invisible to server-side metrics on both ends.</p>
<p><strong>The DNS resolver.</strong> A limited number of in-flight lookups with a cache that can stampede on expiry.</p>
<p><strong>The disk queue.</strong> Storage devices have a queue depth. Exceed it and I/O waits.</p>
<p><strong>The garbage collector.</strong> Not a queue exactly, but the same behaviour: work that accumulates and is paid in a burst.</p>
<p><strong>Rate limiters.</strong> A limiter that delays rather than rejecting is a queue with an enforced service rate.</p>
<h2 id="why-this-matters-more-than-it-sounds">why this matters more than it sounds<a class="anchor" href="#why-this-matters-more-than-it-sounds" aria-label="link to this section">#</a></h2>
<p>Little's Law: <code>L = λW</code>. Items in the system equals arrival rate times time in system. Rearranged, <strong>wait time is queue length over service rate</strong>.</p>
<p>That means your queue depth choice <em>is</em> a latency choice, whether or not you made it deliberately. A connection pool with a 30-second acquisition timeout is a promise that some requests will wait 30 seconds.</p>
<p>And these queues compose. A request waits for a worker, then waits for a connection, then waits for the disk. Each is modest; the sum is the p99 nobody can explain, because each layer's own metrics look fine.</p>
<h2 id="how-to-find-them">how to find them<a class="anchor" href="#how-to-find-them" aria-label="link to this section">#</a></h2>
<p><strong>Measure acquisition, not just use.</strong> Time the wait for the resource separately from the work done with it:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">t0 = time.perf_counter()
with pool.acquire() as conn:
    t1 = time.perf_counter()
    result = conn.execute(query)
span.set_attribute("db.pool_wait_ms", (t1 - t0) * 1000)
span.set_attribute("db.query_ms", (time.perf_counter() - t1) * 1000)</code></pre></div>
<p>Two numbers instead of one. The first one is the one you did not have, and it is frequently the larger.</p>
<p><strong>Add up your span times.</strong> If the parent span is 400 ms and the children total 180 ms, the missing 220 ms is queueing somewhere. That gap is the most useful signal in a trace and almost nobody looks for it.</p>
<p><strong>Check <code>netstat -s</code> for accept queue overflows.</strong> A non-zero and growing "listen queue overflowed" counter means you are dropping connections before your application sees them.</p>
<p><strong>Load test past capacity deliberately.</strong> Push to 3× and watch which metric degrades first. That is your binding queue, and you will not find it at normal load.</p>
<h2 id="the-rule">the rule<a class="anchor" href="#the-rule" aria-label="link to this section">#</a></h2>
<p>Every queue needs three things, and most have none:</p>
<ol><li><strong>A bounded size.</strong> Unbounded is a memory leak with a friendly name.</li><li><strong>A wait metric</strong> — specifically the <em>age of the oldest waiter</em>, which tells you how far behind you are in time rather than in count.</li><li><strong>A decision about overflow.</strong> Reject, drop, or block. Pick one deliberately, because the default is usually "queue forever," and queueing forever means serving requests whose callers gave up ten seconds ago.</li></ol>
<p>Do that for the queues you chose. Then go find the six you did not.</p>]]></content:encoded></item><item><title>The API you ship to yourself</title><link>https://readme.news/the-api-you-ship-to-yourself/</link><guid isPermaLink="true">https://readme.news/the-api-you-ship-to-yourself/</guid><pubDate>Tue, 25 Aug 2026 09:00:00 +0000</pubDate><description>Internal interfaces get none of the care external ones do, and they are the ones you&#x27;ll live with longest.</description><content:encoded><![CDATA[<p>Teams that would never ship a public API without review, versioning and documentation routinely create internal interfaces in a pull request nobody looked at closely, with a name chosen in thirty seconds.</p>
<p>Those internal interfaces then last longer than the public ones, because nothing external forces them to change and nothing internal makes changing them worth the trouble.</p>
<h2 id="why-they-matter-more-than-they-look">why they matter more than they look<a class="anchor" href="#why-they-matter-more-than-they-look" aria-label="link to this section">#</a></h2>
<p><strong>They are the seams your architecture is made of.</strong> A module's public functions <em>are</em> its contract, whether or not you called it a contract.</p>
<p><strong>Everyone reads them.</strong> A public API is read by customers. An internal one is read by every engineer who joins, forever.</p>
<p><strong>They shape what is easy.</strong> An <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> that makes the right thing awkward guarantees people do the wrong thing, and then you have a pattern.</p>
<p><strong>They are hard to change once used.</strong> Not because of compatibility, but because finding and updating every caller is work nobody has budgeted.</p>
<h2 id="what-to-actually-apply">what to actually apply<a class="anchor" href="#what-to-actually-apply" aria-label="link to this section">#</a></h2>
<p>Not the full public-API apparatus. Four things:</p>
<p><strong>1. Name it from the caller's side.</strong></p>
<p>The best test is to write the call before the implementation:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python"># would a caller guess this?
orders.pending_for(customer_id, since=last_sync)

# or this?
OrderService.get_orders_by_customer_with_status_filter(customer_id, "PENDING", last_sync)</code></pre></div>
<p>The first reads like the sentence the caller was thinking. The second reads like the implementation leaked into the name.</p>
<p><strong>2. Make the wrong call hard to write.</strong></p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python"># a caller can get this wrong and never know
def transfer(from_account, to_account, amount): ...
transfer(b, a, 500)   # silently backwards

# a caller cannot
def transfer(*, source: AccountId, dest: AccountId, cents: int): ...</code></pre></div>
<p>Keyword-only arguments, distinct types for things that must not be swapped, and units in the name. Every one of these converts a runtime bug into a compile or call-site error.</p>
<p><strong>3. Return something that cannot be misread.</strong></p>
<p><code>None</code> for "not found" and <code>None</code> for "error" and <code>None</code> for "empty" is three different meanings on one value. A <code>Result</code>, an explicit exception, or distinct return types cost a few lines and remove a category of caller bug.</p>
<p><strong>4. Document the <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a>, not the happy path.</strong></p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def reserve(sku: str, qty: int) -&gt; Reservation:
    """Reserve stock, or raise.

    Raises:
        OutOfStock: qty exceeds available. Callers should offer backorder.
        SkuUnknown: not a real SKU — this is a bug, not a user error.
        LockTimeout: transient, safe to retry with backoff.
    """</code></pre></div>
<p>The signature already said what it takes and returns. What it cannot say is which errors are retryable and which are programmer error — and that is exactly what a caller needs.</p>
<h2 id="the-deprecation-problem">the deprecation problem<a class="anchor" href="#the-deprecation-problem" aria-label="link to this section">#</a></h2>
<p>Internal APIs never get deprecated because there is always something more urgent, and the old one keeps working.</p>
<p>The mechanism that helps is embarrassingly simple: <strong>make the old path noisy in development.</strong></p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def get_orders_by_customer(*args, **kwargs):
    warnings.warn(
        "get_orders_by_customer is replaced by orders.pending_for(); "
        "see ADR-14. Removal targeted 2026-12.",
        DeprecationWarning, stacklevel=2,
    )
    return orders.pending_for(*args, **kwargs)</code></pre></div>
<p>Warnings in test output that name the replacement and a date get acted on. Silent compatibility shims live forever.</p>
<h2 id="the-boundary-that-actually-needs-versioning">the boundary that actually needs versioning<a class="anchor" href="#the-boundary-that-actually-needs-versioning" aria-label="link to this section">#</a></h2>
<p>Most internal interfaces do not need versions — you can find every caller and change them together. That is the entire advantage of being internal and you should use it rather than building machinery to avoid it.</p>
<p>The exception is any interface crossing a <strong>deploy boundary</strong>: service to service, or anything consumed by a client you do not deploy simultaneously. There, old and new run at the same time by definition, and you need the same expand-contract discipline as a schema change — add the new field, support both, migrate callers, remove the old.</p>
<p>The mistake is applying that machinery to an in-process module boundary where a single commit could have updated everything.</p>
<h2 id="the-one-line-version">the one-line version<a class="anchor" href="#the-one-line-version" aria-label="link to this section">#</a></h2>
<p>Design internal interfaces from the call site, make the wrong call unwriteable, document the errors, and change them freely while you still can — because the window where you can find every caller closes quietly, and nobody notices until they need it open.</p>]]></content:encoded></item><item><title>Boring technology, revisited</title><link>https://readme.news/boring-technology-revisited/</link><guid isPermaLink="true">https://readme.news/boring-technology-revisited/</guid><pubDate>Sat, 15 Aug 2026 09:00:00 +0000</pubDate><description>The innovation-token argument is a decade old and still right, with one amendment nobody makes.</description><content:encoded><![CDATA[<p>The "choose boring technology" argument goes like this: you have a small number of innovation tokens. Spend them on the thing that is genuinely your problem. Everything else should be the option so well-understood that its failure modes are documented by strangers.</p>
<p>It has held up for a decade. It is worth restating because a decade is long enough that people have started applying it wrongly, in a specific and predictable way.</p>
<h2 id="why-it-works">why it works<a class="anchor" href="#why-it-works" aria-label="link to this section">#</a></h2>
<p><strong>Known failure modes.</strong> The value of a boring technology is not that it is good. It is that when it breaks at 3 a.m., the <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error message</a> has been pasted into a public forum by someone who then explained the fix.</p>
<p><strong>A hiring pool that already knows it.</strong> Every unusual choice is a thing you must teach every new engineer, forever.</p>
<p><strong>Operational surface you can reason about.</strong> Boring things have runbooks, monitoring integrations, backup tooling, and a decade of accumulated operational wisdom.</p>
<p><strong>Someone else finds the bugs.</strong> A widely-deployed system has been run at scales you will never reach, by people who reported what broke.</p>
<h2 id="the-amendment">the amendment<a class="anchor" href="#the-amendment" aria-label="link to this section">#</a></h2>
<p>Here is what the original argument does not say, and what a decade has taught: <strong>boring is a property of a technology at a moment in time, and it moves in both directions.</strong></p>
<p>Things become boring. Things also stop being boring — a project loses its maintainers, its community fragments, its corporate sponsor reorganises, the ecosystem moves and it does not.</p>
<p>So "choose boring" is not a decision you make once. It is a property you have to re-check, and the failure mode nobody plans for is the technology that was the safe choice in 2019 and is now a liability that everyone forgot to reconsider.</p>
<p><strong>The practical version:</strong> once a year, for each load-bearing dependency, ask — is this still boring? Is it still maintained, still hired for, still the thing a new engineer would expect? A "yes" costs a minute. A "no" is the most valuable thing you will learn that quarter.</p>
<h2 id="how-to-tell-whether-something-is-boring-now">how to tell whether something is boring <em>now</em><a class="anchor" href="#how-to-tell-whether-something-is-boring-now" aria-label="link to this section">#</a></h2>
<p>Not by age. Some old things are unmaintained; some five-year-old things are completely settled.</p>
<p>The signals that actually matter:</p>
<ul><li><strong>Multiple independent maintainers</strong>, ideally from more than one employer.</li><li><strong>A release in the last six months</strong>, and a <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a> that shows maintenance rather than churn.</li><li><strong>You can hire for it.</strong> Search jobs, not GitHub stars.</li><li><strong>Operational documentation written by users</strong>, not just by the vendor.</li><li><strong>A migration path off it</strong>, documented by people who took it. A technology nobody has successfully left is a technology you cannot leave.</li><li><strong>Boring failure modes.</strong> Can you name what happens when it runs out of memory, loses its network, or gets a corrupt config? If nobody has written that down, it is not boring yet.</li></ul>
<h2 id="where-to-spend-the-tokens">where to spend the tokens<a class="anchor" href="#where-to-spend-the-tokens" aria-label="link to this section">#</a></h2>
<p>The argument's real content is that innovation tokens are scarce, and the scarcity is not about technology risk. It is about <strong>attention</strong>.</p>
<p>Every unusual choice consumes attention: to learn it, to operate it, to debug it, to explain it. Attention is the actual constraint on an engineering team, and it is far more limited than the budget.</p>
<p>So spend the tokens where the unusual choice is the product. If your differentiator is a query engine, be adventurous about storage and utterly boring about everything else. If your differentiator is a workflow, use the most conventional stack in existence and put all the attention into the workflow.</p>
<p>The teams that get this wrong are almost never wrong about one big choice. They are wrong about six small ones, each individually defensible, that together consumed all the attention that should have gone into the thing customers pay for.</p>
<h2 id="the-counterweight">the counterweight<a class="anchor" href="#the-counterweight" aria-label="link to this section">#</a></h2>
<p>Taken too far this becomes an argument for never learning anything, and that is its own failure. A team that made every choice in 2016 and never revisited one is not disciplined, it is stuck — and the technology it chose has probably stopped being boring in the meantime.</p>
<p>The synthesis: be adventurous in one place at a time, deliberately, with a written reason. Be boring everywhere else. Re-check annually.</p>
<p>That is a harder discipline than either "always use the new thing" or "never use the new thing," and it is the only one that survives a decade.</p>]]></content:encoded></item><item><title>The service you should not have written</title><link>https://readme.news/the-service-you-should-not-have-written/</link><guid isPermaLink="true">https://readme.news/the-service-you-should-not-have-written/</guid><pubDate>Thu, 13 Aug 2026 09:00:00 +0000</pubDate><description>A checklist for the moment before you create a new deployable, when saying no is still free.</description><content:encoded><![CDATA[<p>Creating a new service feels like progress. It has a clean repository, no legacy, a fresh CI pipeline, and none of the compromises of the thing it is splitting away from.</p>
<p>It is also a permanent commitment that somebody will still be paying for in eight years. Here is the conversation worth having first.</p>
<h2 id="what-a-service-costs-in-full">what a service costs, in full<a class="anchor" href="#what-a-service-costs-in-full" aria-label="link to this section">#</a></h2>
<p>Not the code. The code is the cheap part.</p>
<ul><li>A repository, a build, a deploy pipeline, a rollback path.</li><li>A place to run, sized, with capacity headroom.</li><li>Monitoring, alerting, dashboards, an <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> runbook.</li><li>Secrets, credentials, and their rotation.</li><li>Dependency upgrades and security patching, forever.</li><li>A place in the request path that can now time out, and the circuit breaker and retry policy that implies.</li><li>Documentation, and an owner who is still at the company.</li><li>One more thing a new engineer must learn about.</li></ul>
<p>That list is the same whether the service is four hundred lines or forty thousand. The overhead is fixed, which is why small services are the ones whose economics are worst.</p>
<h2 id="the-questions-in-order">the questions, in order<a class="anchor" href="#the-questions-in-order" aria-label="link to this section">#</a></h2>
<p><strong>1. Could this be a module?</strong></p>
<p>The default answer for new functionality is a well-bounded module in something that already exists. It gets all of the above for free.</p>
<p>The follow-up that matters: <em>if this were a module with an enforced boundary, what would we lose?</em> If the honest answer is "nothing, it would just feel less tidy," write the module.</p>
<p><strong>2. What is the deployment argument?</strong></p>
<p>The strongest reason to extract is that a team needs to release on its own cadence and currently cannot. That is real and it scales with organisation size.</p>
<p>But check it: are they actually blocked, or do they merely deploy together? Two teams that could deploy independently and choose not to have a process problem, not an architecture problem, and a new service will not fix it.</p>
<p><strong>3. What is the scaling argument, with numbers?</strong></p>
<p>"It might need to scale differently" is not an argument. "This component is CPU-bound and spiky while the rest is I/O-bound and steady, and we are currently provisioning for the peak of both" is.</p>
<p>If you cannot state the resource profile that differs, there is no scaling argument.</p>
<p><strong>4. Where does the data live?</strong></p>
<p>The question that sinks most extractions. If the new service needs the same tables the old one does, you have not split anything — you have created two writers to one database, which is worse than one writer, and you now have to coordinate migrations across two deploy cycles.</p>
<p>A service that does not own its data is a distributed monolith with extra latency.</p>
<p><strong>5. What happens when it is down?</strong></p>
<p>If the answer is "the main flow breaks," you have added a failure mode and bought nothing in return. Fault isolation is not free; it is the result of deliberate <a class="xref" href="/timeouts-every-one-of-them/" title="Timeouts: every one of them">timeouts</a>, fallbacks and degradation, and if you are going to build those anyway you could have built them around a module.</p>
<p><strong>6. Who owns it in two years?</strong></p>
<p>Name the team. Not the person — the team. Services outlive their authors, and an ownerless service is the one that runs three major versions behind until it is a security incident.</p>
<h2 id="the-cases-where-the-answer-is-yes">the cases where the answer is yes<a class="anchor" href="#the-cases-where-the-answer-is-yes" aria-label="link to this section">#</a></h2>
<p>Being fair, because the reflex against splitting can be as unexamined as the reflex toward it.</p>
<ul><li>A genuinely different resource profile, with numbers.</li><li>A compliance or data-residency boundary that must be enforced structurally.</li><li>A component that has to be written in a different language for a real reason.</li><li>A team that is demonstrably blocked on someone else's release cadence.</li><li>Something with wildly different availability requirements — a batch job that can be down for an hour next to a checkout path that cannot.</li></ul>
<p>In every one of those, the service is buying something specific that a module cannot.</p>
<h2 id="the-reversibility-test">the reversibility test<a class="anchor" href="#the-reversibility-test" aria-label="link to this section">#</a></h2>
<p>Before you create it, ask: <strong>if this turns out to be wrong, how do we merge it back?</strong></p>
<p>If the answer is "we would not, we would just live with it," you are making a one-way decision on a hunch. Those deserve more scrutiny than they usually get, and the moment before the repository exists is the last time the scrutiny is free.</p>]]></content:encoded></item><item><title>Vendor lock-in: an honest cost model</title><link>https://readme.news/vendor-lock-in-an-honest-cost-model/</link><guid isPermaLink="true">https://readme.news/vendor-lock-in-an-honest-cost-model/</guid><pubDate>Fri, 17 Jul 2026 09:00:00 +0000</pubDate><description>Portability is not free and neither is dependence. A framework for deciding how much abstraction to buy.</description><content:encoded><![CDATA[<p>"Avoid vendor lock-in" is treated as self-evidently good advice. It is not advice, it is a preference, and following it uncritically produces systems that are worse in exchange for optionality nobody will ever exercise.</p>
<p>Here is a way to actually decide.</p>
<h2 id="the-cost-of-lock-in">the cost of lock-in<a class="anchor" href="#the-cost-of-lock-in" aria-label="link to this section">#</a></h2>
<p><strong>Switching cost.</strong> How much engineering time to move off, if you had to.</p>
<p><strong>Pricing power.</strong> A vendor who knows you cannot leave prices accordingly. This is real and it is usually the largest ongoing cost.</p>
<p><strong>Capability ceiling.</strong> You are limited to what they support, on their timeline.</p>
<p><strong>Correlated risk.</strong> They have an outage, you have an outage. They change their terms, you comply. They get acquired and sunset the product, you migrate on their schedule.</p>
<h2 id="the-cost-of-avoiding-lock-in">the cost of avoiding lock-in<a class="anchor" href="#the-cost-of-avoiding-lock-in" aria-label="link to this section">#</a></h2>
<p>This is the half that gets ignored, and it is frequently larger.</p>
<p><strong>The abstraction layer itself.</strong> Code to write, maintain, test, and debug through. It is a permanent tax and it makes stack traces longer.</p>
<p><strong>Lowest common denominator.</strong> Your abstraction can only expose what all candidate providers support. You give up the features that made the good option good.</p>
<p><strong>The abstraction is usually wrong anyway.</strong> It was designed against one provider's model. When you actually try to swap, you discover the abstraction encoded assumptions that do not hold, and you rewrite it.</p>
<p><strong>Delayed value.</strong> Time spent on portability is time not spent on the product.</p>
<h2 id="the-framework">the framework<a class="anchor" href="#the-framework" aria-label="link to this section">#</a></h2>
<p>For each dependency, estimate:</p>
<ol><li><strong>Switching cost</strong> — engineer-weeks to move.</li><li><strong>Probability you switch</strong> in the next three years.</li><li><strong>Cost of the abstraction</strong> — engineer-weeks now, plus ongoing drag.</li></ol>
<p>Then: if <code>switching_cost × probability &lt; abstraction_cost</code>, do not abstract.</p>
<p>The numbers are rough. The exercise still clarifies, because it forces you to state the probability out loud, and stated probabilities are usually much lower than the implied ones people are acting on.</p>
<h2 id="the-categories-worked-through">the categories, worked through<a class="anchor" href="#the-categories-worked-through" aria-label="link to this section">#</a></h2>
<p><strong>Object storage.</strong> Switching cost: low. The S3 API is a de facto standard and every provider implements it. Probability: moderate — people do move for pricing.</p>
<p><strong>Verdict: use the S3 API, do not abstract further.</strong> The API is already the abstraction.</p>
<p><strong>Compute.</strong> Switching cost: moderate if containerized, high if you use provider-specific serverless. Probability: low.</p>
<p><strong>Verdict: containerize</strong> — which is good practice anyway — <strong>and use whatever managed service you want.</strong> Do not build a compute abstraction layer.</p>
<p><strong>Relational database.</strong> Switching cost: high. Probability: low.</p>
<p><strong>Verdict: use the database's features.</strong> Teams that avoid stored procedures, database-specific types, and advanced indexing to stay portable are giving up real capability for an event that will not happen. Postgres-specific SQL is fine. You are not going to migrate to a different engine, and if you do, the SQL dialect will be the smallest part of the pain.</p>
<p><strong>Managed <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a>, streams, and similar.</strong> Switching cost: moderate. The semantics differ enough between providers that a thin abstraction genuinely helps.</p>
<p><strong>Verdict: a thin <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> — publish, subscribe, ack — is worth it.</strong> Not a full abstraction; a boundary.</p>
<p><strong>Authentication.</strong> Switching cost: very high — you have to migrate user credentials, sessions, and integrations. Probability: low, but the consequences of being stuck are severe.</p>
<p><strong>Verdict: use standard protocols.</strong> OIDC and SAML are the abstraction. A provider that supports them is replaceable in principle; one with a proprietary SDK is not.</p>
<p><strong>AI model providers.</strong> Switching cost: low if you kept the interface thin. Probability: <strong>high</strong> — the market is moving fast and the right choice changes quarterly.</p>
<p><strong>Verdict: definitely abstract.</strong> This is the clearest case on the list. A thin interface plus an eval harness in your own repository makes model changes an afternoon. Teams that did this in 2024 have been switching providers casually ever since.</p>
<p><strong>Observability.</strong> Switching cost: moderate to high — instrumentation is everywhere in your code. Probability: moderate, usually driven by cost.</p>
<p><strong>Verdict: OpenTelemetry.</strong> The instrumentation is vendor-neutral, the collector handles routing, and swapping backends is a configuration change.</p>
<h2 id="the-pattern">the pattern<a class="anchor" href="#the-pattern" aria-label="link to this section">#</a></h2>
<p><strong>Abstract where the switching probability is high and the abstraction is cheap.</strong> AI providers, observability backends, object storage.</p>
<p><strong>Do not abstract where switching is unlikely and the abstraction costs you real capability.</strong> Databases, compute platforms, managed services you chose for their specific features.</p>
<p><strong>Use standard protocols wherever they exist.</strong> OIDC, S3, OpenTelemetry, SQL. A standard is an abstraction someone else maintains, and it is always cheaper than yours.</p>
<h2 id="the-thing-that-actually-protects-you">the thing that actually protects you<a class="anchor" href="#the-thing-that-actually-protects-you" aria-label="link to this section">#</a></h2>
<p>Not an abstraction layer. <strong>Your data, in a format you can export, and a documented process for leaving.</strong></p>
<p>Ask, before adopting anything: can I get all my data out, in a usable format, without their cooperation? If yes, you have real optionality regardless of how coupled your code is, because the expensive part of a migration is never the code — it is the data.</p>
<p>If the answer is no, that is a much bigger red flag than any API coupling, and it is the question almost nobody asks during procurement.</p>]]></content:encoded></item><item><title>Connection pooling, explained properly</title><link>https://readme.news/connection-pooling-explained-properly/</link><guid isPermaLink="true">https://readme.news/connection-pooling-explained-properly/</guid><pubDate>Wed, 15 Jul 2026 09:00:00 +0000</pubDate><description>Why your database has 400 connections, why that is bad, and how to size a pool without guessing.</description><content:encoded><![CDATA[<p>Connection pool sizing is done by copying a number from a blog post, and the number is usually wrong in a specific and expensive way.</p>
<h2 id="why-connections-are-expensive">why connections are expensive<a class="anchor" href="#why-connections-are-expensive" aria-label="link to this section">#</a></h2>
<p>In Postgres specifically, each connection is a separate operating system process with its own memory. A few megabytes of baseline, plus work memory for sorting and hashing, plus its share of shared buffer access.</p>
<p>Four hundred connections means four hundred processes. The scheduler is context switching between them, they are contending for the same locks and buffers, and the memory is largely wasted because most of them are idle.</p>
<p>The counterintuitive result, which has been measured many times: <strong>throughput frequently goes down as connection count goes up, past a fairly low threshold.</strong></p>
<p>More connections does not mean more concurrency. It means more contention.</p>
<h2 id="the-actual-number">the actual number<a class="anchor" href="#the-actual-number" aria-label="link to this section">#</a></h2>
<p>A widely used starting formula:</p>
<div class="code"><pre><code>connections = (core_count × 2) + effective_spindle_count</code></pre></div>
<p>For an 8-core machine with SSD storage, that is somewhere around 16 to 20.</p>
<p>That number seems shockingly low to people running pools of 100 or more. It is correct, and the reasoning is straightforward: a query is either using CPU or waiting on I/O. You need enough connections to keep the cores busy and to have some work queued behind I/O waits. Past that, additional connections are queued at the database instead of queued in your pool, and queueing at the database is worse because it consumes resources.</p>
<p><strong>Test it.</strong> Take your load test, run it at pool sizes of 10, 20, 40, 80, and 160, and plot throughput and p99 latency. The curve rises, flattens, and then degrades. Most people are on the degrading side and have never looked.</p>
<h2 id="the-pooler-layers">the pooler layers<a class="anchor" href="#the-pooler-layers" aria-label="link to this section">#</a></h2>
<p>Three, and they do different things.</p>
<p><strong>Application-side pool.</strong> In-process, reuses connections across requests. Every ORM and database driver has one. This is the minimum.</p>
<p><strong>External pooler</strong> — PgBouncer, pgcat, or a cloud provider's equivalent. Sits between your application and the database and multiplexes many client connections onto few server connections.</p>
<p>This is what you need when you have many application instances. Twenty instances with a pool of 20 each is 400 connections to the database, even if each instance is mostly idle. A pooler collapses that to the number the database actually wants.</p>
<p><strong>The database's own limit.</strong> <code>max_connections</code>. Set it lower than you think — it is a safety valve, and setting it high does not make the database faster, it makes the failure mode worse.</p>
<h2 id="pooling-modes-and-the-one-that-bites">pooling modes, and the one that bites<a class="anchor" href="#pooling-modes-and-the-one-that-bites" aria-label="link to this section">#</a></h2>
<p>External poolers have modes and choosing wrong causes subtle correctness bugs.</p>
<p><strong>Session pooling.</strong> A client gets a server connection for the duration of its session. Safe, and provides little multiplexing benefit.</p>
<p><strong>Transaction pooling.</strong> A server connection is assigned per transaction and returned after commit. This is where the big multiplexing win is, and it is what most people want.</p>
<p><strong>The catch:</strong> anything that depends on session state breaks.</p>
<ul><li>Prepared statements (unless the pooler supports them explicitly, which newer ones do)</li><li><code>SET</code> at the session level</li><li>Session-level advisory locks</li><li><code>LISTEN</code>/<code>NOTIFY</code></li><li>Temporary tables</li><li><code>WITH HOLD</code> cursors</li></ul>
<p>If your ORM uses server-side prepared statements by default — many do — you must either disable them or use a pooler that handles them. This is the single most common transaction-pooling problem and it manifests as intermittent errors under load, which is a miserable thing to debug.</p>
<p><strong>Statement pooling.</strong> A connection per statement. Maximum multiplexing, breaks multi-statement transactions. Almost never what you want.</p>
<h2 id="the-settings-that-matter">the settings that matter<a class="anchor" href="#the-settings-that-matter" aria-label="link to this section">#</a></h2>
<p>Beyond size:</p>
<p><strong>Connection timeout.</strong> How long a request waits for a connection before failing. This should be short — a few seconds. A request waiting thirty seconds for a connection has already been abandoned by its caller.</p>
<p><strong>Idle timeout.</strong> Return connections to the database when not in use. Important with many application instances.</p>
<p><strong>Max lifetime.</strong> Recycle connections periodically. This handles the case where a database failover happens and your pool is holding connections to the old primary — without a max lifetime, some pools hold those indefinitely.</p>
<p><strong>Validation.</strong> Test a connection before handing it out, or handle the failure on first use. Something has to deal with connections that died while idle.</p>
<h2 id="the-diagnostic">the diagnostic<a class="anchor" href="#the-diagnostic" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">sql</span><pre><code class="lang-sql">SELECT state, count(*), max(now() - state_change) AS longest
FROM pg_stat_activity
WHERE backend_type = 'client backend'
GROUP BY state;</code></pre></div>
<p>What you are looking for:</p>
<ul><li><strong>Many <code>idle in transaction</code></strong> — the worst state. A transaction is open and doing nothing, holding locks and preventing vacuum. This is an application bug: a transaction opened and not committed, usually because of an early return or an exception path.</li><li><strong>Many <code>idle</code></strong> — pool is oversized. Harmless but wasteful.</li><li><strong>Many <code>active</code> with long durations</strong> — queries are slow; the pool is not the problem.</li></ul>
<p><code>idle in transaction</code> is the one to alert on. It is always a bug and it causes outages that look like database problems and are application problems.</p>
<h2 id="the-summary">the summary<a class="anchor" href="#the-summary" aria-label="link to this section">#</a></h2>
<p>Size your pool from a load test, not from a blog post. It is smaller than you think. Use an external pooler in transaction mode if you have many application instances, and check your prepared statement behavior when you do.</p>]]></content:encoded></item><item><title>Backpressure is the concept your system is missing</title><link>https://readme.news/backpressure-is-the-concept-your-system-is-missing/</link><guid isPermaLink="true">https://readme.news/backpressure-is-the-concept-your-system-is-missing/</guid><pubDate>Fri, 03 Jul 2026 09:00:00 +0000</pubDate><description>When a fast producer meets a slow consumer, something has to give. Deciding what, in advance, is the whole discipline.</description><content:encoded><![CDATA[<p>Every system with a producer and a consumer eventually has a moment where the producer is faster. What happens next is either a design decision you made or an emergent behavior you discover during an incident.</p>
<h2 id="the-four-options">the four options<a class="anchor" href="#the-four-options" aria-label="link to this section">#</a></h2>
<p>There are only four. Every system picks one, explicitly or by accident.</p>
<p><strong>1. Buffer.</strong> Queue the excess.</p>
<p>Works for bursts. Fails for sustained overload, because a buffer is a delay, and an unbounded buffer is a memory leak with a friendly name. The failure mode is that memory grows until the process dies, taking the buffer with it — so you lose everything, at the worst possible moment.</p>
<p><strong>2. Drop.</strong> Discard the excess.</p>
<p>Correct more often than people are comfortable with. Metrics, logs, telemetry, non-critical events — dropping 5% of samples under load is fine and dying is not.</p>
<p>The requirement: <strong>know that you dropped, and how much.</strong> Silent drops are how you get a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> that looks healthy while data is missing.</p>
<p><strong>3. Block.</strong> Make the producer wait.</p>
<p>This is real backpressure. The consumer's slowness propagates upstream, the producer slows down, and the system reaches equilibrium at the consumer's rate.</p>
<p>Correct for internal pipelines where the producer can wait. Dangerous when the producer is a user-facing request handler, because now user requests are blocked on a background process.</p>
<p><strong>4. Reject.</strong> Tell the producer no.</p>
<p>The right answer at a service boundary. A 429 or 503 returned in one millisecond is much better than a request that waits thirty seconds and then times out, because the caller can make a decision — retry later, degrade, or tell the user.</p>
<h2 id="the-wrong-default">the wrong default<a class="anchor" href="#the-wrong-default" aria-label="link to this section">#</a></h2>
<p>Most systems buffer by default, unboundedly, without anyone deciding.</p>
<ul><li>An in-memory list that grows.</li><li>A queue with no maximum length.</li><li>A connection pool that <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> waiters forever.</li><li>A channel with a very large capacity, which is unbounded in practice.</li></ul>
<p>The failure mode is always the same: latency grows, memory grows, and then the process dies. And the requests you were holding were abandoned by their callers ten seconds earlier, so all that work was for nothing.</p>
<p><strong>A queue that is always full is not a buffer. It is a delay you cannot see.</strong></p>
<h2 id="the-practical-rules">the practical rules<a class="anchor" href="#the-practical-rules" aria-label="link to this section">#</a></h2>
<p><strong>Every queue has a maximum size.</strong> Every one. Choose the number by asking: how long should a request be willing to wait? Multiply by the consumer's rate. That is your queue depth.</p>
<p>If you cannot answer that question, the queue is not designed.</p>
<p><strong>Prefer rejecting to queueing at the edge.</strong> When a request arrives and the system is saturated, reject fast. The caller has a timeout; use it as your budget.</p>
<p><strong>Propagate deadlines.</strong> If the caller has 200 ms left, every downstream operation should know that. Work performed after the caller has given up is pure waste, and under overload it is the majority of the work being done.</p>
<div class="code"><span class="code-lang">go</span><pre><code class="lang-go">ctx, cancel := context.WithTimeout(ctx, remaining)
defer cancel()</code></pre></div>
<p><strong>Shed by priority, not randomly.</strong> Under load, serve health checks, serve authenticated users, serve the critical path. Shed background work, analytics, and prefetches. Random shedding means your health checks fail and your orchestrator kills healthy instances, which is a self-inflicted outage.</p>
<p>**Measure queue <em>age</em>, not depth.** Depth tells you how many. Age tells you how far behind. "The oldest item has been waiting four minutes" is actionable in a way that "there are 30,000 items" is not.</p>
<h2 id="the-pattern-that-ties-it-together">the pattern that ties it together<a class="anchor" href="#the-pattern-that-ties-it-together" aria-label="link to this section">#</a></h2>
<p>Little's Law: <code>L = λW</code>. Items in the system equals arrival rate times time in system.</p>
<p>Rearranged: <strong>wait time equals queue length divided by service rate.</strong></p>
<p>If your queue holds 10,000 items and you process 100 per second, the newest item waits 100 seconds. That is not a hypothetical — it is arithmetic, and it means your queue length choice <em>is</em> your latency choice, whether or not you framed it that way.</p>
<p>Pick the latency you can accept, multiply by the service rate, and that is your maximum queue length. Everything past it gets rejected.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>Point a load generator at your service at three times its capacity. Watch:</p>
<ul><li>Does memory grow without bound?</li><li>Does latency grow without bound?</li><li>Does anything get rejected, or does everything just get slower?</li><li>After you stop the load, how long until it recovers?</li></ul>
<p>That last one is the important one. A system that takes twenty minutes to recover from a two-minute overload has a backpressure problem, and it will turn a small incident into a large one on a day you did not choose.</p>]]></content:encoded></item><item><title>Cross-platform is a promise you make to your budget</title><link>https://readme.news/cross-platform-is-a-promise-you-make-to-your-budget/</link><guid isPermaLink="true">https://readme.news/cross-platform-is-a-promise-you-make-to-your-budget/</guid><pubDate>Tue, 30 Jun 2026 09:00:00 +0000</pubDate><description>Write once, run anywhere, debug everywhere. The honest accounting of what each approach actually costs.</description><content:encoded><![CDATA[<p>The cross-platform question — one codebase or several — gets argued as a technical matter and is mostly an organizational one. Here is the honest accounting.</p>
<h2 id="the-actual-trade">the actual trade<a class="anchor" href="#the-actual-trade" aria-label="link to this section">#</a></h2>
<p><strong>Native</strong> gives you: full platform capability, best performance, platform-idiomatic <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>, immediate access to new OS features, and the best debugging tools.</p>
<p>It costs you: two or three codebases, two or three teams, features implemented multiple times, and behavior that diverges over years in ways nobody tracks.</p>
<p><strong>Cross-platform</strong> gives you: one codebase, one team, features implemented once, consistent behavior.</p>
<p>It costs you: a framework layer between you and the platform, a lag before new OS features are available, worse debugging when the problem is in the bridge, and an interface that is either non-idiomatic on every platform or requires platform-specific work anyway.</p>
<h2 id="the-thing-the-arguments-miss">the thing the arguments miss<a class="anchor" href="#the-thing-the-arguments-miss" aria-label="link to this section">#</a></h2>
<p><strong>The largest cost of cross-platform is not performance. It is the <a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">escape hatch</a>.</strong></p>
<p>Everything is fine until you need something the framework does not support. Then you are writing platform-specific native code, plus a bridge, plus a fallback, plus tests for all three — and you have the complexity of native development <em>plus</em> the framework.</p>
<p>This happens. It always happens. The question is how often, and that depends entirely on what your app does.</p>
<p><strong>Low escape-hatch pressure:</strong> content, commerce, forms, dashboards, CRUD, most business applications. These use the standard widget set and standard capabilities. Cross-platform works well.</p>
<p><strong>High escape-hatch pressure:</strong> camera and media processing, background execution, Bluetooth and hardware peripherals, complex custom rendering, deep platform integration (widgets, shortcuts, app extensions), anything real-time.</p>
<p>For high-pressure apps, cross-platform frequently costs more than native, because you pay for the framework and then write the native code anyway.</p>
<p><strong>Be honest about which one you are</strong> before choosing. Most teams that regret their choice were high-pressure and assessed themselves as low.</p>
<h2 id="the-options-honestly">the options, honestly<a class="anchor" href="#the-options-honestly" aria-label="link to this section">#</a></h2>
<p><strong>React Native.</strong> Mature, large ecosystem, native widgets. The new architecture removed the old asynchronous bridge, which was the main performance complaint. Best choice if your team is already React.</p>
<p><strong>Flutter.</strong> Renders its own widgets, which means true visual consistency and a non-native feel that some users notice and most do not. Excellent performance for custom interfaces. Dart is a real adoption cost for a team that does not know it.</p>
<p><strong>Kotlin Multiplatform.</strong> Share business logic, write native UI. This is the approach I find most defensible: the logic layer — networking, models, validation, persistence — is where duplication is most wasteful and least visible to users, and the UI layer is where platform idiom matters most.</p>
<p>Requires two UI implementations, which is the point rather than a limitation.</p>
<p><strong>Web technologies in a wrapper.</strong> Fastest to build if you have web engineers, and the platform feel is the weakest. Fine for content-heavy applications, poor for anything interaction-heavy.</p>
<p><strong>Native.</strong> Still correct for a large category, and the category is smaller than native advocates believe.</p>
<h2 id="the-organizational-question-that-actually-decides-it">the organizational question that actually decides it<a class="anchor" href="#the-organizational-question-that-actually-decides-it" aria-label="link to this section">#</a></h2>
<p><strong>Do you have or can you hire two platform teams?</strong></p>
<p>If yes, native is viable and gives you the best result.</p>
<p>If no — and for most companies below a certain size the answer is no — the choice is between cross-platform and shipping on one platform. Framed that way, the decision is usually easy.</p>
<p><strong>What is your feature velocity?</strong></p>
<p>If you ship a large feature monthly, implementing it twice is a permanent 2× cost on your most expensive activity. If you ship quarterly, the duplication matters less.</p>
<p><strong>How much does platform idiom matter to your users?</strong></p>
<p>For a consumer app competing on polish: a lot. For an internal tool: nothing. For a B2B product where the buyer is not the user: less than you think.</p>
<h2 id="the-hybrid-that-most-people-should-consider">the hybrid that most people should consider<a class="anchor" href="#the-hybrid-that-most-people-should-consider" aria-label="link to this section">#</a></h2>
<p>Share the logic, write the UI natively.</p>
<p>The business logic — API clients, data models, validation, offline storage, sync, analytics — is genuinely identical across platforms and duplicating it produces bugs that exist on one platform and not the other, which are the worst bugs to diagnose.</p>
<p>The UI is where platform conventions matter, where users notice, and where the framework abstraction costs the most.</p>
<p>This is more work than full cross-platform and less than full native, and it puts the sharing where the value is.</p>
<h2 id="the-thing-i-would-tell-someone-deciding">the thing I would tell someone deciding<a class="anchor" href="#the-thing-i-would-tell-someone-deciding" aria-label="link to this section">#</a></h2>
<p>Prototype the hardest thing first.</p>
<p>Not the login screen. The thing you are worried about — the camera flow, the background sync, the complex list, the offline behavior. Build that on your candidate stack, in a week.</p>
<p>You will learn more from that week than from any amount of comparison, and you will learn it while changing your mind is still cheap.</p>
<p>The teams that regret their choice almost always chose based on a comparison article, built the easy part first, and discovered the hard part in month five.</p>]]></content:encoded></item><item><title>Retries: a complete guide to not making it worse</title><link>https://readme.news/retries-a-complete-guide-to-not-making-it-worse/</link><guid isPermaLink="true">https://readme.news/retries-a-complete-guide-to-not-making-it-worse/</guid><pubDate>Mon, 29 Jun 2026 09:00:00 +0000</pubDate><description>The most common way a small incident becomes a large one is a retry policy written without thinking about aggregate behavior.</description><content:encoded><![CDATA[<p>Retries are the most commonly implemented and most commonly wrong piece of resilience engineering. A retry policy that seems obviously correct in isolation is frequently the mechanism that turns a brief degradation into a full outage.</p>
<h2 id="the-failure-mode">the failure mode<a class="anchor" href="#the-failure-mode" aria-label="link to this section">#</a></h2>
<p>A downstream service slows down. Every caller times out. Every caller retries.</p>
<p>The downstream now receives double its normal traffic while already struggling. More requests time out. More retries. The load multiplies with every round.</p>
<p>The original problem might have been a five-second blip. The retry storm keeps the service down for twenty minutes, and it stays down after the original cause is resolved because the queued retries are still arriving.</p>
<p>This is metastable failure and retries are its most common cause.</p>
<h2 id="the-rules">the rules<a class="anchor" href="#the-rules" aria-label="link to this section">#</a></h2>
<p><strong>1. Only retry idempotent operations.</strong></p>
<p>A retried non-idempotent operation charges the card twice. If you need to retry a write, make it idempotent first with an <a class="xref" href="/idempotency-is-the-only-distributed-systems-concept-you-need/" title="Idempotency is the only distributed systems concept you need">idempotency</a> key.</p>
<p><strong>2. Only retry retryable errors.</strong></p>
<p>A 400 will be a 400 next time. Retrying it wastes a request and delays the error the caller needs to see.</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">retry</th><th style="text-align:left">do not retry</th></tr></thead><tbody><tr><td style="text-align:left">connection refused / reset</td><td style="text-align:left">400, 401, 403, 404</td></tr><tr><td style="text-align:left">timeout</td><td style="text-align:left">422</td></tr><tr><td style="text-align:left">502, 503, 504</td><td style="text-align:left">any deterministic validation failure</td></tr><tr><td style="text-align:left">429 (respect <code>Retry-After</code>)</td><td style="text-align:left">501</td></tr></tbody></table></div>
<p>The one people get wrong: <strong>500 is ambiguous.</strong> It might be transient, it might be a deterministic bug. Retrying it once is usually reasonable; retrying it five times is usually pointless.</p>
<p><strong>3. Exponential backoff with jitter. Always jitter.</strong></p>
<p>Without jitter, all your clients retry at the same moments, and you have built a synchronized load generator.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def delay(attempt, base=0.1, cap=30):
    return random.uniform(0, min(cap, base * (2 ** attempt)))</code></pre></div>
<p>That is full jitter, and it is the recommended default. It spreads retries across the whole interval, which is what you want. Half jitter — <code>d/2 + random(0, d/2)</code> — is a reasonable alternative when you want a guaranteed minimum delay.</p>
<p>The version without jitter is the one everyone writes first and it is the one that causes the storm.</p>
<p><strong>4. Cap the attempts and cap the total time.</strong></p>
<p>Three attempts, usually. And a total deadline — if the caller is going to give up after two seconds, retrying at three seconds is pure waste and it is load on a struggling service.</p>
<p>Propagate the deadline. If your caller has 500 ms left, your retry budget is 500 ms, not your configured default.</p>
<p><strong>5. Do not retry at every layer.</strong></p>
<p>This is the one that produces the shocking numbers. Three attempts at the HTTP client, three in the service wrapper, three in the caller, three at the gateway: 3⁴ = 81 requests for one logical call.</p>
<p><strong>Retry at exactly one layer.</strong> Usually the outermost one that has the context to decide. Every other layer fails fast and propagates.</p>
<p>Audit this. Most systems that have grown organically retry at three or four layers and nobody knows.</p>
<p><strong>6. Use a retry budget.</strong></p>
<p>The refinement that actually prevents storms: cap retries as a <em>fraction of total traffic</em>, not per request.</p>
<div class="code"><pre><code>if retries_in_window / requests_in_window &gt; 0.1:
    do_not_retry()</code></pre></div>
<p>When things are healthy, occasional retries are well under the budget and everything works. When things are broken, the budget is exhausted immediately and retries stop entirely — exactly when they would do the most harm.</p>
<p>This single mechanism converts retries from a failure amplifier into a bounded safety net, and it is not widely implemented.</p>
<p><strong>7. Circuit break.</strong></p>
<p>When a downstream is clearly failing, stop calling it. Fail fast, return a cached or degraded response, and probe occasionally to see if it has recovered.</p>
<p>Three states: closed (normal), open (failing fast), half-open (probing). The half-open state must allow only a trickle — if you send full traffic at a recovering service you will knock it over again.</p>
<h2 id="the-server-side">the server side<a class="anchor" href="#the-server-side" aria-label="link to this section">#</a></h2>
<p>The other half, which is usually forgotten.</p>
<p><strong>Send <code>Retry-After</code> on 429 and 503.</strong> Then clients that respect it retry at a time you chose rather than a time they chose.</p>
<p><strong>Shed load rather than queueing it.</strong> A request that will time out anyway should be rejected immediately, not queued. Queueing under overload increases latency for everything without increasing throughput, and the requests you eventually serve have often already been abandoned.</p>
<p><strong>Prioritize.</strong> Under load, serve health checks and critical paths, shed the rest. An unprioritized overload sheds randomly, which means your health checks fail and your orchestrator kills healthy instances.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>Take your service. Make a downstream dependency return 503 for everything. Watch the request rate at the downstream.</p>
<p>If it goes up by more than a small factor, your retry configuration will cause an outage. You have just not had the trigger yet.</p>]]></content:encoded></item><item><title>Configuration is the most under-designed part of your system</title><link>https://readme.news/configuration-is-the-most-under-designed-part-of-your-system/</link><guid isPermaLink="true">https://readme.news/configuration-is-the-most-under-designed-part-of-your-system/</guid><pubDate>Wed, 24 Jun 2026 09:00:00 +0000</pubDate><description>It has no type system, no tests, no review, and it causes a disproportionate share of outages.</description><content:encoded><![CDATA[<p>Look at the last ten significant outages you can remember reading about. A disproportionate number were caused by configuration, not code.</p>
<p>A config file with a wrong value. A <a class="xref" href="/the-graph-nobody-drew/" title="The graph nobody drew">feature flag</a> flipped. A generated file that doubled in size. A DNS record. A permission change.</p>
<p>Configuration receives a fraction of the engineering rigor that code does, and it has comparable power to break things.</p>
<h2 id="why-it-goes-wrong">why it goes wrong<a class="anchor" href="#why-it-goes-wrong" aria-label="link to this section">#</a></h2>
<p><strong>No <a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">type system</a>.</strong> A YAML file will happily contain <code>timeout: 30s</code> where the code expects an integer, or <code>enabled: "false"</code> which is a truthy string.</p>
<p><strong>No tests.</strong> Nobody writes a test for their config.</p>
<p><strong>No review, or perfunctory review.</strong> A config change is "just a value" and gets approved in ten seconds.</p>
<p><strong>Deployed differently from code.</strong> Frequently faster, frequently without staging, frequently without a canary. The safety mechanisms built for code deployment routinely do not cover config.</p>
<p><strong>Environment drift.</strong> Staging and production differ in ways nobody has enumerated, so staging validates a configuration that is not production's.</p>
<p><strong>No rollback story.</strong> "What was this value before" is often unanswerable.</p>
<h2 id="the-fixes">the fixes<a class="anchor" href="#the-fixes" aria-label="link to this section">#</a></h2>
<p><strong>1. Parse and validate at startup, not at use.</strong></p>
<p>The worst failure mode is a config error that manifests three hours later when a rarely-used code path reads a malformed value.</p>
<p>Validate everything at boot. Fail loudly and immediately if anything is wrong.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">class Settings(BaseModel):
    database_url: PostgresDsn
    timeout_ms: int = Field(gt=0, le=60_000)
    max_connections: int = Field(ge=1, le=1000)
    feature_new_checkout: bool = False

settings = Settings(**load_config())   # raises at startup, with a clear message</code></pre></div>
<p>An application that will not start is far better than one that starts and behaves wrongly.</p>
<p><strong>2. Types, with a schema.</strong></p>
<p>Whatever your language, there is a library that turns untyped config into a validated typed object. Use it. This eliminates the entire category of string-that-should-be-a-number bugs.</p>
<p><strong>3. Config changes go through the same pipeline as code.</strong></p>
<p>Version control. Review. Staging. Canary. Rollback.</p>
<p>The argument against is that config changes need to be fast, especially for incident response. That is a real need and the answer is a small explicitly-defined set of emergency levers — kill switches, <a class="xref" href="/rate-limits-are-a-product-decision-not-an-infrastructure-one/" title="Rate limits are a product decision, not an infrastructure one">rate limits</a> — with a fast path, and everything else on the normal pipeline.</p>
<p>Not "all config is fast" because that is how config takes down your service.</p>
<p><strong>4. Validate generated configuration.</strong></p>
<p>If a config file is produced by a program, that program can be wrong. A size check against the previous version, a schema check, a sanity check on record count.</p>
<p>This is the specific failure that has caused several high-profile outages: a generated file changed unexpectedly and propagated globally before anyone looked at it.</p>
<p><strong>5. Make the environment explicit and diff it.</strong></p>
<p>You should be able to answer "how does staging differ from production" with a command. If you cannot, staging is not validating production.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">config-diff staging production</code></pre></div>
<p><strong>6. Log the effective configuration at startup.</strong></p>
<p>Not the file — the resolved values after defaults, overrides, and environment variables are applied. Redact secrets. This is the single most useful thing for debugging "it works on my machine," because the effective config is frequently not what anyone thinks it is.</p>
<p><strong>7. <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">Feature flags</a> need the same discipline as code.</strong></p>
<p>Owner. Expiry. Test both branches. Log the flag state on every event. A flag flip is a production change and should be treated as one.</p>
<h2 id="the-hierarchy-that-works">the hierarchy that works<a class="anchor" href="#the-hierarchy-that-works" aria-label="link to this section">#</a></h2>
<p>Most systems end up with layered configuration and the layering should be explicit and simple:</p>
<div class="code"><pre><code>defaults in code
  ← config file
    ← environment variables
      ← command line flags</code></pre></div>
<p>Later overrides earlier. Log which layer each effective value came from when debugging.</p>
<p><strong>Keep the layers few.</strong> Systems with six overlapping sources of configuration — defaults, file, environment, a service, a database table, a flag system — produce values nobody can trace. Each layer you add makes "why is this value what it is" harder to answer.</p>
<h2 id="secrets">secrets<a class="anchor" href="#secrets" aria-label="link to this section">#</a></h2>
<p>Not the same thing as configuration and should not be in the same place.</p>
<ul><li>Never in the config file. Never in version control. Never in the image.</li><li>A secret manager, injected at runtime.</li><li>Rotated on a schedule, and the rotation must be tested.</li><li>Never logged. Redact by default at the logger, not at each call site — because somebody will forget.</li></ul>
<h2 id="the-general-point">the general point<a class="anchor" href="#the-general-point" aria-label="link to this section">#</a></h2>
<p>Configuration is the input to your program that is most likely to be wrong and least likely to be checked.</p>
<p>Every rigor you apply to code — types, validation, tests, review, staged rollout, rollback — applies to configuration, and applying it costs a day of setup.</p>
<p>The reason it does not happen is that config does not feel like code. It is data, and data feels safe.</p>
<p>It is not data. It is the arguments to your program, and passing wrong arguments to a program is how programs go wrong.</p>]]></content:encoded></item><item><title>Rewrites: when they actually work</title><link>https://readme.news/rewrites-when-they-actually-work/</link><guid isPermaLink="true">https://readme.news/rewrites-when-they-actually-work/</guid><pubDate>Mon, 01 Jun 2026 09:00:00 +0000</pubDate><description>The received wisdom is never rewrite. The received wisdom is mostly right and has three real exceptions.</description><content:encoded><![CDATA[<p>The canonical advice is that rewriting from scratch is the single worst strategic mistake a software company can make. That advice is twenty-five years old and it is mostly still correct.</p>
<p>It has exceptions, and knowing which situation you are in matters more than knowing the rule.</p>
<h2 id="why-rewrites-usually-fail">why rewrites usually fail<a class="anchor" href="#why-rewrites-usually-fail" aria-label="link to this section">#</a></h2>
<p><strong>The old system's behavior is undocumented and load-bearing.</strong> Every strange conditional in a decade-old codebase is a bug report somebody filed. You will not find them by reading the code, because the code does not say why. You will find them by shipping the rewrite and having those bugs re-reported.</p>
<p><strong>The rewrite has no users, so it gets no feedback.</strong> The old system is being exercised by real traffic continuously. The new one is exercised by your test suite, which encodes what you think it should do.</p>
<p><strong>Feature parity is a moving target.</strong> The old system keeps getting features, because the business does not stop. You are chasing a target that recedes, and the chase consumes the time you budgeted for the rewrite.</p>
<p><strong>Nobody can justify continuing past month six.</strong> The rewrite has produced nothing users can see. The pressure to redirect the team to visible work is enormous and usually wins, leaving you with two systems.</p>
<p><strong>The knowledge is in the people, not the code.</strong> And the people who knew have frequently left, which is often why the rewrite was proposed.</p>
<h2 id="the-three-cases-where-it-is-right">the three cases where it is right<a class="anchor" href="#the-three-cases-where-it-is-right" aria-label="link to this section">#</a></h2>
<p><strong>1. The platform is dead.</strong></p>
<p>The framework is unmaintained, the language runtime is past end-of-life and getting CVEs, the vendor discontinued the product, the hardware is unobtainable.</p>
<p>This is not a preference — you are being forced. Rewrite, and be grateful you found out before the security incident.</p>
<p><strong>2. The domain model is fundamentally wrong.</strong></p>
<p>Not "the code is messy." The core abstractions do not correspond to reality, and every feature requires working around them.</p>
<p>The test: can you name a specific business capability that is <em>impossible</em>, not just awkward, in the current model? "We cannot support customers with more than one billing entity, and the fix requires changing what a customer is."</p>
<p>If you cannot name one, you have messy code, not a wrong model, and messy code is fixed by refactoring.</p>
<p><strong>3. The system is small enough that the rewrite is short.</strong></p>
<p>If the whole thing is three weeks of work, the calculus changes entirely. The risk of a three-week rewrite is bounded. Just do it.</p>
<p>The received wisdom is about large systems, and people apply it to small ones where it does not hold.</p>
<h2 id="how-to-do-it-when-you-must">how to do it when you must<a class="anchor" href="#how-to-do-it-when-you-must" aria-label="link to this section">#</a></h2>
<p><strong>Strangler fig, never big bang.</strong></p>
<p>Put a routing layer in front of the old system. Move one endpoint at a time to the new implementation. Route traffic gradually. When everything has moved, delete the old system.</p>
<div class="code"><pre><code>             ┌─→ new service (endpoints A, B)
client → router
             └─→ legacy monolith (everything else)</code></pre></div>
<p>Properties this gives you:</p>
<ul><li><strong>Value ships continuously.</strong> Every migrated endpoint is a delivered improvement.</li><li><strong>Risk is bounded per endpoint.</strong> If one goes wrong, route it back.</li><li><strong>The project survives leadership changes</strong>, because it is producing visible progress the whole time.</li><li><strong>You learn the old system's real behavior</strong> incrementally, at the point where you have to reimplement it.</li></ul>
<p><strong>Run both and compare.</strong> For a while, send traffic to both implementations, return the old one's response, and log the differences. This is the single most effective technique for discovering undocumented behavior, and it finds things no amount of code reading would have.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">old = legacy.compute(req)
try:
    new = rewritten.compute(req)
    if new != old:
        log.warn("divergence", request=req, old=old, new=new)
except Exception as e:
    log.error("new path failed", error=e, request=req)
return old      # old is still authoritative</code></pre></div>
<p>Run that for weeks. The divergence log is your actual specification.</p>
<p><strong>Freeze the old system's features.</strong> If the business keeps adding to it, you will never catch up. This requires an explicit organizational decision and it is usually the hardest part.</p>
<p><strong>Set a deletion date and mean it.</strong> The worst outcome is two systems forever. If the migration stalls at 70%, you have doubled your maintenance surface permanently.</p>
<h2 id="the-question-to-ask-first">the question to ask first<a class="anchor" href="#the-question-to-ask-first" aria-label="link to this section">#</a></h2>
<p>Before any rewrite: <strong>what specifically will be better, and how will you know?</strong></p>
<p>If the answer is "the code will be cleaner," that is not a business outcome and the project will lose its funding to something that is.</p>
<p>If the answer is "features in this area will take one week instead of three, and here are the last five that took three," you have a case.</p>
<p>The rewrites that succeed are the ones where somebody could state the payback in a sentence. The ones that fail are the ones motivated by taste, and taste is not wrong — it is just not fundable.</p>]]></content:encoded></item><item><title>Edge computing, honestly</title><link>https://readme.news/edge-computing-honestly/</link><guid isPermaLink="true">https://readme.news/edge-computing-honestly/</guid><pubDate>Fri, 29 May 2026 09:00:00 +0000</pubDate><description>Running code close to users is a real win for a narrow set of workloads and a complication for everything else.</description><content:encoded><![CDATA[<p>Edge computing has been sold as a general architectural improvement: run your code in hundreds of locations near users, everything gets faster.</p>
<p>The physics is real. The applicability is narrower than the marketing, and the reason is data.</p>
<h2 id="the-physics">the physics<a class="anchor" href="#the-physics" aria-label="link to this section">#</a></h2>
<p>Light in fiber travels roughly 200,000 km/s. New York to London and back is about 55 ms of pure propagation, before any processing. Add TLS handshakes, TCP setup, and real-world routing that is not a great circle, and a cross-Atlantic round trip is frequently 100 ms or more.</p>
<p>Running code 20 ms from the user instead of 120 ms is a genuine improvement and it is not achievable any other way.</p>
<h2 id="the-problem">the problem<a class="anchor" href="#the-problem" aria-label="link to this section">#</a></h2>
<p>Your data is not at the edge. It is in a database, in one region, and if your edge function needs it, you have moved the compute closer to the user and left the round trip in place — plus added a hop.</p>
<div class="code"><pre><code>user → edge (5ms) → origin database (120ms) → edge → user</code></pre></div>
<p>That is slower than the user talking to the origin directly, because you added a hop to a path that was always dominated by the database call.</p>
<p>This is the single most common edge computing mistake and it is easy to make, because the architecture diagram looks right.</p>
<h2 id="what-edge-is-genuinely-good-for">what edge is genuinely good for<a class="anchor" href="#what-edge-is-genuinely-good-for" aria-label="link to this section">#</a></h2>
<p><strong>Anything that needs no origin data:</strong></p>
<ul><li><strong>Redirects and rewrites.</strong> URL normalization, locale routing, legacy path mapping.</li><li><strong>Authentication token validation.</strong> A signed JWT can be verified with a public key at the edge, and an invalid request never reaches your origin. This is a real win — you reject bad traffic at the perimeter.</li><li><strong>A/B test assignment.</strong> Deterministic hash of a cookie into a bucket. No state required.</li><li><strong>Header manipulation.</strong> Security headers, CORS, feature policy.</li><li><strong>Bot filtering and <a class="xref" href="/rate-limiting-the-four-algorithms-and-when-each-is-wrong/" title="Rate limiting: the four algorithms and when each is wrong">rate limiting</a>.</strong> Reject at the edge, before the request costs you anything.</li><li><strong>Personalization of cached content.</strong> Fetch the cached page, inject the user's name from a cookie, return. The expensive part stays cached.</li></ul>
<p><strong>Anything where the data is genuinely replicated to the edge:</strong></p>
<p>Several platforms now offer edge-replicated key-value and SQL storage. If your data is small, read-heavy, and tolerant of replication lag — configuration, <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">feature flags</a>, product catalogs, translations — this works well and the latency win is real.</p>
<p>The constraints are real too: writes go to a primary, replication is eventual, and storage per location is limited.</p>
<h2 id="what-edge-is-bad-for">what edge is bad for<a class="anchor" href="#what-edge-is-bad-for" aria-label="link to this section">#</a></h2>
<p><strong>Anything write-heavy.</strong> Writes need coordination. Coordination needs a primary. The primary is in one place.</p>
<p><strong>Anything requiring strong consistency.</strong> By definition, this needs coordination, which needs round trips, which is what you were trying to avoid.</p>
<p><strong>Anything with a large working set.</strong> You cannot replicate a terabyte to three hundred locations.</p>
<p><strong>Anything computationally heavy.</strong> Edge runtimes have tight CPU and memory limits. They are designed for milliseconds of work per request.</p>
<p><strong>Anything that needs a specific runtime.</strong> Most edge platforms run a constrained JavaScript or <a class="xref" href="/the-component-model-and-the-plugin-problem/" title="The component model and the plugin problem">WebAssembly</a> environment. Your Python dependency with a C extension is not going there.</p>
<h2 id="the-architecture-that-works">the architecture that works<a class="anchor" href="#the-architecture-that-works" aria-label="link to this section">#</a></h2>
<p>Layered, with each layer doing what it is good at:</p>
<div class="code"><pre><code>edge      → auth check, rate limit, routing, cached content, header work
regional  → application logic, caching, session state
origin    → the database, the writes, the truth</code></pre></div>
<p>Most requests are answered at the edge from cache. Some go to a regional application tier. Few reach the origin.</p>
<p>That is a CDN with programmability, which is what edge computing actually is, and framing it that way produces much better decisions than framing it as "serverless everywhere."</p>
<h2 id="the-thing-to-measure-first">the thing to measure first<a class="anchor" href="#the-thing-to-measure-first" aria-label="link to this section">#</a></h2>
<p>Before adopting any of this: <strong>where does your latency actually go?</strong></p>
<p>Break down a typical request:</p>
<ul><li>DNS</li><li>TLS handshake</li><li>Network round trip</li><li>Time to first byte at origin</li><li>Origin processing</li><li>Database time within that</li><li>Response transfer</li></ul>
<p>If origin processing is 400 ms and network is 40 ms, moving compute to the edge addresses 40 ms of a 440 ms problem. Fix the 400 first.</p>
<p>This is the most common reason edge adoption disappoints: it was applied to a latency problem that was not a network problem.</p>
<h2 id="the-honest-summary">the honest summary<a class="anchor" href="#the-honest-summary" aria-label="link to this section">#</a></h2>
<p>Edge is a very good CDN with programmability, and that is genuinely valuable — it lets you do real work at the perimeter that used to require an origin request.</p>
<p>It is not a general application platform, and the platforms selling it as one are selling the constraint as a feature.</p>
<p>Use it for the perimeter. Keep your data where it can be consistent.</p>]]></content:encoded></item><item><title>Idempotency is the only distributed systems concept you need</title><link>https://readme.news/idempotency-is-the-only-distributed-systems-concept-you-need/</link><guid isPermaLink="true">https://readme.news/idempotency-is-the-only-distributed-systems-concept-you-need/</guid><pubDate>Tue, 26 May 2026 09:00:00 +0000</pubDate><description>Not the only one. But if you only internalize one, make it this one, because it makes most of the others survivable.</description><content:encoded><![CDATA[<p>The network will deliver your message twice. Or zero times. Or once, but you will not find out, so you will send it again.</p>
<p>This is not an edge case. It is the normal operating condition of every distributed system, and idempotency is what makes it survivable.</p>
<h2 id="the-fundamental-problem">the fundamental problem<a class="anchor" href="#the-fundamental-problem" aria-label="link to this section">#</a></h2>
<p>You send a request. The connection times out.</p>
<p>Did it succeed?</p>
<p>You do not know. There are three possibilities:</p>
<ol><li>The request never arrived. Retry is correct.</li><li>The request arrived and failed. Retry is correct.</li><li><strong>The request arrived, succeeded, and the response was lost.</strong> Retry charges the card twice.</li></ol>
<p>You cannot distinguish these from the client. This is not a limitation of your tooling; it is a theorem about asynchronous networks. There is no protocol that resolves it.</p>
<p>The only resolution is to make retrying safe.</p>
<h2 id="what-idempotency-means">what idempotency means<a class="anchor" href="#what-idempotency-means" aria-label="link to this section">#</a></h2>
<p>An operation is idempotent if performing it multiple times has the same effect as performing it once.</p>
<p>Naturally idempotent:</p>
<ul><li><code>PUT /user/123 {"name": "Dom"}</code> — set to a value.</li><li><code>DELETE /user/123</code> — the second one finds nothing to do.</li><li><code>SET balance = 100</code> — absolute assignment.</li></ul>
<p>Not idempotent:</p>
<ul><li><code>POST /orders</code> — creates a new one each time.</li><li><code>UPDATE accounts SET balance = balance + 100</code> — relative change.</li><li>Sending an email.</li><li>Incrementing a counter.</li></ul>
<h2 id="the-pattern">the pattern<a class="anchor" href="#the-pattern" aria-label="link to this section">#</a></h2>
<p>For operations that are not naturally idempotent, the client supplies a key and the server remembers it.</p>
<div class="code"><span class="code-lang">http</span><pre><code class="lang-http">POST /payments
Idempotency-Key: 8f14e45f-ea5c-4b0d-9c1a-2f7e3d4a5b6c

{"amount_cents": 4200, "currency": "usd", "source": "card_x"}</code></pre></div>
<p>Server side:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def create_payment(key, request):
    with transaction():
        existing = lookup(key)
        if existing:
            if existing.request_hash != hash(request):
                raise Conflict("key reused with different parameters")
            return existing.response      # replay, do not re-execute

        result = charge_the_card(request)
        store(key, hash(request), result)
        return result</code></pre></div>
<p>Four details that matter and are usually missed:</p>
<p><strong>Store the result, not just the key.</strong> A retry should return the original response, not a "already processed" error. The client's retry should look like a successful first attempt.</p>
<p><strong>Hash the request.</strong> If the same key arrives with different parameters, that is a client bug and you should say so rather than silently returning the wrong result.</p>
<p><strong>Same transaction.</strong> The lookup, the work, and the store must be atomic. Otherwise two concurrent <a class="xref" href="/retries-a-complete-guide-to-not-making-it-worse/" title="Retries: a complete guide to not making it worse">retries</a> both find nothing and both execute.</p>
<p><strong>Expire the keys.</strong> Days, not forever. Storage is not free and a key from a year ago is not going to be retried.</p>
<h2 id="where-else-this-applies">where else this applies<a class="anchor" href="#where-else-this-applies" aria-label="link to this section">#</a></h2>
<p><strong>Message consumers.</strong> At-least-once delivery guarantees duplicates. Same pattern: an idempotency key on the message, a record of processed keys, both in one transaction.</p>
<p><strong>Webhooks you send.</strong> Include an event ID. Your receivers will need it, and if you do not provide one they will invent something worse.</p>
<p><strong>Webhooks you receive.</strong> Assume duplicates. Every major provider retries and several will send the same event twice under normal operation.</p>
<p><strong>Deployment and provisioning.</strong> "Create this resource" should succeed if it already exists in the right state. This is why declarative infrastructure tools work and imperative scripts do not.</p>
<p><strong>Migrations and backfills.</strong> They get interrupted. They must be safe to re-run from the beginning.</p>
<h2 id="the-design-advice">the design advice<a class="anchor" href="#the-design-advice" aria-label="link to this section">#</a></h2>
<p><strong>Prefer absolute over relative.</strong> <code>SET balance = 100</code> is idempotent. <code>ADD 100</code> is not. When you have the choice, take the absolute form.</p>
<p>Where you cannot — and you often cannot, because concurrent updates need relative operations — use a version or a compare-and-set:</p>
<div class="code"><span class="code-lang">sql</span><pre><code class="lang-sql">UPDATE accounts SET balance = balance + 100, version = version + 1
WHERE id = $1 AND version = $2;</code></pre></div>
<p>Zero rows updated means someone else got there first, and you know it.</p>
<p><strong>Let the client generate the key.</strong> The client knows whether this is a new request or a retry. The server cannot tell. A server-generated key defeats the entire purpose.</p>
<p><strong>Document it.</strong> If your API supports idempotency keys, say so prominently, say how long they are honored, and say what happens on key reuse with different parameters. Consumers who do not know about the mechanism will not use it, and then they will double-charge someone and it will be your incident too.</p>
<h2 id="why-this-is-the-one-to-internalize">why this is the one to internalize<a class="anchor" href="#why-this-is-the-one-to-internalize" aria-label="link to this section">#</a></h2>
<p>Most distributed systems failures are one of: a duplicate, a message lost, or an out-of-order arrival.</p>
<p>Idempotency makes duplicates safe. Retries make loss survivable — and retries require idempotency to be safe. Which leaves ordering, which is a narrower problem that fewer systems actually have.</p>
<p>So: one property, correctly implemented, defuses the majority of what goes wrong. That is a very good ratio, and it is why this is the thing to get right before anything else.</p>]]></content:encoded></item><item><title>The queue is the architecture</title><link>https://readme.news/the-queue-is-the-architecture/</link><guid isPermaLink="true">https://readme.news/the-queue-is-the-architecture/</guid><pubDate>Fri, 15 May 2026 09:00:00 +0000</pubDate><description>Most scaling problems are solved by making something asynchronous. Most reliability problems are caused by doing it badly.</description><content:encoded><![CDATA[<p>The single most effective architectural move available to most systems is: take the slow thing out of the request path and put it in a queue.</p>
<p>It is also the move that introduces the most subtle <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a>, and the gap between "we added a queue" and "we added a queue correctly" is large.</p>
<h2 id="what-it-buys">what it buys<a class="anchor" href="#what-it-buys" aria-label="link to this section">#</a></h2>
<p><strong>Latency.</strong> The user gets a response when the work is accepted, not when it is done. A checkout that returns in 80 ms and sends the confirmation email asynchronously is a much better product than one that returns in 900 ms.</p>
<p><strong>Absorbing spikes.</strong> A queue is a buffer. Traffic that would overwhelm a synchronous system accumulates and drains. This is the difference between a slow period and an outage.</p>
<p><strong>Isolation.</strong> If the email provider is down, checkout still works. The messages accumulate and send later.</p>
<p><strong>Retry for free.</strong> A failed message goes back on the queue. A failed synchronous call is a user-visible error.</p>
<h2 id="what-it-costs">what it costs<a class="anchor" href="#what-it-costs" aria-label="link to this section">#</a></h2>
<p><strong>Eventual consistency, everywhere.</strong> The user completed checkout and the confirmation has not arrived. The record exists and the search index does not have it. Every asynchronous boundary introduces a window where the system is inconsistent, and your UI has to be honest about it.</p>
<p><strong>Debugging across the boundary.</strong> A synchronous stack trace tells you the whole story. An asynchronous failure requires correlating a producer, a broker, and a consumer, possibly hours apart.</p>
<p><strong>Ordering.</strong> Most <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> do not guarantee it, or guarantee it only within a partition. If your consumer must process events in order, that is a design constraint that reaches back into how you partition.</p>
<p><strong>Duplicate delivery.</strong> Almost all queues are at-least-once. Your consumer <em>will</em> receive the same message twice. If that is not safe, you have a bug that appears under load, weeks after launch.</p>
<h2 id="the-rules">the rules<a class="anchor" href="#the-rules" aria-label="link to this section">#</a></h2>
<p><strong>1. Consumers must be idempotent. Non-negotiable.</strong></p>
<p>At-least-once delivery means duplicates. The consumer must produce the same result whether it processes a message once or five times.</p>
<p>The usual implementation: a natural <a class="xref" href="/idempotency-is-the-only-distributed-systems-concept-you-need/" title="Idempotency is the only distributed systems concept you need">idempotency</a> key, and a record of processed keys.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def handle(msg):
    key = msg.idempotency_key
    with tx():
        if already_processed(key):
            return
        do_the_work(msg)
        mark_processed(key)</code></pre></div>
<p>The <code>mark_processed</code> must be in the same transaction as the work, or you have moved the race rather than eliminated it.</p>
<p><strong>2. Every queue needs a dead letter queue, and someone must watch it.</strong></p>
<p>A message that fails repeatedly must go somewhere. A DLQ nobody monitors is a place where data goes to be silently lost, which is worse than an error, because errors are visible.</p>
<p>Alert on DLQ depth. Not on it being non-zero — on it growing.</p>
<p><strong>3. Retry with backoff and jitter, and cap the attempts.</strong></p>
<p>Immediate retry on a failing downstream is a denial of service you are performing against yourself. Exponential backoff with jitter, a maximum attempt count, then the DLQ.</p>
<p><strong>4. Monitor queue depth and age, not just throughput.</strong></p>
<p>Throughput looks healthy right up until it does not. The metrics that tell you something is wrong:</p>
<ul><li><strong>Depth</strong> — how many messages are waiting.</li><li><strong>Oldest message age</strong> — the most useful single metric. If it is growing, your consumers cannot keep up, and you know how far behind you are in time rather than in count.</li></ul>
<p><strong>5. Decide what happens when the queue is full.</strong></p>
<p>It will be. Reject the producer, drop messages, or block? Each is right in different cases and the default is usually wrong for you. An unbounded queue is not a solution; it is a memory leak with extra steps.</p>
<p><strong>6. Keep the payload small and the reference stable.</strong></p>
<p>Put an ID in the message, not the whole object. The consumer fetches current state. This avoids stale data in the message and keeps the broker fast.</p>
<p>The exception: if you need the state <em>as it was</em> when the event occurred, put it in the message deliberately, and say so.</p>
<p><strong>7. Version your message schema from day one.</strong></p>
<p>Producers and consumers deploy independently. Old consumers will see new messages. Include a version field. Make additive changes only, or handle both shapes.</p>
<h2 id="the-thing-that-surprises-people">the thing that surprises people<a class="anchor" href="#the-thing-that-surprises-people" aria-label="link to this section">#</a></h2>
<p><strong>Queues do not reduce load. They defer it.</strong></p>
<p>If your consumer processes 100 messages per second and you produce 150, you do not have a working system with a buffer. You have a system that is failing slowly, and the queue depth graph is a countdown.</p>
<p>A queue absorbs <em>bursts</em>. It does not fix a sustained capacity deficit, and the failure mode when you use it that way is a queue that grows for six hours and then an incident where you are simultaneously behind and unable to catch up.</p>
<p>Alert on the age, watch the trend, and size the consumers for the sustained rate.</p>]]></content:encoded></item><item><title>Search is hard and you should probably not build it</title><link>https://readme.news/search-is-hard-and-you-should-probably-not-build-it/</link><guid isPermaLink="true">https://readme.news/search-is-hard-and-you-should-probably-not-build-it/</guid><pubDate>Wed, 29 Apr 2026 09:00:00 +0000</pubDate><description>Relevance ranking is a specialist discipline. Here&#x27;s the decision tree, and what to do at each level.</description><content:encoded><![CDATA[<p>Every product eventually adds a search box. The distance between "a search box that works" and "a search box users trust" is much larger than it looks, and most teams discover this after committing.</p>
<h2 id="the-levels">the levels<a class="anchor" href="#the-levels" aria-label="link to this section">#</a></h2>
<p><strong>Level 0: <code>LIKE '%query%'</code>.</strong></p>
<p>Works for tiny datasets. No ranking, no stemming, no typo tolerance, and a full table scan on every query.</p>
<p>Fine for an admin tool with a thousand rows. Not fine for anything a customer touches.</p>
<p><strong>Level 1: your database's full-text search.</strong></p>
<p>Postgres <code>tsvector</code>, MySQL full-text, SQLite FTS5. You get stemming, stop words, ranking, and index support.</p>
<div class="code"><span class="code-lang">sql</span><pre><code class="lang-sql">ALTER TABLE articles ADD COLUMN search tsvector
  GENERATED ALWAYS AS (
    setweight(to_tsvector('english', coalesce(title,'')), 'A') ||
    setweight(to_tsvector('english', coalesce(body,'')), 'B')
  ) STORED;

CREATE INDEX ON articles USING GIN (search);

SELECT id, title, ts_rank(search, q) AS rank
FROM articles, websearch_to_tsquery('english', $1) q
WHERE search @@ q
ORDER BY rank DESC LIMIT 20;</code></pre></div>
<p>The <code>setweight</code> calls are the part people miss: a match in the title should outrank a match in the body, and without weighting it does not.</p>
<p><strong>This handles most applications.</strong> If you have under a few million documents and your users search for terms that appear in them, stop here. One system, no synchronization problem, joins to your relational data.</p>
<p><strong>Level 2: a dedicated search engine.</strong></p>
<p>Elasticsearch, OpenSearch, Typesense, Meilisearch. You get: typo tolerance, faceting, synonyms, custom analyzers, distributed scaling, and much better relevance tuning.</p>
<p>The cost is a second system with a synchronization problem. Your search index is now eventually consistent with your database, and every write path must update both. That inconsistency will produce bugs — a deleted item still appearing in results is the classic — and handling it correctly is real work.</p>
<p><strong>Level 3: hybrid semantic search.</strong></p>
<p>Vector embeddings alongside keyword search, combined with reciprocal rank fusion or a learned reranker.</p>
<p>This handles the case where the user's words are not the document's words. "How do I cancel" should find "Terminating your subscription."</p>
<p><strong>Important:</strong> hybrid, not pure vector. Pure semantic search is bad at exact matches — product codes, error numbers, names, function names — precisely the queries where users are most certain about what they want and least tolerant of a wrong answer.</p>
<h2 id="what-makes-search-actually-good">what makes search actually good<a class="anchor" href="#what-makes-search-actually-good" aria-label="link to this section">#</a></h2>
<p>The engine is the easy part. Relevance is the hard part, and it is mostly not about the algorithm.</p>
<p><strong>Weight your fields.</strong> Title beats body. Exact phrase beats individual terms. Recent beats old, for content where recency matters.</p>
<p><strong>Use behavioral signals.</strong> What users clicked on for similar queries is the strongest relevance signal available, and it requires logging queries and clicks from day one. Retrofitting this means starting your data collection from zero.</p>
<p><strong>Handle the empty result.</strong> "No results for X" is a failure. Show something: did-you-mean, related content, popular items, a way to browse. An empty page is where users leave.</p>
<p><strong>Handle the head queries manually.</strong> A small number of queries make up a large share of volume. Look at your top hundred, check what they return, and pin the correct answer where it is wrong. This is unglamorous, takes an afternoon, and improves perceived quality more than any algorithmic change.</p>
<p><strong>Log everything.</strong> Query, result count, position clicked, whether anything was clicked. The searches that return nothing and the searches where nobody clicks are your improvement backlog, delivered for free.</p>
<h2 id="the-instrumentation-that-matters">the instrumentation that matters<a class="anchor" href="#the-instrumentation-that-matters" aria-label="link to this section">#</a></h2>
<p>Three metrics:</p>
<ol><li><strong>Zero-result rate.</strong> Should be low. Every zero-result query is a user who did not find what they wanted.</li><li><strong>Click-through rate</strong>, and the position clicked. If people consistently click the fifth result, your ranking is wrong.</li><li><strong>Query refinement rate.</strong> Users who search, then immediately search again with different words, did not find it the first time.</li></ol>
<p>Most teams have none of these and are tuning relevance by intuition.</p>
<h2 id="the-recommendation">the recommendation<a class="anchor" href="#the-recommendation" aria-label="link to this section">#</a></h2>
<p>Start at level 1. Postgres full-text search with weighted fields covers more applications than people expect, and it does not introduce a synchronization problem.</p>
<p>Move to level 2 when you have a specific complaint you cannot fix — typo tolerance, faceting at scale, or performance. Move to level 3 when you have evidence that users search for concepts rather than terms.</p>
<p>And whatever level you are at: log the queries. That data is the input to every future improvement and you cannot get it retroactively.</p>]]></content:encoded></item><item><title>The three kinds of technical debt</title><link>https://readme.news/the-three-kinds-of-technical-debt/</link><guid isPermaLink="true">https://readme.news/the-three-kinds-of-technical-debt/</guid><pubDate>Fri, 24 Apr 2026 09:00:00 +0000</pubDate><description>Deliberate, accidental, and structural. They need completely different responses and everyone calls them the same thing.</description><content:encoded><![CDATA[<p>"Technical debt" has become a phrase that means "code I do not like," which has made it useless in the conversations where it matters — the ones where you are asking for time to fix something.</p>
<p>Three distinct things wear the label. They have different causes, different costs, and different correct responses.</p>
<h2 id="1-deliberate-debt">1. deliberate debt<a class="anchor" href="#1-deliberate-debt" aria-label="link to this section">#</a></h2>
<p>You knew the better solution. You chose the faster one on purpose, for a reason.</p>
<p>"We are hard-coding this to hit the deadline. We will parameterize it in Q3."</p>
<p>This is the original meaning of the metaphor and it is a legitimate engineering tool. Borrowing against future effort to ship now is frequently correct, especially when you are not yet sure the thing will survive.</p>
<p><strong>The failure is not taking on the debt. It is not recording it.</strong></p>
<p>What makes deliberate debt manageable:</p>
<ul><li>A comment at the site, explaining the trade-off and the condition that would trigger the fix.</li><li>A ticket, linked from the comment.</li><li>A named condition: "when we have more than five customers on this path" is much better than "later."</li></ul>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python"># DEBT: hardcoded to the US tax table to ship for the March launch.
# Parameterize when we take our first non-US customer. See ENG-4417.
TAX_RATE = 0.0725</code></pre></div>
<p>That comment costs thirty seconds and it is the difference between debt and mystery.</p>
<h2 id="2-accidental-debt">2. accidental debt<a class="anchor" href="#2-accidental-debt" aria-label="link to this section">#</a></h2>
<p>You did not know better at the time. The requirements changed. The library you chose turned out to be wrong. The abstraction fit the problem you had and not the one you have now.</p>
<p>This is the largest category and it is not anyone's fault. It is the natural consequence of building things under uncertainty.</p>
<p><strong>The response is refactoring as part of ordinary work</strong>, not as a project.</p>
<p>The practice that works: when you touch a file, leave it slightly better. Not a rewrite — rename the confusing variable, extract the function that is doing two things, delete the dead branch. Small, continuous, in the same commit as the feature.</p>
<p>Refactoring projects — a quarter dedicated to cleanup — mostly fail. They are unfunded after the first month, they conflict with in-flight work, and they produce large risky changes with no user-visible benefit to justify them.</p>
<p>Continuous small improvement compounds and is nearly invisible in the process. That is the whole trick.</p>
<h2 id="3-structural-debt">3. structural debt<a class="anchor" href="#3-structural-debt" aria-label="link to this section">#</a></h2>
<p>The architecture is wrong for what the system now does.</p>
<p>The monolith needs to be split, or the microservices need to be merged. The data model does not represent the domain. The synchronous design cannot support the scale. The framework choice from 2018 is blocking everything.</p>
<p>This is the expensive category and it is qualitatively different from the other two, because <strong>you cannot fix it incrementally without a plan.</strong> Small improvements to a wrong structure make the wrong structure more entrenched.</p>
<p><strong>The response is a real project, with a real justification, sized honestly.</strong></p>
<p>That requires:</p>
<ul><li><strong>A specific cost statement.</strong> Not "the architecture is bad." "Every new feature in this area takes three times as long as an equivalent feature elsewhere, and here are the last four examples with dates."</li><li><strong>A specific benefit.</strong> What becomes possible or fast afterward.</li><li><strong>A path with intermediate value.</strong> A twelve-month rewrite with no deliverable until month twelve will be cancelled in month seven. Structure it so each phase ships something.</li><li><strong>A strangler pattern, not a rewrite.</strong> Build the new alongside the old, migrate incrementally, delete the old. Big-bang rewrites have a well-documented failure rate and the reason is always the same: the old system's behavior includes a decade of undocumented edge cases nobody enumerated.</li></ul>
<h2 id="how-to-talk-about-it">how to talk about it<a class="anchor" href="#how-to-talk-about-it" aria-label="link to this section">#</a></h2>
<p>The conversation that fails: "we need to spend time on technical debt."</p>
<p>The conversation that works: "the last four features in the billing area each took about three weeks. Equivalent features elsewhere take one. The difference is the data model, specifically that subscriptions and invoices share a table. Fixing it is about six weeks and would bring billing features back to normal velocity. We have eleven billing features on the roadmap."</p>
<p>The second version has: a measurement, a cause, a cost, and a payback period. It is a business case, and business cases get funded.</p>
<p>The first version is a complaint, and complaints do not.</p>
<h2 id="the-thing-nobody-says">the thing nobody says<a class="anchor" href="#the-thing-nobody-says" aria-label="link to this section">#</a></h2>
<p>Some debt should never be paid.</p>
<p>Code in a system that will be retired, or that nobody has changed in three years, or that works and has no pending requirements — leave it. Ugly code that is stable and unmodified costs you nothing. It is not debt if you never pay interest on it.</p>
<p>The question is never "is this code good." It is "is this code costing us anything." A lot of what gets called technical debt is code that offends someone's taste in a module nobody touches, and refactoring it is a hobby, not engineering.</p>]]></content:encoded></item><item><title>The retrieval question, revisited</title><link>https://readme.news/the-retrieval-question-revisited/</link><guid isPermaLink="true">https://readme.news/the-retrieval-question-revisited/</guid><pubDate>Fri, 17 Apr 2026 09:00:00 +0000</pubDate><description>Context windows kept growing and retrieval did not die. It changed shape. Where the line actually sits now.</description><content:encoded><![CDATA[<p>Two years ago the argument was that growing context windows would make retrieval pipelines obsolete. I made a version of that argument and was partly wrong.</p>
<p>Here is where the line actually is now, with the reasoning rather than the conclusion, because the line keeps moving and the reasoning does not.</p>
<h2 id="the-four-variables">the four variables<a class="anchor" href="#the-four-variables" aria-label="link to this section">#</a></h2>
<p>Whether to retrieve or to stuff the context depends on four things:</p>
<p><strong>Corpus size.</strong> If everything fits, stuff it. "Fits" now means a large document set — a codebase, a product's full documentation, a year of one team's tickets. It does not mean an enterprise's entire document store.</p>
<p><strong>Query pattern.</strong> If you ask many questions against the same corpus, <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> the prefix makes stuffing cheap. If every query touches a different slice of a large corpus, retrieval wins.</p>
<p><strong>Latency budget.</strong> Prefill on a very large context takes real time even when cached. If you need sub-second responses, a small retrieved context is faster.</p>
<p><strong>Citation requirement.</strong> If you must show which source supported a claim, retrieval gives you that structurally. Stuffing requires trusting the model's attribution, which is less reliable than people assume.</p>
<h2 id="the-decision-concretely">the decision, concretely<a class="anchor" href="#the-decision-concretely" aria-label="link to this section">#</a></h2>
<p><strong>Stuff the context when:</strong> the corpus is under a few hundred thousand tokens, you query it repeatedly (so caching applies), and you want the model to see relationships between distant parts.</p>
<p>The clearest example remains a codebase. Retrieval over code performs worse than whole-repository context because code's meaning lives in the relationships between files, and chunking destroys exactly that.</p>
<p><strong>Retrieve when:</strong> the corpus is large, queries are diverse, latency matters, or you need citations.</p>
<p><strong>Do both when:</strong> you have a large corpus with a hot subset. Stuff the hot subset — the style guide, the schema, the core documents — and retrieve from the tail.</p>
<h2 id="what-actually-changed-retrieval-became-a-tool">what actually changed: retrieval became a tool<a class="anchor" href="#what-actually-changed-retrieval-became-a-tool" aria-label="link to this section">#</a></h2>
<p>The important architectural shift is not about size. It is that retrieval moved from a <strong>preprocessing step</strong> to a <strong>tool the model calls</strong>.</p>
<p>Old shape:</p>
<div class="code"><pre><code>query → embed → search → rerank → stuff top-k → generate</code></pre></div>
<p>One retrieval, before generation, with <code>k</code> fixed by you.</p>
<p>New shape:</p>
<div class="code"><pre><code>query → model reasons → calls search tool → reads results
      → reasons → searches again with a better query → reads
      → generates answer with citations</code></pre></div>
<p>The model decides what to search for, evaluates whether the results answered the question, and searches again with a refined query if not.</p>
<p>This is dramatically better and it is better for a specific reason: <strong>the first query is usually not the right query.</strong> A user asks about "the timeout issue." The right search is for the specific component's retry configuration, which you only know to search for after reading something else.</p>
<p>A fixed one-shot retrieval cannot do that. An agent with a search tool can.</p>
<h2 id="what-this-means-for-your-pipeline">what this means for your pipeline<a class="anchor" href="#what-this-means-for-your-pipeline" aria-label="link to this section">#</a></h2>
<p><strong>Embeddings matter less.</strong> When the model can iterate, a mediocre first retrieval is recoverable. When you had one shot, embedding quality was everything.</p>
<p><strong>Keyword search came back.</strong> Hybrid search — BM25 plus vectors — consistently outperforms pure vector search, and for a model that can iterate, plain keyword search is often sufficient. Exact terms matter: error codes, function names, product names. Embeddings are bad at exact match and always were.</p>
<p>If you built a pure-vector pipeline in 2023, adding keyword search is probably your biggest available quality improvement.</p>
<p><strong>Chunking matters less, and differently.</strong> Instead of chunking for retrieval, store documents whole and retrieve whole documents when they fit. Chunk only what is too large, and chunk on structural boundaries — sections, functions, headings — rather than by token count.</p>
<p><strong>Metadata filtering matters more.</strong> The model can specify constraints: this project, this date range, this author. Filtering is cheap, precise, and it is frequently what the query actually needed. Make sure your index supports it.</p>
<h2 id="the-practical-setup">the practical setup<a class="anchor" href="#the-practical-setup" aria-label="link to this section">#</a></h2>
<p>For most applications:</p>
<ol><li><strong>Postgres with <code>pgvector</code> plus full-text search.</strong> One system, hybrid search, metadata filters, and joins to your relational data.</li><li><strong>Expose search as a tool</strong>, not as a preprocessing step. Let the model iterate.</li><li><strong>Return whole documents</strong> where they fit; chunk on structure where they do not.</li><li><strong>Cache the stable prefix</strong> — the instructions, the schema, the core reference material.</li><li><strong>Measure retrieval quality separately from answer quality.</strong> If the answer is wrong, you need to know whether the right document was retrieved. Most teams cannot answer that and debug blind.</li></ol>
<p>That last one is the single most useful piece of instrumentation in a retrieval system and almost nobody has it.</p>
<h2 id="the-thing-i-got-wrong">the thing I got wrong<a class="anchor" href="#the-thing-i-got-wrong" aria-label="link to this section">#</a></h2>
<p>I said the RAG infrastructure category was solving a temporary problem. The capability got absorbed, as predicted. The infrastructure repositioned rather than disappearing, which I did not predict.</p>
<p>That is the third time I have watched this exact pattern and failed to apply the lesson. Infrastructure around a model limitation rarely dies. It moves to whatever the model still cannot do, which is usually one layer out.</p>]]></content:encoded></item><item><title>The API versioning decision</title><link>https://readme.news/the-api-versioning-decision/</link><guid isPermaLink="true">https://readme.news/the-api-versioning-decision/</guid><pubDate>Wed, 15 Apr 2026 09:00:00 +0000</pubDate><description>URL versions, header versions, or no versions. Each is a bet on how your consumers behave, and one of them is usually wrong.</description><content:encoded><![CDATA[<p>Every API eventually needs to change in a way that breaks someone. How you handle that is a decision made once, early, that you live with for the life of the product.</p>
<h2 id="the-options">the options<a class="anchor" href="#the-options" aria-label="link to this section">#</a></h2>
<p><strong>URL versioning.</strong> <code>/v1/users</code>, <code>/v2/users</code>.</p>
<p>Explicit, visible, cacheable, easy to route. Anyone can see which version a request targets by reading the URL, which matters more than it sounds — for debugging, for logs, for support conversations.</p>
<p>The cost: it encourages big-bang versions. <code>/v2</code> implies everything changed, so teams batch breaking changes into a major version, and consumers face a large migration instead of several small ones.</p>
<p><strong>Header versioning.</strong> <code>Accept: application/vnd.example.v2+json</code> or a custom header.</p>
<p>Cleaner URLs, more RESTful in a purist sense, allows finer granularity.</p>
<p>The cost: invisible. You cannot paste a URL and know what it returns. <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">Caching</a> requires <code>Vary</code> handling that intermediaries get wrong. Debugging is harder. In practice most teams that choose this regret the invisibility.</p>
<p><strong>Date-based versioning.</strong> <code>API-Version: 2026-04-15</code>.</p>
<p>The consumer pins to a date. Every breaking change gets a new date. Consumers upgrade by changing the date and reading the <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a> for the intervening period.</p>
<p>This is what several of the better-run APIs do, and it is my recommendation. The advantage is that breaking changes can be small and frequent rather than large and rare, and each one is individually documented.</p>
<p>The cost: you must maintain compatibility shims for every version you support, which is real engineering work and requires discipline about how long you support them.</p>
<p><strong>No versioning: only additive changes.</strong> Never break anything. Add fields, never remove or change them. Add endpoints, never change existing ones.</p>
<p>This works, is the least work for consumers, and is genuinely achievable for many APIs. The cost is accumulating cruft — deprecated fields that must be populated forever, endpoints nobody should use but that cannot be removed.</p>
<h2 id="what-actually-counts-as-breaking">what actually counts as breaking<a class="anchor" href="#what-actually-counts-as-breaking" aria-label="link to this section">#</a></h2>
<p>Less obvious than it looks. Breaking:</p>
<ul><li>Removing a field or an endpoint.</li><li>Renaming anything.</li><li>Changing a type — even integer to string for an ID, which people do.</li><li>Adding a required request field.</li><li>Changing validation to be stricter.</li><li>Changing an error code for an existing condition.</li><li>Changing default values.</li><li>Changing pagination behavior.</li><li>Changing the order of an array that consumers might depend on.</li></ul>
<p>Non-breaking, usually:</p>
<ul><li>Adding an optional request field.</li><li>Adding a response field. <strong>Usually.</strong> A consumer with strict schema validation that rejects unknown fields will break, which is why your documentation should state explicitly that clients must tolerate unknown fields.</li></ul>
<p>That last one is worth putting in writing early. "Clients MUST ignore unrecognized fields" in your API documentation, from day one, converts a whole category of future breaking change into a non-breaking one.</p>
<h2 id="the-deprecation-process-that-works">the deprecation process that works<a class="anchor" href="#the-deprecation-process-that-works" aria-label="link to this section">#</a></h2>
<ol><li><strong>Announce</strong>, with a date, in the changelog and in the response headers.</li></ol>
<div class="code"><pre><code>Deprecation: Sat, 15 Apr 2026 00:00:00 GMT
Sunset: Wed, 15 Oct 2026 00:00:00 GMT
Link: &lt;https://docs.example.com/migrate/v2&gt;; rel="deprecation"</code></pre></div>
<ol><li><strong>Measure who is still using it.</strong> You need per-consumer usage metrics on the</li></ol>
<p>deprecated path or you are guessing. This is the step everyone skips and it is what makes the rest possible.</p>
<ol><li><strong>Contact the remaining users directly.</strong> Not a blog post. An email to the</li></ol>
<p>specific integrations still calling it.</p>
<ol><li><strong>Brownouts.</strong> Before the sunset, disable the endpoint for short windows —</li></ol>
<p>an hour, then a day — with clear <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a>. This surfaces the consumers who ignored every notice, while they still have time to fix it.</p>
<ol><li><strong>Sunset</strong>, with a clear error explaining what to do.</li></ol>
<p>The brownout step is the one that separates deprecations that work from deprecations that get postponed four times.</p>
<h2 id="the-internal-api-exception">the internal API exception<a class="anchor" href="#the-internal-api-exception" aria-label="link to this section">#</a></h2>
<p>For an API consumed only inside your organization, versioning is usually over-engineering. You can find every consumer and change them.</p>
<p>Use a shared schema, break things when needed, coordinate the change, move on. The versioning machinery exists because you cannot coordinate with strangers, and if you can coordinate, skip it.</p>
<p>The failure mode is treating an internal API as external — building versioning, deprecation policies, and compatibility shims for three consumers you could have just updated.</p>
<h2 id="the-recommendation">the recommendation<a class="anchor" href="#the-recommendation" aria-label="link to this section">#</a></h2>
<p>Date-based versioning for a public API, with a documented support window and a changelog that lists every change with its date.</p>
<p>Additive-only for an internal API, with a shared schema and the discipline to coordinate when you must break something.</p>
<p>And in both cases: state in your documentation, prominently, that clients must ignore unknown fields. It costs nothing now and saves a version bump later.</p>]]></content:encoded></item><item><title>You still do not need Kubernetes</title><link>https://readme.news/you-still-do-not-need-kubernetes/</link><guid isPermaLink="true">https://readme.news/you-still-do-not-need-kubernetes/</guid><pubDate>Mon, 13 Apr 2026 09:00:00 +0000</pubDate><description>It&#x27;s excellent software solving a real problem that most teams do not have. The honest threshold, and what to do below it.</description><content:encoded><![CDATA[<p>Kubernetes is genuinely good software. It solves a real problem well. It has an enormous ecosystem and a large pool of people who know it.</p>
<p>It is also, for a majority of the teams running it, a substantial amount of complexity in exchange for benefits they do not receive, and saying so is still mildly heretical.</p>
<h2 id="the-problem-it-actually-solves">the problem it actually solves<a class="anchor" href="#the-problem-it-actually-solves" aria-label="link to this section">#</a></h2>
<p>Kubernetes was built for: many services, many teams, heterogeneous workloads, on a fleet of machines, where you want bin-packing efficiency and declarative self-healing, and where the platform is operated by people whose job that is.</p>
<p>If you have all of those, it is the right answer and there is no close second.</p>
<h2 id="what-it-costs">what it costs<a class="anchor" href="#what-it-costs" aria-label="link to this section">#</a></h2>
<p><strong>A permanent learning tax.</strong> Pods, deployments, services, ingresses, configmaps, secrets, persistent volume claims, storage classes, service accounts, roles, network policies, resource quotas, and a YAML dialect for each. Every engineer who deploys anything must learn a meaningful fraction of it.</p>
<p><strong>Operational surface.</strong> Control plane upgrades, node upgrades, CNI plugin, CSI driver, ingress controller, cert manager, metrics server, log shipper. Each is a component that can break and that must be upgraded on someone else's schedule.</p>
<p><strong>Debugging distance.</strong> "Why is my service not reachable" has a dozen possible answers across five layers, and diagnosing it requires understanding all of them.</p>
<p><strong>Cost, frequently.</strong> A managed control plane plus nodes sized for the platform's own overhead plus the observability stack it needs is often more than the equivalent capacity on simpler infrastructure.</p>
<p><strong>Resume-driven adoption.</strong> This is real and worth naming. Kubernetes on your CV is worth money. That is a genuine incentive pointed away from the simplest solution that works.</p>
<h2 id="the-honest-threshold">the honest threshold<a class="anchor" href="#the-honest-threshold" aria-label="link to this section">#</a></h2>
<p>You probably want Kubernetes if:</p>
<ul><li>More than roughly fifteen to twenty distinct services, deployed independently.</li><li>More than a handful of teams that need to deploy without coordinating.</li><li>You have someone whose job includes operating the platform, not as a side task.</li><li>Genuinely heterogeneous workloads with different scaling characteristics.</li><li>Multi-tenancy requirements with real isolation needs.</li></ul>
<p>You probably do not if:</p>
<ul><li>Under ten services.</li><li>One or two teams.</li><li>Nobody owns the platform.</li><li>Traffic is predictable.</li><li>You are running one application with a database.</li></ul>
<h2 id="what-to-do-below-the-threshold">what to do below the threshold<a class="anchor" href="#what-to-do-below-the-threshold" aria-label="link to this section">#</a></h2>
<p>The options are better than they were, and all of them are boring:</p>
<p><strong>A platform-as-a-service.</strong> Push code, it runs. This is the correct answer for a very large number of applications and the reason people avoid it is usually aesthetic.</p>
<p><strong>Containers on a managed container service</strong> without the orchestrator — the various "run this container, scale it, load balance it" products every cloud offers. You get containers, autoscaling, and rolling deploys without the platform.</p>
<p><strong>A couple of servers and a process manager.</strong> Systemd units, a reverse proxy with automatic certificates, and a deploy script. This runs an enormous amount of traffic, is trivially debuggable, and every engineer already understands it.</p>
<p><strong>Docker Compose on one machine.</strong> For staging, for internal tools, for anything where a single host is enough. Unfashionable, works.</p>
<h2 id="the-migration-path-argument">the migration-path argument<a class="anchor" href="#the-migration-path-argument" aria-label="link to this section">#</a></h2>
<p>"We will need Kubernetes eventually, so we should start now."</p>
<p>This is the most common justification and it is usually wrong, for two reasons.</p>
<p>First, the complexity cost is paid every day from now until then, and "eventually" frequently never arrives.</p>
<p>Second, containerizing your application is the actual hard part of a future migration, and you can do that without an orchestrator. A containerized application running under a simple process manager can move to Kubernetes later in a couple of weeks.</p>
<p>Build the container. Skip the platform until you have the problem it solves.</p>
<h2 id="the-position-i-will-defend">the position I will defend<a class="anchor" href="#the-position-i-will-defend" aria-label="link to this section">#</a></h2>
<p>The teams I have seen most successfully run Kubernetes are large organizations with dedicated <a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">platform teams</a>, where it is genuinely the right tool.</p>
<p>The teams I have seen most damaged by it are small ones where a single engineer set it up, that engineer left, and the remaining team is operating a system nobody understands and is afraid to touch.</p>
<p>That second failure is common, expensive, and entirely predictable from the staffing at adoption time. If nobody's job is going to be operating the platform, you should not have a platform.</p>]]></content:encoded></item><item><title>Rate limiting: the four algorithms and when each is wrong</title><link>https://readme.news/rate-limiting-the-four-algorithms-and-when-each-is-wrong/</link><guid isPermaLink="true">https://readme.news/rate-limiting-the-four-algorithms-and-when-each-is-wrong/</guid><pubDate>Mon, 06 Apr 2026 09:00:00 +0000</pubDate><description>Token bucket, leaky bucket, fixed window, sliding window. They fail differently and the differences matter.</description><content:encoded><![CDATA[<p>Rate limiting looks like a solved problem until you have to pick an algorithm, at which point the choice determines your failure mode.</p>
<h2 id="fixed-window">fixed window<a class="anchor" href="#fixed-window" aria-label="link to this section">#</a></h2>
<p>Count requests per fixed interval. Reset the counter at the boundary.</p>
<div class="code"><pre><code>key: user:1234:minute:2026-04-06T14:23
incr, expire 60s, reject if &gt; limit</code></pre></div>
<p><strong>Pros:</strong> trivial to implement, one counter, minimal memory.</p>
<p><strong>Cons:</strong> the boundary problem. A client can send the full limit at 14:23:59 and the full limit again at 14:24:00 — double the intended rate, in one second, legally.</p>
<p><strong>Use when:</strong> the limit is generous relative to the burst you can absorb, and simplicity matters more than precision. Most internal services.</p>
<h2 id="sliding-window-log">sliding window log<a class="anchor" href="#sliding-window-log" aria-label="link to this section">#</a></h2>
<p>Store a timestamp per request. Count the ones inside the window.</p>
<p><strong>Pros:</strong> exactly correct. No boundary artifacts.</p>
<p><strong>Cons:</strong> memory proportional to the limit times the number of clients. A limit of 10,000 per hour per user across a million users is a lot of timestamps.</p>
<p><strong>Use when:</strong> limits are small, precision matters, and client count is bounded. API keys with strict quotas.</p>
<h2 id="sliding-window-counter">sliding window counter<a class="anchor" href="#sliding-window-counter" aria-label="link to this section">#</a></h2>
<p>The practical compromise. Keep counters for the current and previous window, interpolate based on how far into the current window you are.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def allowed(now, limit, window):
    cur_start = now - (now % window)
    elapsed = (now - cur_start) / window
    estimate = prev_count * (1 - elapsed) + cur_count
    return estimate &lt; limit</code></pre></div>
<p><strong>Pros:</strong> two counters per client, no boundary problem in practice, cheap.</p>
<p><strong>Cons:</strong> an approximation. Can be slightly wrong at the edges under a very uneven request distribution.</p>
<p><strong>Use when:</strong> this is the default. It is what most production rate limiters actually do and it is almost always the right choice.</p>
<h2 id="token-bucket">token bucket<a class="anchor" href="#token-bucket" aria-label="link to this section">#</a></h2>
<p>A bucket holds tokens, refilled at a constant rate up to a maximum. Each request consumes one. Empty bucket, request rejected.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def allowed(bucket, now, rate, capacity):
    bucket.tokens = min(capacity, bucket.tokens + (now - bucket.last) * rate)
    bucket.last = now
    if bucket.tokens &gt;= 1:
        bucket.tokens -= 1
        return True
    return False</code></pre></div>
<p><strong>Pros:</strong> allows bursts up to the bucket capacity while enforcing a long-run average. This matches how real clients behave — idle, then a burst of activity — much better than a strict window.</p>
<p><strong>Cons:</strong> two parameters instead of one, and people set them without thinking about what the burst allowance means.</p>
<p><strong>Use when:</strong> you want to allow legitimate bursts. Most user-facing APIs. This is the second-best default after sliding window counter and is better when burstiness is expected.</p>
<h2 id="leaky-bucket">leaky bucket<a class="anchor" href="#leaky-bucket" aria-label="link to this section">#</a></h2>
<p>A queue drained at a constant rate. Requests enter the queue; overflow is rejected.</p>
<p><strong>Pros:</strong> output rate is perfectly smooth, which is what you want when you are protecting a downstream system that cannot handle bursts at all.</p>
<p><strong>Cons:</strong> adds latency — requests wait in the queue. Not appropriate for interactive traffic.</p>
<p><strong>Use when:</strong> you are shaping traffic to a fixed-capacity downstream, like a third-party API with a hard rate limit, or a legacy system that falls over above a threshold.</p>
<h2 id="the-parts-everyone-gets-wrong">the parts everyone gets wrong<a class="anchor" href="#the-parts-everyone-gets-wrong" aria-label="link to this section">#</a></h2>
<p><strong>Rate limiting the wrong key.</strong> Limiting by IP breaks for users behind NAT and does nothing against a distributed attacker. Limit by authenticated identity where you have one, and by IP only as a coarse pre-auth defense.</p>
<p><strong>Not telling the client anything useful.</strong> A 429 with no information is hostile. Send the standard headers:</p>
<div class="code"><pre><code>RateLimit-Limit: 100
RateLimit-Remaining: 0
RateLimit-Reset: 42
Retry-After: 42</code></pre></div>
<p>A client that knows when to retry will retry then. A client that does not will hammer you.</p>
<p><strong>No jitter on the client side.</strong> If a thousand clients are told to retry in 42 seconds, they all retry at exactly the same moment. Always jitter, and say so in your API documentation.</p>
<p><strong>Rate limiting at the wrong layer.</strong> Limiting at the application means the request already consumed a connection, a thread, and possibly a database query. For abuse protection, limit at the edge. For fairness between legitimate users, limit in the application where you know who they are.</p>
<p><strong>Ignoring cost variance.</strong> Not all requests are equal. A search that scans a million rows and a <a class="xref" href="/your-monitoring-is-measuring-the-wrong-nines/" title="Your monitoring is measuring the wrong nines">health check</a> are both "one request." Weight by cost, or rate-limit expensive endpoints separately, or you will limit the wrong thing.</p>
<p><strong>Global counters in a distributed system.</strong> Exact global rate limiting requires coordination on every request, which is a latency and availability problem. The usual answer is per-node limits with the total divided across nodes, accepting some imprecision, or a shared store with local <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> and periodic reconciliation.</p>
<p>Decide which imprecision you can live with, deliberately, rather than discovering it during an incident.</p>
<h2 id="the-one-line-recommendation">the one-line recommendation<a class="anchor" href="#the-one-line-recommendation" aria-label="link to this section">#</a></h2>
<p>Sliding window counter for fairness, token bucket where bursts are legitimate, limit by authenticated identity, always send <code>Retry-After</code>, always jitter.</p>]]></content:encoded></item><item><title>The database you should have chosen</title><link>https://readme.news/the-database-you-should-have-chosen/</link><guid isPermaLink="true">https://readme.news/the-database-you-should-have-chosen/</guid><pubDate>Mon, 30 Mar 2026 09:00:00 +0000</pubDate><description>Postgres. That&#x27;s the article. Here&#x27;s the longer version, including the cases where it&#x27;s wrong.</description><content:encoded><![CDATA[<p>For a new application, the default database choice is Postgres, and the burden of proof is on anything else.</p>
<p>This is not a controversial position anymore, which is itself notable, because a decade ago it was.</p>
<h2 id="what-it-absorbed">what it absorbed<a class="anchor" href="#what-it-absorbed" aria-label="link to this section">#</a></h2>
<p>The reason the argument ended is that Postgres kept adding the things people left it for:</p>
<p><strong>JSON.</strong> <code>jsonb</code> with indexing, operators, and path queries. The document database use case — schemaless-ish data with nested structure — is handled well enough that the specialized option is rarely worth a second system.</p>
<p><strong>Full-text search.</strong> Built in, with ranking, stemming, and index support. Not as good as a dedicated search engine at large scale or with complex relevance requirements. Good enough for the search box on most applications, and one fewer system to operate.</p>
<p><strong>Vector search.</strong> <code>pgvector</code> with HNSW indexing. Again: not the best available at extreme scale, and entirely adequate for most retrieval workloads, with the enormous advantage that your vectors sit next to your metadata and you can filter and join in one query.</p>
<p>That last point is underrated. A dedicated vector database that cannot join to your relational data means you do the join in application code, badly.</p>
<p><strong>Time series.</strong> Partitioning, BRIN indexes, and extensions cover a lot of the use case.</p>
<p><strong>Geospatial.</strong> PostGIS remains the best geospatial implementation in any database, period.</p>
<p><strong><a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">Queues</a>.</strong> <code>SELECT ... FOR UPDATE SKIP LOCKED</code> gives you a correct, transactional job queue in about twenty lines. For anything under high-thousands of jobs per second — which is most systems — you do not need a dedicated queue, and the transactional property (enqueue in the same transaction as the state change) removes an entire class of consistency bug.</p>
<h2 id="the-operational-reality">the operational reality<a class="anchor" href="#the-operational-reality" aria-label="link to this section">#</a></h2>
<p><strong>One system to operate.</strong> Every additional datastore is a backup strategy, a monitoring setup, an upgrade path, a failure mode, an <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> runbook, and a set of consistency questions between it and everything else.</p>
<p>The cost of a second datastore is not the license. It is the permanent operational surface, and teams consistently underestimate it by a large factor.</p>
<p><strong>Transactions across your data.</strong> If your users and your documents and your embeddings are in one database, a change to all three is one transaction. Split across three systems, it is a distributed consistency problem you now own.</p>
<h2 id="when-it-is-genuinely-wrong">when it is genuinely wrong<a class="anchor" href="#when-it-is-genuinely-wrong" aria-label="link to this section">#</a></h2>
<p>Being honest about this matters, because "just use Postgres" as a reflex is the same failure mode as any other reflex.</p>
<p><strong>Extreme write throughput on a single logical dataset.</strong> Postgres scales writes vertically, well, up to a point. Past that point you need sharding, and Postgres's sharding story is real but less mature than systems designed for it from the start. If you genuinely need millions of writes per second, look elsewhere.</p>
<p><strong>Analytical queries over very large datasets.</strong> Postgres is a row store. Scanning a billion rows to compute an aggregate is what column stores are for, and the difference is orders of magnitude. Use a column store — several integrate cleanly with Postgres.</p>
<p><strong>Global multi-region with low write latency everywhere.</strong> Postgres replication is primary-based. If you need writes accepted in multiple regions with low latency, that is a distributed database problem and Postgres is not one.</p>
<p><strong>Extremely high-volume <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a>.</strong> Redis exists and is better at being a cache. This is a legitimate second system.</p>
<p><strong>Embedded / <a class="xref" href="/pixel-10-and-the-on-device-model-as-a-platform-feature/" title="Pixel 10 and the on-device model as a platform feature">on-device</a>.</strong> SQLite. Different problem, different answer.</p>
<h2 id="the-thing-to-actually-do">the thing to actually do<a class="anchor" href="#the-thing-to-actually-do" aria-label="link to this section">#</a></h2>
<p>Start with Postgres. Put everything in it. Measure.</p>
<p>When something is genuinely the bottleneck — and you will know, because you will have the metrics — extract that one thing to a specialized system with a clear reason.</p>
<p>That order matters. Teams that start with five specialized systems on the theory that they will need them end up operating five systems for a workload one would have handled, and the complexity is permanent.</p>
<h2 id="the-version-note">the version note<a class="anchor" href="#the-version-note" aria-label="link to this section">#</a></h2>
<p>Whatever you are running, be less than two major versions behind. Postgres's release quality is high, the upgrades are usually boring, and the performance improvements between versions are substantial and free.</p>
<p>The people who have bad Postgres upgrade experiences are the ones who skipped four versions and tried to do it in one jump during an outage.</p>]]></content:encoded></item>
</channel>
</rss>
