<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — craft</title>
<link>https://readme.news/tags/craft/</link>
<atom:link href="https://readme.news/tags/craft/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged craft.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>The dashboard nobody looks at</title><link>https://readme.news/the-dashboard-nobody-looks-at/</link><guid isPermaLink="true">https://readme.news/the-dashboard-nobody-looks-at/</guid><pubDate>Mon, 28 Sep 2026 09:00:00 +0000</pubDate><description>Most dashboards are built during an incident and never opened again. A few are worth keeping, and they look different.</description><content:encoded><![CDATA[<p>Open your monitoring tool and count the dashboards. Now check when each was last viewed — most tools will tell you.</p>
<p>The distribution is always the same: three dashboards carry all the traffic and forty have not been opened in a year.</p>
<h2 id="why-they-accumulate">why they accumulate<a class="anchor" href="#why-they-accumulate" aria-label="link to this section">#</a></h2>
<p>Dashboards get created during incidents. Someone needs to see a specific correlation right now, builds a view, the incident ends, and the dashboard stays — specific to a problem that has been fixed, named something like "orders debug 2," owned by nobody.</p>
<p>Nothing removes them, so they accumulate until finding the useful one requires knowing its name in advance.</p>
<h2 id="the-three-that-earn-their-place">the three that earn their place<a class="anchor" href="#the-three-that-earn-their-place" aria-label="link to this section">#</a></h2>
<p><strong>The service dashboard.</strong> One per service, same layout for every service, so that anyone can read any service's dashboard without orientation. Rate, errors, duration, saturation. Six to eight panels, above the fold, no scrolling.</p>
<p>The consistency is the point. When every service looks the same, an engineer responding to an unfamiliar service already knows where to look.</p>
<p><strong>The journey dashboard.</strong> One per critical user journey — checkout, signup, search. Not per service: per outcome, end to end, across every service involved.</p>
<p>This is the one most organisations lack, and it is the one that answers the question that matters during an incident: <em>are users able to do the thing?</em> Every service being green while checkout is broken is a common and confusing state, and only a journey view shows it.</p>
<p><strong>The capacity dashboard.</strong> Resource headroom over weeks, not minutes. Reviewed on a schedule, not during incidents. This is the only one meant for planning rather than response.</p>
<h2 id="what-makes-a-panel-useful">what makes a panel useful<a class="anchor" href="#what-makes-a-panel-useful" aria-label="link to this section">#</a></h2>
<p><strong>A threshold drawn on it.</strong> A line with no reference is a shape. A line with "target: 200ms" drawn across it is information. Every latency and error panel should show what "fine" is.</p>
<p><strong>Comparison to last week.</strong> Most values are meaningless in isolation. 4,000 requests per minute — is that high? The same panel with last week's line overlaid answers instantly.</p>
<p><strong>A unit and a direction.</strong> Label it. State whether up is good. This sounds trivial and half of all panels fail it.</p>
<p><strong>Percentiles, not averages,</strong> for anything latency-shaped. And show p50 next to p99: the gap between them is often the real signal.</p>
<p><strong>Fewer panels.</strong> A dashboard with forty panels is a data dump. During an incident nobody scrolls. Eight panels that fit on a screen beat forty that do not.</p>
<h2 id="the-annotations-that-make-it-work">the annotations that make it work<a class="anchor" href="#the-annotations-that-make-it-work" aria-label="link to this section">#</a></h2>
<p>The highest-value dashboard feature is the least used: <strong>deploy markers</strong>.</p>
<p>Vertical lines on every time series showing when a deploy happened. Most tools support it and most teams have not wired it up.</p>
<p>With them, "did this start after the deploy?" is answered by looking. Without them it is answered by opening another tab, finding the deploy log, and correlating timestamps by hand — during an incident, under pressure.</p>
<p>Add config changes and <a class="xref" href="/the-graph-nobody-drew/" title="The graph nobody drew">feature flag</a> flips to the same annotation stream and you have covered the causes of most self-inflicted incidents.</p>
<h2 id="the-maintenance-rule">the maintenance rule<a class="anchor" href="#the-maintenance-rule" aria-label="link to this section">#</a></h2>
<p><strong>Delete dashboards nobody opens.</strong> Quarterly, using the view counts your tool already records. If someone misses one, they can rebuild it, and the rebuild will be better because it will reflect the current system.</p>
<p><strong>Give every dashboard an owner and a purpose in its description.</strong> "Used by <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> to triage checkout failures" tells the next person whether to trust it. An unlabelled dashboard is one nobody dares delete and nobody quite believes.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>During your next incident, notice which dashboard you actually opened first.</p>
<p>That is your real service dashboard. Whether it is the one you designated as such, and how much of what you needed was on it, is the most honest review of your monitoring you will ever get — and it costs nothing but paying attention for five minutes while you are already there.</p>]]></content:encoded></item><item><title>Least privilege, actually applied</title><link>https://readme.news/least-privilege-actually-applied/</link><guid isPermaLink="true">https://readme.news/least-privilege-actually-applied/</guid><pubDate>Fri, 25 Sep 2026 09:00:00 +0000</pubDate><description>Everyone agrees with the principle. Almost nobody has checked what their service can actually reach.</description><content:encoded><![CDATA[<p>Least privilege is one of those principles nobody argues with and few implement, because implementing it means doing the tedious work of finding out what permissions are actually used.</p>
<p>Here is the tedious work, in a manageable order.</p>
<h2 id="the-audit-in-one-afternoon">the audit, in one afternoon<a class="anchor" href="#the-audit-in-one-afternoon" aria-label="link to this section">#</a></h2>
<p>For each service, answer four questions. Write the answers down.</p>
<p><strong>1. What identity does it run as?</strong> Not "a service account" — which one. A surprising number of services run as something shared, or as a human's credentials that were used for the initial deploy and never replaced.</p>
<p><strong>2. What can that identity do?</strong> Enumerate the actual permissions. In most cloud platforms this is one API call. The answer is usually much broader than anyone expects, because permissions accumulate and nothing removes them.</p>
<p><strong>3. What does it actually use?</strong> This is the gap. Cloud providers log every authorised call — CloudTrail, audit logs, equivalents elsewhere. Ninety days of logs tell you precisely which permissions were exercised.</p>
<p><strong>4. What is the delta?</strong> Everything granted and never used is a candidate for removal, and that set is typically most of the grant.</p>
<p>That fourth number is the finding. In every audit I have seen, a service uses somewhere between 5% and 20% of what it is permitted to do.</p>
<h2 id="the-specific-over-grants-to-look-for">the specific over-grants to look for<a class="anchor" href="#the-specific-over-grants-to-look-for" aria-label="link to this section">#</a></h2>
<p><strong>Wildcards.</strong> <code>s3:*</code> on <code>*</code>. Almost always the result of "it wasn't working so I broadened it until it did," which is a debugging technique that becomes a permanent security posture.</p>
<p><strong>Write access where only read is used.</strong> Extremely common for anything reading configuration or reference data.</p>
<p><strong>Access to everything of a type.</strong> A service that needs one bucket having access to all buckets. Scope to the resource, and where possible to a prefix within it.</p>
<p><strong>Permissions for a feature that was removed.</strong> The code went; the grant stayed.</p>
<p><strong>Human roles used by machines.</strong> A deploy pipeline running as a role designed for an engineer's console access — which usually includes the ability to change permissions, and that is the one that turns a compromise into a takeover.</p>
<p><strong><code>iam:*</code> or its equivalent.</strong> The permission to grant permissions. Anything with this is effectively an administrator regardless of what else it has. Treat it as its own category and audit it separately.</p>
<h2 id="the-direction-to-move">the direction to move<a class="anchor" href="#the-direction-to-move" aria-label="link to this section">#</a></h2>
<p><strong>Short-lived over long-lived.</strong> Workload identity — the container proves what it is and receives a token that expires in minutes — removes the stealable credential entirely. This is the single highest-value change on this list, and it is architectural rather than a matter of tightening a policy.</p>
<p><strong>Scoped over broad.</strong> One bucket and one prefix, not the account.</p>
<p><strong>Read over write.</strong> Split the identity if the same service does both, so the read path cannot write.</p>
<p><strong>Deny at the boundary.</strong> A service in a network that cannot reach the internet cannot exfiltrate, whatever its IAM permissions say. Network policy and identity policy are independent layers and both matter.</p>
<h2 id="the-thing-that-makes-it-stick">the thing that makes it stick<a class="anchor" href="#the-thing-that-makes-it-stick" aria-label="link to this section">#</a></h2>
<p>An audit is a snapshot. Permissions creep back the next time something does not work at 6 p.m. on a Friday.</p>
<p>Two mechanisms hold the line:</p>
<p><strong>Permissions in code, reviewed like code.</strong> If a grant requires a pull request, the broad one gets a comment. If it is a console click, it does not.</p>
<p><strong>A scheduled review.</strong> Quarterly, using the same used-versus-granted delta. Twenty minutes per service, and it catches both the emergency grant nobody reverted and the feature that was deleted.</p>
<h2 id="the-honest-framing-for-a-sceptical-audience">the honest framing for a sceptical audience<a class="anchor" href="#the-honest-framing-for-a-sceptical-audience" aria-label="link to this section">#</a></h2>
<p>Least privilege does not prevent compromise. It bounds what a compromise costs.</p>
<p>That is the argument to make when someone asks why it is worth the effort: the question is not whether a credential will leak — a dependency, a laptop, a log file, a misconfigured bucket, eventually one will. The question is whether the answer to "what could they do with it" is "read one bucket" or "anything."</p>
<p>Those two answers are the difference between an incident report and a breach notification, and the work that separates them is an afternoon per service.</p>]]></content:encoded></item><item><title>Reading an RFC</title><link>https://readme.news/reading-an-rfc/</link><guid isPermaLink="true">https://readme.news/reading-an-rfc/</guid><pubDate>Wed, 23 Sep 2026 09:00:00 +0000</pubDate><description>Specifications look impenetrable and follow strict conventions. Knowing the conventions makes them the best documentation available.</description><content:encoded><![CDATA[<p>Most developers have never read a specification end to end. They search for the bit they need, misread it, and implement something that works against one server.</p>
<p>RFCs are dense but they are not difficult, and they follow conventions that make them scannable once you know them.</p>
<h2 id="the-keywords-are-load-bearing">the keywords are load-bearing<a class="anchor" href="#the-keywords-are-load-bearing" aria-label="link to this section">#</a></h2>
<p>RFC 2119 defines a small vocabulary, and in a specification these words are technical terms, not English:</p>
<ul><li><strong>MUST</strong> / <strong>REQUIRED</strong> / <strong>SHALL</strong> — absolute requirement. Violating it makes your implementation non-conforming.</li><li><strong>MUST NOT</strong> / <strong>SHALL NOT</strong> — absolute prohibition.</li><li><strong>SHOULD</strong> / <strong>RECOMMENDED</strong> — there may be valid reasons to ignore this, but understand the implications first. In practice: everyone does it, and if you do not, something will break eventually.</li><li><strong>SHOULD NOT</strong> — same, inverted.</li><li><strong>MAY</strong> / <strong>OPTIONAL</strong> — genuinely optional. Critically: <strong>an implementation that does not do it must interoperate with one that does</strong>, and vice versa.</li></ul>
<p>When reading, look for the capitalised words first. They are the actual requirements; everything around them is explanation. Skimming an RFC by jumping between MUSTs is a legitimate and efficient way to read one.</p>
<h2 id="the-structure-is-consistent">the structure is consistent<a class="anchor" href="#the-structure-is-consistent" aria-label="link to this section">#</a></h2>
<p><strong>Abstract</strong> — one paragraph. Read it to decide whether you want this document.</p>
<p><strong>Introduction / Terminology</strong> — read the terminology. Specifications define ordinary-looking words precisely, and misreading one is the most common source of implementation bugs.</p>
<p><strong>The body</strong> — the actual protocol.</p>
<p><strong>Security Considerations</strong> — mandatory in every RFC and consistently the most interesting section. It is where the authors say what they could not fix, what attacks apply, and what implementers get wrong. If you read one section, read this one.</p>
<p><strong>IANA Considerations</strong> — registries. Where you find the canonical list of registered values, which is often the thing you actually needed.</p>
<p><strong>ABNF grammar</strong> — usually an appendix. The precise syntax, in a formal grammar. If you are writing a parser, this is authoritative and the prose is not.</p>
<h2 id="the-things-that-trip-people-up">the things that trip people up<a class="anchor" href="#the-things-that-trip-people-up" aria-label="link to this section">#</a></h2>
<p><strong>Updated and obsoleted.</strong> An RFC is never edited. It is replaced. Always check the header for "Obsoleted by" — implementing an obsolete specification is a common and embarrassing mistake, and HTTP/1.1 alone has been re-specified several times.</p>
<p><strong>Errata.</strong> Published RFCs have errata filed against them. Check them; some are substantive corrections.</p>
<p><strong>Not all RFCs are standards.</strong> The series includes Informational, Experimental, Best Current Practice, Historic, and April Fools jokes. The status is in the header. "It's an RFC" is not the same as "it's a standard."</p>
<p><strong>The grammar wins over the prose.</strong> Where the ABNF and the description disagree, the ABNF is normative. This is the single most useful thing to know when a specification appears ambiguous.</p>
<h2 id="the-reading-strategy">the reading strategy<a class="anchor" href="#the-reading-strategy" aria-label="link to this section">#</a></h2>
<p>For a specification you need to implement:</p>
<ol><li>Abstract, to confirm it is the right document.</li><li>Check the header for obsoletions, then the errata.</li><li>Terminology section, properly.</li><li>Security Considerations — early, not last. It tells you what the hard parts are before you have written anything.</li><li>Skim for MUST and MUST NOT to get the shape of the requirements.</li><li>Then the relevant sections in detail, with the ABNF beside you.</li></ol>
<p>That is maybe ninety minutes for a substantial specification, and it is dramatically cheaper than the alternative, which is discovering the requirements one interoperability bug at a time.</p>
<h2 id="why-bother">why bother<a class="anchor" href="#why-bother" aria-label="link to this section">#</a></h2>
<p>Because the specification is the only source that is actually authoritative. Blog posts, <a class="xref" href="/stack-overflows-traffic-fell-off-a-cliff-and-it-is-not-coming-back/" title="Stack Overflow&#x27;s traffic fell off a cliff and it is not coming back">Stack Overflow</a> answers and library documentation are all someone's interpretation, and interpretations drift.</p>
<p>When two implementations disagree — and they will — the specification is what settles it. Being the person on the team who can read one and say "we are wrong, section 4.2 says the header is case-insensitive" is a genuinely useful thing to be, and it costs one afternoon of learning the conventions.</p>]]></content:encoded></item><item><title>The migration that never finished</title><link>https://readme.news/the-migration-that-never-finished/</link><guid isPermaLink="true">https://readme.news/the-migration-that-never-finished/</guid><pubDate>Mon, 21 Sep 2026 09:00:00 +0000</pubDate><description>Half-completed migrations are the most expensive state a system can be in, and the most common one.</description><content:encoded><![CDATA[<p>Every mature codebase has at least one: the migration that got to 70% and stopped. The new system handles most traffic, the old one handles the awkward remainder, and both are maintained forever.</p>
<p>This is worse than either finishing or never starting, and it is the default outcome unless something prevents it.</p>
<h2 id="why-it-costs-more-than-both">why it costs more than both<a class="anchor" href="#why-it-costs-more-than-both" aria-label="link to this section">#</a></h2>
<p><strong>Two systems to maintain.</strong> Every change lands twice, or lands in one and silently diverges in the other.</p>
<p><strong>Two sets of bugs</strong>, plus a third set caused by the interaction.</p>
<p><strong>Nobody knows which is authoritative.</strong> New engineers ask; the answer is "it depends."</p>
<p><strong>The benefit never arrives.</strong> The reason for the migration — delete the old thing, get the performance, simplify the model — is only realised at 100%. At 70% you have paid the full cost and collected none of the return.</p>
<p><strong>It gets harder over time.</strong> The remaining 30% is the hard 30%: the weird integrations, the customer with the bespoke arrangement, the code nobody understands. And it gets harder as the people who understood the original migration leave.</p>
<h2 id="why-it-happens">why it happens<a class="anchor" href="#why-it-happens" aria-label="link to this section">#</a></h2>
<p>Not laziness. The incentives genuinely point this way:</p>
<p><strong>The easy 70% delivers most of the visible benefit.</strong> The graph goes up, the demo works, the announcement is made. The remaining 30% has no visible reward.</p>
<p><strong>The hard cases are hard for a reason.</strong> They were skipped because someone did not know how to handle them, and that has not changed.</p>
<p><strong>Priorities move.</strong> The migration was urgent in Q1. In Q3 there is a launch.</p>
<p><strong>Nobody owns the finish.</strong> The person who drove it moved on, and completion was never anyone's explicit goal.</p>
<h2 id="the-mechanisms-that-actually-work">the mechanisms that actually work<a class="anchor" href="#the-mechanisms-that-actually-work" aria-label="link to this section">#</a></h2>
<p><strong>Name the deletion, not the migration.</strong> The project is not "migrate to the new pricing service." It is <strong>"delete the old pricing service."</strong> The deliverable is the deletion, and the migration is how you get there. This one rewording changes what people track and what counts as done.</p>
<p><strong>Set a date, publicly, at the start.</strong> Not "by end of year" — a date, in the plan, with the deletion as the milestone. Dates without deletions slip silently; a date attached to "and then this code is gone" is checkable.</p>
<p><strong>Make the old path visibly worse.</strong> Log a warning on every use. Add latency — genuinely, deliberately. Put a banner in the internal tool. The old path being comfortable is why nobody leaves it.</p>
<p><strong>Count the stragglers on a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>.</strong> "Requests still on the legacy path" as a number that goes down, reviewed weekly. Anything not measured stalls at whatever level nobody notices.</p>
<p><strong>Do the hard cases first.</strong> Backwards from the usual instinct and correct. The easy 70% will always be doable; the hard 30% is what determines whether the project is possible at all. Finding out in month one that a case cannot be migrated is a much better outcome than finding out in month nine.</p>
<p><strong>Budget the finish before starting.</strong> If you cannot fund the last 30%, do not start. A migration you cannot complete is worse than the system you have.</p>
<h2 id="the-decision-worth-making-explicitly">the decision worth making explicitly<a class="anchor" href="#the-decision-worth-making-explicitly" aria-label="link to this section">#</a></h2>
<p>Sometimes finishing genuinely is not worth it. The remaining cases are rare, the old path works, and the effort is better spent elsewhere.</p>
<p>That is a legitimate call — but it has to be <em>made</em>, written down, and the state made permanent rather than provisional:</p>
<blockquote><p>The legacy importer stays for CSV uploads from the four enterprise customers using it. New integrations use the API. We are not migrating them; the importer is now a supported, frozen component with an owner, not a migration in progress. Reviewed 2026-09-21, revisit 2028.</p></blockquote>
<p>That is a completely different thing from a stalled migration, even though the code looks identical. One is a decision with an owner. The other is a mess nobody has admitted to.</p>
<p>The cost of the second one is not the code. It is that every engineer who touches that area has to reconstruct which way things are supposed to be going, and nobody can tell them.</p>]]></content:encoded></item><item><title>Choosing an ID</title><link>https://readme.news/choosing-an-id/</link><guid isPermaLink="true">https://readme.news/choosing-an-id/</guid><pubDate>Fri, 18 Sep 2026 09:00:00 +0000</pubDate><description>Integers, UUIDs, ULIDs, or something with meaning. The decision is permanent and it leaks into everything.</description><content:encoded><![CDATA[<p>Primary key type is one of the few decisions that is genuinely hard to reverse. It propagates into every foreign key, every URL, every log line, every external integration, and every client that ever stored one.</p>
<p>Worth twenty minutes up front.</p>
<h2 id="the-options">the options<a class="anchor" href="#the-options" aria-label="link to this section">#</a></h2>
<p><strong>Auto-increment integer.</strong> Compact, fast, perfect index locality, human-readable in logs.</p>
<p>The problems are real. It leaks volume — <code>/orders/48213</code> tells a competitor how many orders you have taken. It requires a round trip to the database to learn the ID, which blocks client-side generation and batching. And it makes merging data from two systems a genuine ordeal, because both start at 1.</p>
<p><strong>UUIDv4.</strong> Random, globally unique, generatable anywhere without coordination, leaks nothing.</p>
<p>The cost is index behaviour. Random values scatter inserts across the entire B-tree, which destroys cache locality, inflates the index, and causes page splits on every insert. On a large, write-heavy table this is a genuine and measurable problem, not a theoretical one.</p>
<p><strong>UUIDv7.</strong> Time-ordered UUID: a millisecond timestamp prefix, then randomness. Keeps global uniqueness and client-side generation, restores insert locality because new rows land at the end of the index.</p>
<p>This is the default I would now recommend for most new systems. It fixes the one serious problem with v4 while keeping everything that made v4 attractive. Postgres has <code>uuidv7()</code> built in; most languages have a library.</p>
<p>The trade: it leaks creation time. Usually fine, occasionally not.</p>
<p><strong>ULID / KSUID and friends.</strong> Same idea as v7 — sortable, time-prefixed — with a more compact text encoding. Fine choices. UUIDv7 has the advantage of being a standard with native database support, which matters more than encoding length.</p>
<p><strong>Natural keys.</strong> Email, ISBN, SKU. Almost always a mistake as a primary key, because "naturally unique and never changes" turns out to be false: people change email addresses, and standards get revised. Use them as unique constraints, not as the identifier everything else points at.</p>
<h2 id="the-two-identifier-pattern">the two-identifier pattern<a class="anchor" href="#the-two-identifier-pattern" aria-label="link to this section">#</a></h2>
<p>Frequently the right answer is not to choose:</p>
<ul><li><strong>Internal key</strong>: <code>bigint</code>, auto-increment. Used for foreign keys and joins. Compact, fast, never exposed.</li><li><strong>External ID</strong>: UUIDv7 or a prefixed string. Used in URLs, APIs, logs, support conversations. Unique index on it.</li></ul>
<p>You get join performance and index locality internally, and no information leakage externally. The cost is one extra column and remembering which is which.</p>
<h2 id="prefix-your-external-ids">prefix your external IDs<a class="anchor" href="#prefix-your-external-ids" aria-label="link to this section">#</a></h2>
<p>If you expose identifiers, prefix them by type:</p>
<div class="code"><pre><code>cus_01J9F3K2M4N5P6Q7R8S9T0V1W2
ord_01J9F3K8X1Y2Z3A4B5C6D7E8F9</code></pre></div>
<p>This is a small thing with an outsized payoff:</p>
<ul><li>A support engineer can tell what an ID refers to without asking.</li><li>Passing a customer ID where an order ID belongs becomes detectable, and can be rejected at the API boundary rather than producing a confusing not-found.</li><li>Logs and error reports become self-describing.</li><li>You can rotate the encoding later without ambiguity.</li></ul>
<p>Several well-run APIs do this and it is consistently one of the things developers say they appreciate about them.</p>
<h2 id="the-practical-rules">the practical rules<a class="anchor" href="#the-practical-rules" aria-label="link to this section">#</a></h2>
<p><strong>Never expose auto-increment integers publicly.</strong> Volume leakage plus trivial enumeration.</p>
<p><strong>Do not use UUIDv4 as a clustered primary key</strong> on a table that will get large and write-heavy. If you already have, v7 for new tables and consider whether the old one needs a migration — usually it does not, but measure the index bloat before deciding.</p>
<p><strong>Store UUIDs as a native <code>uuid</code> type</strong>, not as text. Sixteen bytes versus thirty-six, and correct comparison semantics.</p>
<p><strong>Decide the case and format once</strong>, write it down, and validate at the boundary. Half your system accepting hyphenated lowercase and the other half accepting bare uppercase is a bug that surfaces years later in one integration.</p>
<p><strong>Never reuse an ID.</strong> Ever. Deleted means gone; the identifier is retired with it. Reuse turns a stale reference from a clean not-found into silent corruption pointing at the wrong record.</p>]]></content:encoded></item><item><title>Why your tests are slow</title><link>https://readme.news/why-your-tests-are-slow/</link><guid isPermaLink="true">https://readme.news/why-your-tests-are-slow/</guid><pubDate>Wed, 16 Sep 2026 09:00:00 +0000</pubDate><description>Six causes, ranked by how often they are the real problem, and the fix for each.</description><content:encoded><![CDATA[<p>A test suite that takes twenty minutes is not run locally. A suite that is not run locally is a suite that fails in CI, which means the feedback loop is now measured in pipeline runs.</p>
<p>Here is where the time actually goes, roughly in order of how often each is the dominant cause.</p>
<h2 id="1-everything-talks-to-a-real-database">1. everything talks to a real database<a class="anchor" href="#1-everything-talks-to-a-real-database" aria-label="link to this section">#</a></h2>
<p>The most common cause by a wide margin. Each test sets up a database, inserts fixtures, runs, and tears down.</p>
<p>Fixes, in order of return:</p>
<p><strong>Roll back instead of truncating.</strong> Wrap each test in a transaction and roll it back. Orders of magnitude faster than deleting rows, and it needs no cleanup code.</p>
<p><strong>Share the schema, not the data.</strong> Create the schema once per run, not per test.</p>
<p><strong>Use a tmpfs for the test database.</strong> The data does not need to survive; durable writes are pure cost. On Postgres, a data directory in memory plus <code>fsync=off</code> and <code>synchronous_commit=off</code> is dramatically faster and completely inappropriate for anything but tests.</p>
<p><strong>Move the logic out of the database's reach.</strong> The deepest fix: if business logic is a pure function of its inputs, its tests do not need a database at all. Tests that are hard to write without infrastructure are telling you about your design.</p>
<h2 id="2-the-suite-is-serial">2. the suite is serial<a class="anchor" href="#2-the-suite-is-serial" aria-label="link to this section">#</a></h2>
<p>Most runners parallelise and most projects have not turned it on, usually because it exposed shared state once and someone reverted it.</p>
<p>That shared state is a bug. Tests that pass only in a particular order have an ordering dependency, and ordering dependencies in tests usually mirror a real one in the code.</p>
<p>Turn on parallelism, fix what breaks, and you will find at least one genuine issue.</p>
<h2 id="3-sleeps">3. sleeps<a class="anchor" href="#3-sleeps" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">time.sleep(2)   # wait for the worker to pick it up</code></pre></div>
<p>Every one of these is pure latency, and they are always tuned to the slowest machine anyone has run on. Fifty of them is a hundred seconds of doing nothing.</p>
<p>Replace with polling on the actual condition:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">wait_until(lambda: job.reload().status == "done", timeout=5)</code></pre></div>
<p>Same worst case, typically a hundredth of the elapsed time, and it fails with a useful message instead of a mysterious assertion.</p>
<h2 id="4-the-fixtures-are-enormous">4. the fixtures are enormous<a class="anchor" href="#4-the-fixtures-are-enormous" aria-label="link to this section">#</a></h2>
<p>A test that needs one user builds an organisation, three teams, forty users and a year of history, because it reuses the "standard" fixture.</p>
<p>Build the minimum. Factories with sensible defaults, overridden per test, beat shared fixture files that grow to serve every case.</p>
<h2 id="5-it-is-doing-real-network-io">5. it is doing real network I/O<a class="anchor" href="#5-it-is-doing-real-network-io" aria-label="link to this section">#</a></h2>
<p>Tests hitting real HTTP endpoints — even internal ones, even mock servers over a socket — pay connection setup and scheduling per call.</p>
<p>Intercept at the client layer rather than the network layer. Most languages have a way to stub the HTTP client in-process, and it is both faster and more deterministic than a local server.</p>
<p>The exception is contract tests, which exist to catch exactly what stubs hide. Keep those, run them separately, and do not let them into the fast suite.</p>
<h2 id="6-everything-is-an-end-to-end-test">6. everything is an end-to-end test<a class="anchor" href="#6-everything-is-an-end-to-end-test" aria-label="link to this section">#</a></h2>
<p>Browser-driving tests are two to three orders of magnitude slower than unit tests. A suite made mostly of them is slow no matter what you do.</p>
<p>The ratio that works: a large number of fast unit tests, a moderate number of integration tests around real boundaries, and a small number of end-to-end tests covering the handful of journeys that must never break.</p>
<p>Inverting that ratio is the most expensive testing mistake a team can make, and it is usually the result of not trusting the lower layers rather than a deliberate choice.</p>
<h2 id="the-measurement-first">the measurement first<a class="anchor" href="#the-measurement-first" aria-label="link to this section">#</a></h2>
<p>Before fixing anything, get the distribution:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">pytest --durations=25
go test ./... -json | jq -r 'select(.Action=="pass") | "\(.Elapsed) \(.Test)"' | sort -rn | head -25</code></pre></div>
<p>Almost always, a small number of tests dominate. Fixing the slowest twenty is usually the whole job, and it is an afternoon rather than a project.</p>
<h2 id="the-target">the target<a class="anchor" href="#the-target" aria-label="link to this section">#</a></h2>
<p><strong>Under 10 seconds for the unit suite</strong>, run on every save. Under two minutes for everything that gates a merge.</p>
<p>The ten-second number is not arbitrary — it is roughly the threshold past which people stop running tests as they work and start batching. Once that happens you have lost the feedback loop, and the suite's speed stops being a convenience question and starts being a quality one.</p>]]></content:encoded></item><item><title>The graph nobody drew</title><link>https://readme.news/the-graph-nobody-drew/</link><guid isPermaLink="true">https://readme.news/the-graph-nobody-drew/</guid><pubDate>Mon, 14 Sep 2026 09:00:00 +0000</pubDate><description>Your service dependency graph exists whether or not anyone has written it down. Writing it down takes an afternoon.</description><content:encoded><![CDATA[<p>Ask an engineering team to list what their service depends on to serve a request. You will get the database and two or three obvious APIs.</p>
<p>The real list is usually fifteen things, and the difference between the two lists is where the surprising outages come from.</p>
<h2 id="the-full-inventory">the full inventory<a class="anchor" href="#the-full-inventory" aria-label="link to this section">#</a></h2>
<p>For one service, one request path:</p>
<ul><li><strong>Datastores.</strong> Primary, replicas, cache, search index, object storage.</li><li><strong>Internal services</strong> it calls synchronously.</li><li><strong>Third-party APIs</strong> — payments, email, geocoding, <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">feature flags</a>, auth.</li><li><strong>Infrastructure control planes.</strong> Service discovery, DNS, the secret manager, the config service. These are the ones nobody lists and they are the ones that take everything down.</li><li><strong>Identity.</strong> The token issuer, the JWKS endpoint it fetches keys from.</li><li><strong>Observability.</strong> If your telemetry agent blocks when the collector is down — and some do — it is in your request path.</li><li><strong>The container registry</strong>, at start-up. Fine until you need to scale during an incident and cannot pull an image.</li><li><strong>Certificate infrastructure</strong>, including OCSP and the ACME endpoint.</li><li><strong>NTP.</strong> Rarely, spectacularly.</li></ul>
<p>The pattern: the dependencies people list are the ones they <em>call</em>. The ones that hurt are the ones they <em>need in order to start, authenticate, or scale.</em></p>
<h2 id="the-two-questions-per-dependency">the two questions per dependency<a class="anchor" href="#the-two-questions-per-dependency" aria-label="link to this section">#</a></h2>
<p>For each item, two answers, written down:</p>
<p><strong>What happens when it is slow?</strong> Slow is worse than down and much more common. Down fails fast; slow holds your connections, exhausts your pool, and turns one dependency's bad day into your outage.</p>
<p><strong>What happens when it is unavailable?</strong> Three possible answers, and only one is a decision:</p>
<ul><li><em>We fail.</em> Fine, if the dependency is genuinely essential — a database for a write path.</li><li><em>We degrade.</em> Serve stale cache, hide the feature, use a default. This is the answer for most non-essential dependencies and it requires code that usually does not exist yet.</li><li><em>We do not know.</em> This is the real answer for most dependencies, and it is the finding.</li></ul>
<h2 id="the-classification-that-matters">the classification that matters<a class="anchor" href="#the-classification-that-matters" aria-label="link to this section">#</a></h2>
<p>Sort every dependency into two buckets:</p>
<p><strong>Hard</strong> — the request genuinely cannot be served without it. Should be a very short list. Each one caps your maximum availability at its own.</p>
<p><strong>Soft</strong> — the request can be served in degraded form. Every soft dependency needs a timeout, a fallback, and a circuit breaker, or it is a hard dependency that nobody has admitted to.</p>
<p>The arithmetic is the reason this matters: five hard dependencies at 99.9% each gives you at best 99.5%, before any failure of your own. If your availability target is higher than that, some of those dependencies have to become soft, and that is an engineering project rather than a target you can declare.</p>
<h2 id="the-feature-flag-trap">the feature flag trap<a class="anchor" href="#the-feature-flag-trap" aria-label="link to this section">#</a></h2>
<p>Worth naming specifically because it catches good teams. Feature flag services are adopted as a safety mechanism — a way to turn things off when they go wrong.</p>
<p>Then the flag SDK becomes a hard dependency on the request path, and when the flag service has an outage your service does too. The tool you adopted to make incidents smaller has made one bigger.</p>
<p>The fix is standard for this whole category and rarely applied: cache flag values locally, evaluate from the cache, refresh in the background, and ship a default in the binary. Then the flag service being down means flags are stale, which is survivable.</p>
<h2 id="how-to-draw-it">how to draw it<a class="anchor" href="#how-to-draw-it" aria-label="link to this section">#</a></h2>
<p>Do not buy anything. Open a file:</p>
<div class="code"><pre><code>orders-api
  HARD  postgres-primary        write path; no fallback
  HARD  auth-jwks               cached 1h, so 1h of grace
  SOFT  redis                   cache; miss → primary, 3x latency
  SOFT  inventory-svc           timeout 300ms → show "check availability"
  SOFT  pricing-svc             timeout 200ms → list price, no promo
  SOFT  flags-svc               local cache + baked defaults
  BOOT  vault                   secrets at start; running pods unaffected
  BOOT  registry                image pull; blocks scale-up only</code></pre></div>
<p>Twenty minutes per service. The <code>BOOT</code> category — needed to start but not to serve — is the one that catches people out during incidents, because those dependencies are invisible right up until you need to replace a pod.</p>
<p>Then, for every <code>SOFT</code> line, check that the fallback described actually exists in code. That check is where the real findings are, and it is usually not the timeout that is missing but the fallback behind it.</p>]]></content:encoded></item><item><title>Timeouts: every one of them</title><link>https://readme.news/timeouts-every-one-of-them/</link><guid isPermaLink="true">https://readme.news/timeouts-every-one-of-them/</guid><pubDate>Thu, 10 Sep 2026 09:00:00 +0000</pubDate><description>A request crosses a dozen components with a timeout each, and almost nobody has ever added them up.</description><content:encoded><![CDATA[<p>Every layer of your stack has a timeout. Most of them are defaults. Almost nobody has written them down in one place and checked that they make sense together.</p>
<p>The result is systems where the client gives up at 10 seconds, the server keeps working for 60, and the database holds a lock for 300.</p>
<h2 id="the-inventory">the inventory<a class="anchor" href="#the-inventory" aria-label="link to this section">#</a></h2>
<p>For a single HTTP request, in rough order:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">layer</th><th style="text-align:left">typical default</th></tr></thead><tbody><tr><td style="text-align:left">browser / client library</td><td style="text-align:left">30s or none</td></tr><tr><td style="text-align:left">DNS resolution</td><td style="text-align:left">5s per attempt</td></tr><tr><td style="text-align:left">TCP connect</td><td style="text-align:left">20–75s (OS)</td></tr><tr><td style="text-align:left">TLS handshake</td><td style="text-align:left">inherits connect</td></tr><tr><td style="text-align:left">load balancer idle</td><td style="text-align:left">60s</td></tr><tr><td style="text-align:left">reverse proxy read</td><td style="text-align:left">60s</td></tr><tr><td style="text-align:left">application server request</td><td style="text-align:left">often none</td></tr><tr><td style="text-align:left">HTTP client to downstream</td><td style="text-align:left">often none</td></tr><tr><td style="text-align:left">connection pool acquire</td><td style="text-align:left">30s</td></tr><tr><td style="text-align:left">database statement</td><td style="text-align:left">often none</td></tr><tr><td style="text-align:left">database lock wait</td><td style="text-align:left">often none</td></tr></tbody></table></div>
<p>Two things stand out. <strong>"Often none" appears five times</strong> — most application-level timeouts are unset by default. And <strong>the values are not coordinated</strong> with each other in any way.</p>
<h2 id="the-rule-that-fixes-most-of-it">the rule that fixes most of it<a class="anchor" href="#the-rule-that-fixes-most-of-it" aria-label="link to this section">#</a></h2>
<p><strong>Timeouts must decrease as you go deeper.</strong></p>
<div class="code"><pre><code>client 10s
  └─ load balancer 9s
       └─ application 8s
            └─ downstream call 3s (with 1 retry → 6s worst case)
                 └─ connection acquire 1s
                      └─ database statement 2s</code></pre></div>
<p>Each layer must be shorter than its caller, with room for <a class="xref" href="/retries-a-complete-guide-to-not-making-it-worse/" title="Retries: a complete guide to not making it worse">retries</a>. If an inner layer can outlast its caller, the caller gives up while the work continues — which means you are burning capacity on results nobody will receive. Under load, that is most of your capacity.</p>
<h2 id="deadline-propagation">deadline propagation<a class="anchor" href="#deadline-propagation" aria-label="link to this section">#</a></h2>
<p>The better version of the rule: do not configure each timeout independently. Pass the deadline down.</p>
<div class="code"><span class="code-lang">go</span><pre><code class="lang-go">// caller has 10s; every downstream inherits what remains
ctx, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()

// 3s into the request, this call gets at most 7s, automatically
resp, err := downstream.Fetch(ctx, id)</code></pre></div>
<p>Go's <code>context</code>, gRPC's deadlines and equivalents elsewhere exist for this. Once you propagate deadlines, a slow first step automatically shortens the budget for later ones, and work never continues past the point where anyone is waiting.</p>
<p>This is the single highest-value change available in most service codebases, and it is usually a few days of threading a parameter through.</p>
<h2 id="the-ones-people-forget">the ones people forget<a class="anchor" href="#the-ones-people-forget" aria-label="link to this section">#</a></h2>
<p><strong>Lock wait timeouts.</strong> A transaction waiting on a row lock with no timeout waits forever. Set <code>lock_timeout</code> in Postgres, <code>innodb_lock_wait_timeout</code> in MySQL.</p>
<p><strong>DDL statements.</strong> A migration that cannot get its lock will queue behind a long transaction — and everything else <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> behind it. <code>SET lock_timeout = '3s'</code> before DDL is the line that prevents a large share of migration incidents.</p>
<p><strong>Idle-in-transaction.</strong> A connection that opened a transaction and went away holds locks and blocks vacuum indefinitely. <code>idle_in_transaction_session_timeout</code> is the safety net.</p>
<p><strong>Client-side connect vs. read.</strong> These are different, and most libraries let you set them separately. A short connect timeout with a long read timeout is usually what you want.</p>
<p><strong>DNS.</strong> Rarely configured, occasionally the whole problem.</p>
<h2 id="write-them-down">write them down<a class="anchor" href="#write-them-down" aria-label="link to this section">#</a></h2>
<p>One table, in the repository, listing every timeout in the request path with its current value and where it is configured.</p>
<p>The exercise takes an afternoon and reliably finds at least one inversion — an inner layer waiting longer than the outer one — and at least two places where the value is a framework default nobody chose.</p>
<p>That document then becomes something you can review when latency changes, rather than a set of numbers scattered across six config files and three languages.</p>
<h2 id="the-failure-mode-this-prevents">the failure mode this prevents<a class="anchor" href="#the-failure-mode-this-prevents" aria-label="link to this section">#</a></h2>
<p>Without coordinated timeouts, a slow dependency does not degrade your service — it exhausts it. Requests pile up waiting on something their callers abandoned long ago, connections stay held, the pool empties, and healthy requests start failing for want of a connection.</p>
<p>That is a full outage caused by one slow downstream, and the difference between it and a brief latency blip is entirely whether the numbers were coordinated.</p>]]></content:encoded></item><item><title>The debugging notebook</title><link>https://readme.news/the-debugging-notebook/</link><guid isPermaLink="true">https://readme.news/the-debugging-notebook/</guid><pubDate>Sat, 05 Sep 2026 09:00:00 +0000</pubDate><description>Writing down what you tried is the cheapest debugging technique there is, and almost nobody does it.</description><content:encoded><![CDATA[<p>Two hours into a hard bug, most people cannot remember which of the six things they tried actually changed the behaviour, or whether they have already checked the thing they are about to check again.</p>
<p>The fix is a text file. It costs nothing and it is the difference between an investigation and a random walk.</p>
<h2 id="what-goes-in-it">what goes in it<a class="anchor" href="#what-goes-in-it" aria-label="link to this section">#</a></h2>
<div class="code"><pre><code>14:02  Symptom: 502s on /api/orders, ~4% of requests, started ~13:40.
       Not correlated with deploy (last one 11:20).

14:08  FACT: only POST. GET is clean. (checked: 30min of access logs)
14:15  FACT: all failures have Content-Length &gt; 64KB.
       → hypothesis: body size limit somewhere
14:22  Checked nginx client_max_body_size = 10m. Not it.
14:31  Checked the app's own limit — 100MB. Not it.
14:40  FACT: failures all hit pod-7 and pod-11 (of 12). Others clean.
       → hypothesis abandoned: not size, size just correlates with
         which client sends big bodies, and that client is sticky
14:52  pod-7 and pod-11 started 13:38. Two minutes before symptoms.
       → new hypothesis: bad config in the new pods
15:01  CONFIRMED: those pods have UPSTREAM_TIMEOUT unset → defaulting to 1s</code></pre></div>
<p>An hour of work, ten lines. Note what it captures: facts with how they were checked, hypotheses explicitly raised and explicitly abandoned, and the moment a correlation turned out to be a red herring.</p>
<h2 id="why-it-works">why it works<a class="anchor" href="#why-it-works" aria-label="link to this section">#</a></h2>
<p><strong>It stops you re-checking things.</strong> The single biggest time sink in a long debugging session is doing the same check twice because you cannot remember the result.</p>
<p><strong>It forces hypotheses to be explicit.</strong> Written down, "it's a size limit" is a claim you can test and discard. Held in your head, it quietly steers every subsequent action.</p>
<p><strong>It separates fact from guess.</strong> The most common way a debugging session goes wrong is a guess getting promoted to a fact through repetition. Labelling them differently prevents it.</p>
<p><strong>It makes handoff possible.</strong> When you hit the end of your day, or the end of your knowledge, the notebook is the handoff. Without it the next person starts from zero.</p>
<p><strong>It becomes the postmortem.</strong> The timeline is already written, with the dead ends included — which are the most instructive part and the first thing lost to memory.</p>
<h2 id="the-three-prompts">the three prompts<a class="anchor" href="#the-three-prompts" aria-label="link to this section">#</a></h2>
<p>When stuck, the notebook gives you somewhere to answer these:</p>
<p><strong>"What changed?"</strong> Deploys, config, data volume, dependency versions, traffic shape, time of day. Most bugs in a previously-working system are caused by a change, and enumerating them beats staring at code.</p>
<p><strong>"What do I actually know?"</strong> Read back your own FACT lines. Frequently the answer is that you know much less than you thought, and one of your assumptions was never verified.</p>
<p><strong>"What would prove me wrong?"</strong> Not what would confirm the theory — what would kill it. If you cannot name that, the theory is not yet a hypothesis.</p>
<h2 id="the-format-does-not-matter">the format does not matter<a class="anchor" href="#the-format-does-not-matter" aria-label="link to this section">#</a></h2>
<p>A scratch file, a comment thread on the ticket, a channel where you narrate to yourself. What matters is that it is written, timestamped, and distinguishes what you checked from what you suspect.</p>
<p>The one thing that does not work is keeping it in your head, because the whole point is that your head is where the confusion is.</p>
<h2 id="the-compounding-version">the compounding version<a class="anchor" href="#the-compounding-version" aria-label="link to this section">#</a></h2>
<p>Keep the notebooks. A directory of them, named by date and symptom.</p>
<p>Six months in, you have a searchable record of how this system actually fails, which is knowledge that otherwise exists only as a vague feeling in whoever was on call. When the same class of bug returns — and it does — the search takes a minute and the fix takes five.</p>
<p>That is the cheapest institutional memory available to an engineering team, and it is built one text file at a time.</p>]]></content:encoded></item><item><title>The interface is the product</title><link>https://readme.news/the-interface-is-the-product/</link><guid isPermaLink="true">https://readme.news/the-interface-is-the-product/</guid><pubDate>Mon, 31 Aug 2026 09:00:00 +0000</pubDate><description>Users cannot see your architecture. They can see the six seconds it takes to do the thing they came for.</description><content:encoded><![CDATA[<p>Engineers evaluate software by its internals: the architecture, the correctness, the elegance of the data model. Users evaluate it by the sequence of actions required to get what they came for.</p>
<p>These correlate less than we would like, and the gap is where a lot of otherwise good software fails.</p>
<h2 id="what-users-actually-experience">what users actually experience<a class="anchor" href="#what-users-actually-experience" aria-label="link to this section">#</a></h2>
<p>Not your service boundaries. Not your consistency model. Not the fact that you handle a partition correctly.</p>
<ul><li>How many steps to do the common thing.</li><li>Whether it responds instantly or after a spinner.</li><li>Whether an error tells them what to do.</li><li>Whether it does the same thing twice in a row.</li><li>Whether it remembers what they told it last time.</li></ul>
<p>Every one of those is an interface property, and every one is achievable on top of an ugly implementation — or destroyed on top of a beautiful one.</p>
<h2 id="the-trade-that-is-usually-made-backwards">the trade that is usually made backwards<a class="anchor" href="#the-trade-that-is-usually-made-backwards" aria-label="link to this section">#</a></h2>
<p>There is a real tension between internal cleanliness and external simplicity, and it comes up constantly:</p>
<ul><li>The clean data model exposes three concepts where users think in one.</li><li>The correct API makes the caller specify things they do not care about.</li><li>The properly-separated services mean the UI has to make four calls and handle each failing independently.</li><li>The general solution has eleven configuration options; the specific one would have had zero.</li></ul>
<p>The default resolution is to protect the internals and push the complexity outward, because the internals are what engineers look at and defend in review.</p>
<p>That is backwards. <strong>The interface is used far more times than the implementation is read.</strong> Complexity at the boundary is multiplied by every user and every call; complexity inside is paid by the people who chose it.</p>
<h2 id="what-this-looks-like-in-practice">what this looks like in practice<a class="anchor" href="#what-this-looks-like-in-practice" aria-label="link to this section">#</a></h2>
<p><strong>Collapse concepts at the boundary.</strong> If users think of one thing and your model has three, expose one and do the mapping internally. Yes, that is a lossy abstraction and you will occasionally have to break it. That is the correct place to put the pain.</p>
<p><strong>Default everything.</strong> Every required parameter is a decision you have forced on someone who has less context than you. Make it optional with a sensible default, and let the people who genuinely need control find the option.</p>
<p><strong>Make the common path one step.</strong> Count the actions for the thing 90% of users do 90% of the time. If it is more than two, that is the work.</p>
<p><strong>Absorb the failure.</strong> If your architecture means four things can fail independently, the interface should not surface four independent failures. Retry, degrade, or present one coherent state.</p>
<p><strong>Make it fast where they are waiting.</strong> A local read, an optimistic update, a cached response. Users cannot tell the difference between "fast because it is well-engineered" and "fast because you cheated" — and neither can anyone else.</p>
<h2 id="the-counterweight-honestly">the counterweight, honestly<a class="anchor" href="#the-counterweight-honestly" aria-label="link to this section">#</a></h2>
<p>This is not an argument for shipping a nice facade over a broken system. The internals are what make the interface <em>keep</em> working — under load, at the edges, after a year of changes. Software that is lovely to use and impossible to change dies just as reliably, only slower.</p>
<p>The claim is narrower: <strong>when the two genuinely conflict, and they do, the interface should usually win</strong> — because it is what the software is <em>for</em>, and because interface decisions are much more expensive to reverse than internal ones. You can rewrite the storage layer. You cannot easily take back a concept you taught a million users.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>Watch someone use your software for the first time without helping them.</p>
<p>Count the moments they hesitate. Each one is a place where your model and theirs diverged, and no amount of internal quality closes that gap.</p>
<p>That exercise is uncomfortable, takes twenty minutes, and consistently produces a better backlog than any amount of architectural discussion.</p>]]></content:encoded></item><item><title>The API you ship to yourself</title><link>https://readme.news/the-api-you-ship-to-yourself/</link><guid isPermaLink="true">https://readme.news/the-api-you-ship-to-yourself/</guid><pubDate>Tue, 25 Aug 2026 09:00:00 +0000</pubDate><description>Internal interfaces get none of the care external ones do, and they are the ones you&#x27;ll live with longest.</description><content:encoded><![CDATA[<p>Teams that would never ship a public API without review, versioning and documentation routinely create internal interfaces in a pull request nobody looked at closely, with a name chosen in thirty seconds.</p>
<p>Those internal interfaces then last longer than the public ones, because nothing external forces them to change and nothing internal makes changing them worth the trouble.</p>
<h2 id="why-they-matter-more-than-they-look">why they matter more than they look<a class="anchor" href="#why-they-matter-more-than-they-look" aria-label="link to this section">#</a></h2>
<p><strong>They are the seams your architecture is made of.</strong> A module's public functions <em>are</em> its contract, whether or not you called it a contract.</p>
<p><strong>Everyone reads them.</strong> A public API is read by customers. An internal one is read by every engineer who joins, forever.</p>
<p><strong>They shape what is easy.</strong> An <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> that makes the right thing awkward guarantees people do the wrong thing, and then you have a pattern.</p>
<p><strong>They are hard to change once used.</strong> Not because of compatibility, but because finding and updating every caller is work nobody has budgeted.</p>
<h2 id="what-to-actually-apply">what to actually apply<a class="anchor" href="#what-to-actually-apply" aria-label="link to this section">#</a></h2>
<p>Not the full public-API apparatus. Four things:</p>
<p><strong>1. Name it from the caller's side.</strong></p>
<p>The best test is to write the call before the implementation:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python"># would a caller guess this?
orders.pending_for(customer_id, since=last_sync)

# or this?
OrderService.get_orders_by_customer_with_status_filter(customer_id, "PENDING", last_sync)</code></pre></div>
<p>The first reads like the sentence the caller was thinking. The second reads like the implementation leaked into the name.</p>
<p><strong>2. Make the wrong call hard to write.</strong></p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python"># a caller can get this wrong and never know
def transfer(from_account, to_account, amount): ...
transfer(b, a, 500)   # silently backwards

# a caller cannot
def transfer(*, source: AccountId, dest: AccountId, cents: int): ...</code></pre></div>
<p>Keyword-only arguments, distinct types for things that must not be swapped, and units in the name. Every one of these converts a runtime bug into a compile or call-site error.</p>
<p><strong>3. Return something that cannot be misread.</strong></p>
<p><code>None</code> for "not found" and <code>None</code> for "error" and <code>None</code> for "empty" is three different meanings on one value. A <code>Result</code>, an explicit exception, or distinct return types cost a few lines and remove a category of caller bug.</p>
<p><strong>4. Document the <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a>, not the happy path.</strong></p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def reserve(sku: str, qty: int) -&gt; Reservation:
    """Reserve stock, or raise.

    Raises:
        OutOfStock: qty exceeds available. Callers should offer backorder.
        SkuUnknown: not a real SKU — this is a bug, not a user error.
        LockTimeout: transient, safe to retry with backoff.
    """</code></pre></div>
<p>The signature already said what it takes and returns. What it cannot say is which errors are retryable and which are programmer error — and that is exactly what a caller needs.</p>
<h2 id="the-deprecation-problem">the deprecation problem<a class="anchor" href="#the-deprecation-problem" aria-label="link to this section">#</a></h2>
<p>Internal APIs never get deprecated because there is always something more urgent, and the old one keeps working.</p>
<p>The mechanism that helps is embarrassingly simple: <strong>make the old path noisy in development.</strong></p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def get_orders_by_customer(*args, **kwargs):
    warnings.warn(
        "get_orders_by_customer is replaced by orders.pending_for(); "
        "see ADR-14. Removal targeted 2026-12.",
        DeprecationWarning, stacklevel=2,
    )
    return orders.pending_for(*args, **kwargs)</code></pre></div>
<p>Warnings in test output that name the replacement and a date get acted on. Silent compatibility shims live forever.</p>
<h2 id="the-boundary-that-actually-needs-versioning">the boundary that actually needs versioning<a class="anchor" href="#the-boundary-that-actually-needs-versioning" aria-label="link to this section">#</a></h2>
<p>Most internal interfaces do not need versions — you can find every caller and change them together. That is the entire advantage of being internal and you should use it rather than building machinery to avoid it.</p>
<p>The exception is any interface crossing a <strong>deploy boundary</strong>: service to service, or anything consumed by a client you do not deploy simultaneously. There, old and new run at the same time by definition, and you need the same expand-contract discipline as a schema change — add the new field, support both, migrate callers, remove the old.</p>
<p>The mistake is applying that machinery to an in-process module boundary where a single commit could have updated everything.</p>
<h2 id="the-one-line-version">the one-line version<a class="anchor" href="#the-one-line-version" aria-label="link to this section">#</a></h2>
<p>Design internal interfaces from the call site, make the wrong call unwriteable, document the errors, and change them freely while you still can — because the window where you can find every caller closes quietly, and nobody notices until they need it open.</p>]]></content:encoded></item><item><title>Shell scripts that outlive you</title><link>https://readme.news/shell-scripts-that-outlive-you/</link><guid isPermaLink="true">https://readme.news/shell-scripts-that-outlive-you/</guid><pubDate>Sat, 22 Aug 2026 09:00:00 +0000</pubDate><description>Six lines at the top of a bash script are the difference between a tool and a trap.</description><content:encoded><![CDATA[<p>Every codebase has a <code>scripts/</code> directory. Most of the files in it were written in ten minutes, work correctly on exactly one machine, and fail in ways that produce no error and no output.</p>
<p>Shell is a fine language for gluing programs together. It is a terrible language for doing it <em>safely</em> unless you tell it to be, and telling it to be takes six lines.</p>
<h2 id="the-preamble">the preamble<a class="anchor" href="#the-preamble" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">#!/usr/bin/env bash
set -euo pipefail
IFS=$'\n\t'</code></pre></div>
<p>What each one prevents:</p>
<p><strong><code>set -e</code></strong> — exit on any command that fails. Without it, a script continues merrily after <code>cd /nonexistent</code> and then runs the rest of its commands in the wrong directory. This is how a cleanup script deletes the wrong thing.</p>
<p><strong><code>set -u</code></strong> — error on an undefined variable. Without it, <code>rm -rf "$BUILD_DIR/"</code> with an unset <code>BUILD_DIR</code> expands to <code>rm -rf /</code>. This has happened to real people, at real companies, more than once.</p>
<p><strong><code>set -o pipefail</code></strong> — a pipeline fails if <em>any</em> stage fails, not just the last one. Without it, <code>curl bad-url | jq .</code> succeeds, because <code>jq</code> was happy with the empty input.</p>
<p><strong><code>IFS=$'\n\t'</code></strong> — stop splitting on spaces. This is what makes filenames with spaces stop being a source of bugs.</p>
<p>Four lines. They convert an entire category of silent wrong behaviour into loud failure.</p>
<h2 id="the-next-four-things">the next four things<a class="anchor" href="#the-next-four-things" aria-label="link to this section">#</a></h2>
<p><strong>Quote every expansion.</strong> <code>"$var"</code>, not <code>$var</code>. Always, including inside <code>[[ ]]</code> where it usually does not matter, because "usually" is not a rule anyone remembers correctly.</p>
<p><strong>Use <code>"${var:?message}"</code> for required inputs.</strong></p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">: "${DATABASE_URL:?DATABASE_URL is required}"</code></pre></div>
<p>One line, fails immediately with a useful message rather than three steps later with a confusing one.</p>
<p><strong>Make it idempotent.</strong> A script that is safe to run twice is a script that is safe to run at all. <code>mkdir -p</code>, <code>rm -f</code>, check-before-create. The second run is the one that happens during an incident when nobody is sure whether the first one worked.</p>
<p><strong>Clean up with <code>trap</code>.</strong></p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT</code></pre></div>
<p>Now the temporary directory is removed whether the script succeeds, fails, or is interrupted.</p>
<h2 id="the-usability-part">the usability part<a class="anchor" href="#the-usability-part" aria-label="link to this section">#</a></h2>
<p>A script that other people run needs the same courtesy as any other <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">usage() {
  cat &lt;&lt;'EOF'
Usage: deploy.sh ENVIRONMENT [--dry-run]

  ENVIRONMENT   staging | production
  --dry-run     print what would happen, change nothing

Requires: awscli &gt;= 2, jq. Reads DEPLOY_ROLE from the environment.
EOF
}
[[ $# -eq 0 || "${1:-}" == "-h" ]] &amp;&amp; { usage; exit 0; }</code></pre></div>
<p><strong>Add a dry-run mode to anything destructive.</strong> It costs one conditional and it is the difference between a script people trust and a script people read three times before running.</p>
<p><strong>Echo what you are about to do.</strong> <code>set -x</code> is the crude version and it is better than silence. A script that prints "deleting 4 objects from s3://bucket/prefix/" before doing it lets a human catch the mistake.</p>
<h2 id="when-to-stop-using-shell">when to stop using shell<a class="anchor" href="#when-to-stop-using-shell" aria-label="link to this section">#</a></h2>
<p>Shell is right for: calling other programs in sequence, moving files, gluing a pipeline together. Under about a hundred lines.</p>
<p>Switch to a real language when you need:</p>
<ul><li><strong>Data structures.</strong> Bash arrays are a trap and associative arrays are worse.</li><li><strong>Any arithmetic beyond counting.</strong></li><li><strong>Error handling with recovery</strong>, rather than exit-on-failure.</li><li><strong>Parsing anything structured.</strong> If you are pulling JSON apart with <code>sed</code>, stop.</li><li><strong>Tests.</strong> You can test shell, and almost nobody does, which tells you something.</li></ul>
<p>Python or Go for anything past that line. The rewrite is an hour and it pays back the first time somebody has to change it.</p>
<h2 id="the-check-that-costs-nothing">the check that costs nothing<a class="anchor" href="#the-check-that-costs-nothing" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">shellcheck scripts/*.sh</code></pre></div>
<p>Run it in CI. It catches unquoted expansions, useless <code>cat</code>, subshell variable scoping, and roughly a dozen other things that produce silent wrong behaviour.</p>
<p>Every shell script I have ever run it against had at least one real finding. It takes five minutes to add and it is the highest-value linting available for the least-linted language in most repositories.</p>]]></content:encoded></item><item><title>The cost of a 500</title><link>https://readme.news/the-cost-of-a-500/</link><guid isPermaLink="true">https://readme.news/the-cost-of-a-500/</guid><pubDate>Thu, 20 Aug 2026 09:00:00 +0000</pubDate><description>Error rates get tracked as a percentage. Users experience them as a specific person having a specific bad day.</description><content:encoded><![CDATA[<p>A 0.1% error rate sounds like nothing. It is a rounding error on a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>, it is well inside almost any SLO, and it is the kind of number that gets a green tick in a weekly review.</p>
<p>Run the arithmetic on it and the picture is different.</p>
<h2 id="the-arithmetic">the arithmetic<a class="anchor" href="#the-arithmetic" aria-label="link to this section">#</a></h2>
<p>Ten million requests a day at 0.1% is <strong>ten thousand failed requests</strong>. If a user session averages twenty requests, that is roughly five hundred sessions with at least one failure — every day.</p>
<p>But the failures are not evenly spread, and that is the part that matters. Errors cluster: by endpoint, by tenant, by client version, by region, by input shape. A 0.1% global rate is frequently one customer at 40% and everyone else at 0.001%.</p>
<p>For that customer, your service is broken. Your dashboard says 99.9%.</p>
<h2 id="the-slices-that-reveal-it">the slices that reveal it<a class="anchor" href="#the-slices-that-reveal-it" aria-label="link to this section">#</a></h2>
<p>Aggregate error rate is nearly useless on its own. The same number means completely different things depending on the distribution beneath it, and you cannot tell which without slicing:</p>
<ul><li><strong>By customer or tenant.</strong> The single highest-value slice in any multi-tenant system. One broken integration hides perfectly in a global average.</li><li><strong>By endpoint.</strong> A 0.1% average across forty endpoints can be 4% on the one that matters.</li><li><strong>By client version.</strong> Errors concentrated in one mobile build means a bad release, not a server problem.</li><li><strong>By region.</strong> Often a networking or dependency issue, not application logic.</li><li><strong>By authenticated vs anonymous.</strong> Frequently a different code path entirely.</li></ul>
<p>If your telemetry cannot answer "which customers saw errors in the last hour," that is the gap to close before adding any more dashboards.</p>
<h2 id="the-errors-that-are-not-500s">the errors that are not 500s<a class="anchor" href="#the-errors-that-are-not-500s" aria-label="link to this section">#</a></h2>
<p>Counting HTTP status codes undercounts real failure, sometimes badly:</p>
<ul><li><strong>200 with an empty result</strong> because a downstream timed out and the code swallowed it. The most dangerous category — invisible to every status-based metric.</li><li><strong>200 with a partial result.</strong> The page rendered, three of the eight widgets are missing.</li><li><strong>Client-side failures</strong> that never reach your server at all: a bad bundle, a CSP violation, a network drop mid-request.</li><li><strong>Slow enough to be a failure.</strong> A request that succeeds after 45 seconds has failed as far as the user is concerned, and it is counted as a success.</li><li><strong>The retry that worked.</strong> The user saw a spinner for eight seconds. Your success rate is 100%.</li></ul>
<p><strong>The fix is measuring outcomes rather than responses.</strong> Did the order get placed? Did the message send? Those are answerable from your own data and they count the failures that status codes miss.</p>
<h2 id="what-a-single-failure-actually-costs">what a single failure actually costs<a class="anchor" href="#what-a-single-failure-actually-costs" aria-label="link to this section">#</a></h2>
<p>Worth being concrete, because "0.1%" abstracts it away:</p>
<ul><li>The user <a class="xref" href="/retries-a-complete-guide-to-not-making-it-worse/" title="Retries: a complete guide to not making it worse">retries</a>. If it fails again, some fraction leave.</li><li>Some fraction contact support. Each of those is real money and an engineer's attention.</li><li>Some fraction never come back, and you will not attribute that to today.</li><li>If it is a payment or a submission, someone may not know whether it went through — which generates duplicates, which generates a different class of problem.</li></ul>
<p>None of that appears on the error-rate graph, which is why the graph being green is not the same as things being fine.</p>
<h2 id="the-practical-program">the practical program<a class="anchor" href="#the-practical-program" aria-label="link to this section">#</a></h2>
<p><strong>Alert on the slice, not the aggregate.</strong> "Any single tenant above 5% for ten minutes" catches things the global rate never will.</p>
<p><strong>Sample real failures and read them.</strong> Not the counts — the actual requests, with their inputs and their traces. An hour a week reading real failed requests teaches you more about your system than any dashboard.</p>
<p><strong>Put a request ID on the error page.</strong> So a user's complaint becomes a trace lookup rather than an investigation.</p>
<p><strong>Track outcomes for your critical journeys.</strong> Checkout completed, message delivered, deploy finished. That number is the one to put in the weekly review, not the status-code ratio.</p>
<h2 id="the-framing-that-changes-behaviour">the framing that changes behaviour<a class="anchor" href="#the-framing-that-changes-behaviour" aria-label="link to this section">#</a></h2>
<p>Stop saying "we're at 99.9%."</p>
<p>Start saying "<strong>about five hundred people had a broken session yesterday.</strong>"</p>
<p>Same data. Completely different meeting.</p>]]></content:encoded></item><item><title>How to run a spike</title><link>https://readme.news/how-to-run-a-spike/</link><guid isPermaLink="true">https://readme.news/how-to-run-a-spike/</guid><pubDate>Tue, 18 Aug 2026 09:00:00 +0000</pubDate><description>Timeboxed investigation is the cheapest way to buy certainty, and most teams do it in a way that produces neither certainty nor code.</description><content:encoded><![CDATA[<p>A spike is a timeboxed investigation whose output is a decision, not a feature. Done well it is the highest-leverage day in a project. Done the way most teams do it, it produces a half-finished prototype and an opinion nobody trusts.</p>
<p>The difference is entirely in the setup.</p>
<h2 id="write-the-question-down-first">write the question down first<a class="anchor" href="#write-the-question-down-first" aria-label="link to this section">#</a></h2>
<p>A spike without a written question becomes exploration, and exploration has no end condition.</p>
<p>The question has to be answerable and specific:</p>
<ul><li>Bad: "look into whether we can use X"</li><li>Good: "can X ingest 50k events/second on a single node with our event shape, and what does its failure behaviour look like when the disk fills?"</li></ul>
<p>The second one tells you when you are finished, what to measure, and what would count as a no. The first one runs until someone gets bored.</p>
<h2 id="state-the-decision-it-feeds">state the decision it feeds<a class="anchor" href="#state-the-decision-it-feeds" aria-label="link to this section">#</a></h2>
<p>A spike exists to unblock a decision. Name the decision explicitly: <em>we will use X for the ingest path, or we will stay with Y and shard it.</em></p>
<p>If you cannot name the decision, you are not running a spike, you are learning something — which is fine and should be called that, because it has a different budget.</p>
<h2 id="set-the-box-and-mean-it">set the box, and mean it<a class="anchor" href="#set-the-box-and-mean-it" aria-label="link to this section">#</a></h2>
<p>Two days is the usual right answer. Long enough to get past setup, short enough that being wrong is cheap.</p>
<p>The rule that makes it work: <strong>when the box ends, you stop and report, even if you are nearly there.</strong> "Nearly there" is where spikes go to become three-week projects. If the answer genuinely needs more time, that is itself a finding — report it and ask for a second box with a narrower question.</p>
<h2 id="the-output-is-a-document-not-a-branch">the output is a document, not a branch<a class="anchor" href="#the-output-is-a-document-not-a-branch" aria-label="link to this section">#</a></h2>
<p>This is the part teams skip. The deliverable is one page:</p>
<div class="code"><span class="code-lang">markdown</span><pre><code class="lang-markdown">## Question
Can X ingest 50k events/s on one node with our event shape?

## Answer
Yes for the steady state, no for our burst profile.
32k/s sustained, 61k/s for ~40s before the write buffer backs up.

## Evidence
Harness: spike/x-ingest/ (throwaway). Ran on c7g.4xlarge, 20 min,
production event sample from 2026-08-10. Numbers: spike/x-ingest/results.md

## What surprised me
Backpressure is silent — it drops rather than erroring, and the drop
counter is only in a debug endpoint. That is a production risk with X
regardless of throughput.

## Recommendation
Stay with Y and shard. Revisit X if their 3.x backpressure work lands.

## What I did not test
Multi-node, recovery after disk-full, or the managed offering.</code></pre></div>
<p>Six sections, twenty minutes to write, and it is still useful in a year when somebody asks why the decision went the way it did. The "what I did not test" section is the one that keeps the document honest.</p>
<h2 id="throw-the-code-away">throw the code away<a class="anchor" href="#throw-the-code-away" aria-label="link to this section">#</a></h2>
<p>Spike code is written to answer a question, not to be maintained. It has no tests, no error handling, hard-coded credentials and a <code>main</code> function that does everything.</p>
<p>That is correct — the speed comes from those omissions. The failure is letting it become the foundation of the real implementation, at which point you have shipped a prototype and will spend a year discovering what it does not handle.</p>
<p>Put it in a directory named <code>spike/</code>, or a branch you never merge. Say in the document where it is. Then write the real thing properly, informed by what you learned.</p>
<p>The strongest signal that a team's spikes are working: the spike code never ships.</p>
<h2 id="when-a-spike-is-the-wrong-tool">when a spike is the wrong tool<a class="anchor" href="#when-a-spike-is-the-wrong-tool" aria-label="link to this section">#</a></h2>
<ul><li><strong>When the answer is already written down.</strong> Read the documentation, the <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a>, the issue tracker. A surprising amount of "we need to try it" is answerable in an hour of reading.</li><li><strong>When the question is about production behaviour under real load.</strong> A spike cannot tell you that. A canary can.</li><li><strong>When the decision is reversible and cheap.</strong> Then just pick one. A spike to choose between two libraries you could swap in a day costs more than being wrong.</li></ul>
<p>Spikes are for expensive, hard-to-reverse decisions where the deciding information does not exist yet. That is a narrower set than it feels like, and it is exactly where two days is a bargain.</p>]]></content:encoded></item><item><title>Boring technology, revisited</title><link>https://readme.news/boring-technology-revisited/</link><guid isPermaLink="true">https://readme.news/boring-technology-revisited/</guid><pubDate>Sat, 15 Aug 2026 09:00:00 +0000</pubDate><description>The innovation-token argument is a decade old and still right, with one amendment nobody makes.</description><content:encoded><![CDATA[<p>The "choose boring technology" argument goes like this: you have a small number of innovation tokens. Spend them on the thing that is genuinely your problem. Everything else should be the option so well-understood that its failure modes are documented by strangers.</p>
<p>It has held up for a decade. It is worth restating because a decade is long enough that people have started applying it wrongly, in a specific and predictable way.</p>
<h2 id="why-it-works">why it works<a class="anchor" href="#why-it-works" aria-label="link to this section">#</a></h2>
<p><strong>Known failure modes.</strong> The value of a boring technology is not that it is good. It is that when it breaks at 3 a.m., the <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error message</a> has been pasted into a public forum by someone who then explained the fix.</p>
<p><strong>A hiring pool that already knows it.</strong> Every unusual choice is a thing you must teach every new engineer, forever.</p>
<p><strong>Operational surface you can reason about.</strong> Boring things have runbooks, monitoring integrations, backup tooling, and a decade of accumulated operational wisdom.</p>
<p><strong>Someone else finds the bugs.</strong> A widely-deployed system has been run at scales you will never reach, by people who reported what broke.</p>
<h2 id="the-amendment">the amendment<a class="anchor" href="#the-amendment" aria-label="link to this section">#</a></h2>
<p>Here is what the original argument does not say, and what a decade has taught: <strong>boring is a property of a technology at a moment in time, and it moves in both directions.</strong></p>
<p>Things become boring. Things also stop being boring — a project loses its maintainers, its community fragments, its corporate sponsor reorganises, the ecosystem moves and it does not.</p>
<p>So "choose boring" is not a decision you make once. It is a property you have to re-check, and the failure mode nobody plans for is the technology that was the safe choice in 2019 and is now a liability that everyone forgot to reconsider.</p>
<p><strong>The practical version:</strong> once a year, for each load-bearing dependency, ask — is this still boring? Is it still maintained, still hired for, still the thing a new engineer would expect? A "yes" costs a minute. A "no" is the most valuable thing you will learn that quarter.</p>
<h2 id="how-to-tell-whether-something-is-boring-now">how to tell whether something is boring <em>now</em><a class="anchor" href="#how-to-tell-whether-something-is-boring-now" aria-label="link to this section">#</a></h2>
<p>Not by age. Some old things are unmaintained; some five-year-old things are completely settled.</p>
<p>The signals that actually matter:</p>
<ul><li><strong>Multiple independent maintainers</strong>, ideally from more than one employer.</li><li><strong>A release in the last six months</strong>, and a <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a> that shows maintenance rather than churn.</li><li><strong>You can hire for it.</strong> Search jobs, not GitHub stars.</li><li><strong>Operational documentation written by users</strong>, not just by the vendor.</li><li><strong>A migration path off it</strong>, documented by people who took it. A technology nobody has successfully left is a technology you cannot leave.</li><li><strong>Boring failure modes.</strong> Can you name what happens when it runs out of memory, loses its network, or gets a corrupt config? If nobody has written that down, it is not boring yet.</li></ul>
<h2 id="where-to-spend-the-tokens">where to spend the tokens<a class="anchor" href="#where-to-spend-the-tokens" aria-label="link to this section">#</a></h2>
<p>The argument's real content is that innovation tokens are scarce, and the scarcity is not about technology risk. It is about <strong>attention</strong>.</p>
<p>Every unusual choice consumes attention: to learn it, to operate it, to debug it, to explain it. Attention is the actual constraint on an engineering team, and it is far more limited than the budget.</p>
<p>So spend the tokens where the unusual choice is the product. If your differentiator is a query engine, be adventurous about storage and utterly boring about everything else. If your differentiator is a workflow, use the most conventional stack in existence and put all the attention into the workflow.</p>
<p>The teams that get this wrong are almost never wrong about one big choice. They are wrong about six small ones, each individually defensible, that together consumed all the attention that should have gone into the thing customers pay for.</p>
<h2 id="the-counterweight">the counterweight<a class="anchor" href="#the-counterweight" aria-label="link to this section">#</a></h2>
<p>Taken too far this becomes an argument for never learning anything, and that is its own failure. A team that made every choice in 2016 and never revisited one is not disciplined, it is stuck — and the technology it chose has probably stopped being boring in the meantime.</p>
<p>The synthesis: be adventurous in one place at a time, deliberately, with a written reason. Be boring everywhere else. Re-check annually.</p>
<p>That is a harder discipline than either "always use the new thing" or "never use the new thing," and it is the only one that survives a decade.</p>]]></content:encoded></item><item><title>The service you should not have written</title><link>https://readme.news/the-service-you-should-not-have-written/</link><guid isPermaLink="true">https://readme.news/the-service-you-should-not-have-written/</guid><pubDate>Thu, 13 Aug 2026 09:00:00 +0000</pubDate><description>A checklist for the moment before you create a new deployable, when saying no is still free.</description><content:encoded><![CDATA[<p>Creating a new service feels like progress. It has a clean repository, no legacy, a fresh CI pipeline, and none of the compromises of the thing it is splitting away from.</p>
<p>It is also a permanent commitment that somebody will still be paying for in eight years. Here is the conversation worth having first.</p>
<h2 id="what-a-service-costs-in-full">what a service costs, in full<a class="anchor" href="#what-a-service-costs-in-full" aria-label="link to this section">#</a></h2>
<p>Not the code. The code is the cheap part.</p>
<ul><li>A repository, a build, a deploy pipeline, a rollback path.</li><li>A place to run, sized, with capacity headroom.</li><li>Monitoring, alerting, dashboards, an <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> runbook.</li><li>Secrets, credentials, and their rotation.</li><li>Dependency upgrades and security patching, forever.</li><li>A place in the request path that can now time out, and the circuit breaker and retry policy that implies.</li><li>Documentation, and an owner who is still at the company.</li><li>One more thing a new engineer must learn about.</li></ul>
<p>That list is the same whether the service is four hundred lines or forty thousand. The overhead is fixed, which is why small services are the ones whose economics are worst.</p>
<h2 id="the-questions-in-order">the questions, in order<a class="anchor" href="#the-questions-in-order" aria-label="link to this section">#</a></h2>
<p><strong>1. Could this be a module?</strong></p>
<p>The default answer for new functionality is a well-bounded module in something that already exists. It gets all of the above for free.</p>
<p>The follow-up that matters: <em>if this were a module with an enforced boundary, what would we lose?</em> If the honest answer is "nothing, it would just feel less tidy," write the module.</p>
<p><strong>2. What is the deployment argument?</strong></p>
<p>The strongest reason to extract is that a team needs to release on its own cadence and currently cannot. That is real and it scales with organisation size.</p>
<p>But check it: are they actually blocked, or do they merely deploy together? Two teams that could deploy independently and choose not to have a process problem, not an architecture problem, and a new service will not fix it.</p>
<p><strong>3. What is the scaling argument, with numbers?</strong></p>
<p>"It might need to scale differently" is not an argument. "This component is CPU-bound and spiky while the rest is I/O-bound and steady, and we are currently provisioning for the peak of both" is.</p>
<p>If you cannot state the resource profile that differs, there is no scaling argument.</p>
<p><strong>4. Where does the data live?</strong></p>
<p>The question that sinks most extractions. If the new service needs the same tables the old one does, you have not split anything — you have created two writers to one database, which is worse than one writer, and you now have to coordinate migrations across two deploy cycles.</p>
<p>A service that does not own its data is a distributed monolith with extra latency.</p>
<p><strong>5. What happens when it is down?</strong></p>
<p>If the answer is "the main flow breaks," you have added a failure mode and bought nothing in return. Fault isolation is not free; it is the result of deliberate <a class="xref" href="/timeouts-every-one-of-them/" title="Timeouts: every one of them">timeouts</a>, fallbacks and degradation, and if you are going to build those anyway you could have built them around a module.</p>
<p><strong>6. Who owns it in two years?</strong></p>
<p>Name the team. Not the person — the team. Services outlive their authors, and an ownerless service is the one that runs three major versions behind until it is a security incident.</p>
<h2 id="the-cases-where-the-answer-is-yes">the cases where the answer is yes<a class="anchor" href="#the-cases-where-the-answer-is-yes" aria-label="link to this section">#</a></h2>
<p>Being fair, because the reflex against splitting can be as unexamined as the reflex toward it.</p>
<ul><li>A genuinely different resource profile, with numbers.</li><li>A compliance or data-residency boundary that must be enforced structurally.</li><li>A component that has to be written in a different language for a real reason.</li><li>A team that is demonstrably blocked on someone else's release cadence.</li><li>Something with wildly different availability requirements — a batch job that can be down for an hour next to a checkout path that cannot.</li></ul>
<p>In every one of those, the service is buying something specific that a module cannot.</p>
<h2 id="the-reversibility-test">the reversibility test<a class="anchor" href="#the-reversibility-test" aria-label="link to this section">#</a></h2>
<p>Before you create it, ask: <strong>if this turns out to be wrong, how do we merge it back?</strong></p>
<p>If the answer is "we would not, we would just live with it," you are making a one-way decision on a hunch. Those deserve more scrutiny than they usually get, and the moment before the repository exists is the last time the scrutiny is free.</p>]]></content:encoded></item><item><title>Reading a flame graph</title><link>https://readme.news/reading-a-flame-graph/</link><guid isPermaLink="true">https://readme.news/reading-a-flame-graph/</guid><pubDate>Tue, 11 Aug 2026 09:00:00 +0000</pubDate><description>The single most useful performance visualisation, and the four shapes worth recognising in one.</description><content:encoded><![CDATA[<p>A flame graph answers one question extremely well: <em>where is the time going?</em> Most people who look at one have never been told how to read it, so they squint at a colourful pile of rectangles and conclude that profiling is hard.</p>
<p>It is not. There are four rules and four shapes.</p>
<h2 id="the-four-rules">the four rules<a class="anchor" href="#the-four-rules" aria-label="link to this section">#</a></h2>
<p><strong>The x-axis is not time.</strong> This is the rule everyone gets wrong. Left-to-right is alphabetical or arbitrary — it is <em>not</em> chronological. A frame on the left did not happen before a frame on the right.</p>
<p><strong>Width is total time.</strong> A frame's width is the proportion of samples that included it. Wide means expensive. That is the entire message.</p>
<p><strong>Height is stack depth.</strong> A frame sitting on another means it was called by it. Tall is not bad; tall just means deep call stacks.</p>
<p><strong>Colour is usually meaningless.</strong> In most tools it is random, chosen to make adjacent frames distinguishable. Do not read anything into it unless the tool explicitly says otherwise — some use it for language or module.</p>
<p>That is it. Wide is expensive, stacked means called-by, colour is decoration.</p>
<h2 id="the-four-shapes">the four shapes<a class="anchor" href="#the-four-shapes" aria-label="link to this section">#</a></h2>
<p><strong>A wide plateau near the top.</strong> One function, doing real work, dominating. This is the good case: a clear, single hotspot. Optimise that function or call it less.</p>
<p><strong>A wide plateau near the bottom, narrowing above.</strong> The time is spread across many children. There is no single hotspot; the cost is the whole subtree. Look one level up — usually the answer is "call this subtree fewer times" rather than "make some leaf faster".</p>
<p><strong>A staircase.</strong> Deep, narrow, repeating structure. Often recursion, sometimes a framework's middleware chain. Each frame is cheap; the depth is the cost. Look for a way to shorten the chain rather than to optimise any frame in it.</p>
<p><strong>Many thin spikes with nothing dominant.</strong> Death by a thousand cuts, or your profile is too short. Check the sample count first — a profile of two hundred samples looks like this regardless of the workload. If the sample count is healthy and it still looks like this, the program is genuinely uniform and your wins are architectural rather than local.</p>
<h2 id="the-practical-workflow">the practical workflow<a class="anchor" href="#the-practical-workflow" aria-label="link to this section">#</a></h2>
<p><strong>Profile the thing you care about, under load that resembles production.</strong> A profile of a cold process running a synthetic benchmark measures start-up and your benchmark harness.</p>
<p><strong>Profile long enough.</strong> Seconds, not milliseconds. You are sampling; you need samples.</p>
<p><strong>Look at the widest frame you did not expect.</strong> Not the widest frame — the widest <em>surprising</em> one. Everyone's profile is dominated by something obvious and irreducible. The win is in the frame that has no business being 8% of your runtime.</p>
<p><strong>Check for time you cannot see.</strong> A flame graph of CPU samples shows CPU. If your program spends most of its wall clock waiting on a socket, an on-CPU profile will be nearly empty and entirely misleading. That is what off-CPU profiling and tracing are for, and confusing the two is the most common way people conclude a profiler is lying to them.</p>
<h2 id="differential-flame-graphs">differential flame graphs<a class="anchor" href="#differential-flame-graphs" aria-label="link to this section">#</a></h2>
<p>Underused and excellent. Profile before, profile after, render the difference — frames that got wider are coloured one way, narrower the other.</p>
<p>This turns "did my change help?" from an argument about noisy averages into a picture. It is also the fastest way to find a performance regression introduced between two releases: profile both, diff, look at what turned red.</p>
<h2 id="the-honest-caveat">the honest caveat<a class="anchor" href="#the-honest-caveat" aria-label="link to this section">#</a></h2>
<p>A flame graph tells you where time went. It does not tell you whether that time was necessary, and it will happily show you a perfectly optimised hot loop that should not have been called at all.</p>
<p>The largest performance wins are almost always "stop doing this work" rather than "do this work faster," and no profiler can suggest that. It can only show you where to point the question.</p>]]></content:encoded></item><item><title>What "done" actually means</title><link>https://readme.news/what-done-actually-means/</link><guid isPermaLink="true">https://readme.news/what-done-actually-means/</guid><pubDate>Sat, 08 Aug 2026 09:00:00 +0000</pubDate><description>Every team has a definition of done and most of them are a checklist nobody reads. Here is the version that changes behaviour.</description><content:encoded><![CDATA[<p>Ask five engineers on the same team whether a ticket is done and you will get three answers: the code is written, the code is merged, the code is in production. All three are defensible and they differ by days.</p>
<p>That gap is where most delivery confusion lives, and it is cheap to close.</p>
<h2 id="the-honest-definition">the honest definition<a class="anchor" href="#the-honest-definition" aria-label="link to this section">#</a></h2>
<p>A change is done when <strong>the outcome it was meant to produce is observable, and the team would find out if it stopped.</strong></p>
<p>Not "the code is merged." Merged code that is not deployed is inventory. Deployed code that nobody looked at is a hypothesis.</p>
<p>That framing sounds demanding and mostly is not, because for a large share of work the observation is trivial — the feature is behind a flag, the flag is on for the team, someone clicked it. The point is that <em>someone did</em>.</p>
<h2 id="the-checklist-that-actually-helps">the checklist that actually helps<a class="anchor" href="#the-checklist-that-actually-helps" aria-label="link to this section">#</a></h2>
<p>Most definition-of-done checklists fail because they list activities rather than properties. "Tests written" is an activity. "The change is covered by a test that would fail if it regressed" is a property, and it is checkable.</p>
<p>The version I have seen work is short:</p>
<ul><li><strong>The change does what the ticket asked</strong>, verified by someone other than the author — a test, a reviewer, or the person who asked for it.</li><li><strong>It is in production</strong>, or in a deliberate queue with a named release date.</li><li><strong>A regression would be caught</strong> by a test, a metric, or an alert. If none of those apply, that is a decision, not an oversight.</li><li><strong>Its failure mode is known.</strong> Someone can say what happens when it breaks.</li><li><strong>It is documented where the next person will look</strong> — which is usually the code, sometimes the runbook, rarely the wiki.</li><li><strong>The flag, the branch, and the dead code are cleaned up</strong>, or ticketed with a date.</li></ul>
<p>Six lines. If a change satisfies them, it is done in a way that survives the author going on holiday.</p>
<h2 id="the-two-failure-modes">the two failure modes<a class="anchor" href="#the-two-failure-modes" aria-label="link to this section">#</a></h2>
<p><strong>Done-done-done.</strong> Teams that discover the ambiguity often respond by inventing stages: "dev done", "QA done", "really done". This makes the confusion explicit without removing it, and adds ceremony. The fix is one definition, not three adjectives.</p>
<p><strong>Definition as theatre.</strong> A twenty-item checklist copied from a blog post, pasted into the wiki, referenced never. If your definition of done is not enforced by something — a pull request template, a merge check, a question someone actually asks in standup — it is decoration.</p>
<p>The test: pick a ticket closed last week and walk the list. If it fails two items and nobody noticed, the definition is not operating.</p>
<h2 id="the-part-about-flags-and-branches">the part about flags and branches<a class="anchor" href="#the-part-about-flags-and-branches" aria-label="link to this section">#</a></h2>
<p>The most common way "done" quietly is not: the feature ships behind a flag at 10%, works, and then nobody finishes the rollout. Three months later the flag is still at 10%, the old code path is still there, and the ticket has been closed since March.</p>
<p>That is not done. It is half-deployed with the cleanup unfunded, and it is how codebases accumulate the dual implementations that make everything else harder.</p>
<p><strong>Make the rollout part of the ticket, not a follow-up.</strong> A change is not finished when it is enabled for some users; it is finished when the decision has been made either way and the losing path is deleted.</p>
<h2 id="why-it-is-worth-the-argument">why it is worth the argument<a class="anchor" href="#why-it-is-worth-the-argument" aria-label="link to this section">#</a></h2>
<p>A shared definition of done is what makes "how much is left" answerable. Without it, the burn-down is measuring something nobody agrees on, estimates are uncomparable between people, and "almost done" means whatever the speaker wants.</p>
<p>It costs one conversation and a paragraph in the repository. It is the cheapest process improvement available to most teams and it is skipped because it sounds like process rather than engineering.</p>]]></content:encoded></item><item><title>The load-bearing comment</title><link>https://readme.news/the-load-bearing-comment/</link><guid isPermaLink="true">https://readme.news/the-load-bearing-comment/</guid><pubDate>Tue, 04 Aug 2026 09:00:00 +0000</pubDate><description>Most comments restate the code and rot. A few carry information the code cannot express, and those are worth defending.</description><content:encoded><![CDATA[<p>The advice that comments should explain <em>why</em> rather than <em>what</em> is correct and too abstract to act on. Here is the concrete version: a comment earns its place when it records information that could not have been derived from reading the code, and that someone would otherwise have to rediscover.</p>
<p>Everything else is decoration that will drift out of sync.</p>
<h2 id="the-four-that-earn-their-place">the four that earn their place<a class="anchor" href="#the-four-that-earn-their-place" aria-label="link to this section">#</a></h2>
<p><strong>The constraint that came from outside.</strong></p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python"># The vendor's API rejects batches over 500 even though their docs say 1000.
# Confirmed with their support, ticket #48812, March 2026.
BATCH_SIZE = 500</code></pre></div>
<p>Nothing in the code can tell you this. Without it, someone raises the number back to 1000 in eighteen months, ships it, and spends a day debugging.</p>
<p><strong>The alternative that was tried and failed.</strong></p>
<div class="code"><span class="code-lang">go</span><pre><code class="lang-go">// Tried sync.Map here first. It was slower for this access pattern — writes
// dominate and the keys are hot, which is the case sync.Map is worst at.
// Benchmark: bench/cache_test.go, BenchmarkHotKeys.</code></pre></div>
<p>This is the highest-value comment there is, because it stops the next person repeating your experiment. It converts a day of their work into a paragraph.</p>
<p><strong>The non-obvious ordering.</strong></p>
<div class="code"><span class="code-lang">javascript</span><pre><code class="lang-javascript">// Must run before initAuth(): the session store reads the feature flag cache,
// which is populated here. Reversing these fails silently — the user gets the
// default flag set rather than an error.</code></pre></div>
<p>"Fails silently" is the part that matters. A comment marking a trap is worth more than one marking a happy path.</p>
<p><strong>The deliberate weirdness.</strong></p>
<div class="code"><span class="code-lang">rust</span><pre><code class="lang-rust">// Deliberately not using the iterator: this is on the hot path and the
// bounds checks cost ~8% on the profile. See PERF-221.</code></pre></div>
<p>Without this, a well-meaning reviewer "cleans it up" and quietly regresses performance. Deliberate ugliness that isn't labelled reads as accidental ugliness, and accidental ugliness gets fixed.</p>
<h2 id="the-ones-that-do-not">the ones that do not<a class="anchor" href="#the-ones-that-do-not" aria-label="link to this section">#</a></h2>
<p><strong>Restating the line.</strong> <code>// increment counter</code> above <code>counter += 1</code>. This is the canonical example and it is still everywhere.</p>
<p><strong>Section headers in a long function.</strong> <code>// ---- validation ----</code> is a request for an <code>extract function</code> refactor wearing a disguise.</p>
<p><strong>Commented-out code.</strong> Delete it. Git has it. A block of dead code with no note is a question nobody can answer: is this a work in progress, a rollback, or something someone forgot?</p>
<p><strong>Changelogs in the file header.</strong> <code>// 2019-03-04 JS: added retry logic</code>. That is what the commit log is for, and it is accurate there.</p>
<p><strong>Comments that duplicate the type.</strong> <code>// returns a list of user IDs</code> above <code>fn user_ids() -&gt; Vec&lt;UserId&gt;</code>. The signature already said it, and the signature cannot go stale.</p>
<h2 id="the-decay-problem">the decay problem<a class="anchor" href="#the-decay-problem" aria-label="link to this section">#</a></h2>
<p>Every comment is a claim that is not checked by anything. The code changes; the comment does not; now it is actively misleading, and a misleading comment is worse than no comment because it is trusted.</p>
<p>Three things reduce the decay:</p>
<p><strong>Put it as close to the fact as possible.</strong> A comment above the line it describes survives longer than one at the top of the file describing behaviour forty lines down.</p>
<p><strong>Prefer things that are checked.</strong> A well-named function, a type, a test with a descriptive name, an assertion — all of these carry the same information and break when they become false. Reach for those first; comment only what none of them can express.</p>
<p><strong>Reference something durable.</strong> A ticket, a benchmark file, a commit hash, a support ticket number, a link to the vendor's documentation. Then a reader who doubts the comment can check it rather than guess.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>Before writing a comment, ask: <strong>could I encode this in a name, a type, or a test instead?</strong></p>
<p>If yes, do that — it will be checked, and the comment will not.</p>
<p>If no — if the information genuinely lives outside the code, in a vendor's behaviour or a decision someone made or an experiment that failed — write it down. That is exactly the case comments exist for, and those comments will still be earning their keep in a decade.</p>]]></content:encoded></item><item><title>Two hundred pieces in</title><link>https://readme.news/two-hundred-pieces-in/</link><guid isPermaLink="true">https://readme.news/two-hundred-pieces-in/</guid><pubDate>Mon, 03 Aug 2026 09:00:00 +0000</pubDate><description>What writing in public taught me about engineering, what I was most wrong about, and why the archive is the point.</description><content:encoded><![CDATA[<p>This is the two hundredth piece published here. That seems like a reasonable moment to write about the writing rather than about the subject.</p>
<h2 id="what-it-did-to-my-engineering">what it did to my engineering<a class="anchor" href="#what-it-did-to-my-engineering" aria-label="link to this section">#</a></h2>
<p><strong>It forced me to actually understand things.</strong> You can hold a fuzzy model of a technology in your head indefinitely and it feels like knowledge. The moment you try to explain it in a paragraph, the fuzziness becomes visible.</p>
<p>I have abandoned drafts because I discovered, four hundred words in, that I did not understand the thing well enough to write about it. Every one of those was more educational than the pieces I finished.</p>
<p><strong>It made me check things.</strong> Writing "X is faster than Y" in public means someone will ask for numbers. Knowing that in advance changes how you form the belief in the first place.</p>
<p><strong>It made me notice my own patterns.</strong> Reading two hundred pieces of my own writing back, the recurring themes are obvious to me now and were invisible while writing: verification over generation, boring over clever, measurement over intuition, and a persistent suspicion of anything that requires you to trust rather than check.</p>
<p>I did not set out with a thesis. It assembled itself.</p>
<p><strong>It taught me to be wrong in public</strong>, which is a skill and is uncomfortable and is the only way to find out you were wrong quickly.</p>
<h2 id="what-i-have-been-most-wrong-about">what I have been most wrong about<a class="anchor" href="#what-i-have-been-most-wrong-about" aria-label="link to this section">#</a></h2>
<p>I graded myself in December and the pattern held for the following eight months.</p>
<p><strong>I am reliably right about technical trajectories and reliably wrong about adoption.</strong></p>
<p>Local models got good; people did not switch, because hosted models got cheap faster than I expected. RAG got less necessary; the infrastructure repositioned instead of dying, which I have now failed to predict three separate times. Registry security controls were obviously needed; they arrived after the incident rather than before.</p>
<p>The lesson I keep relearning: the technology is the easy part to forecast, and the technology was never the hard part. Human and organizational behavior is where the uncertainty lives and where the consequences land.</p>
<p><strong>I under-predict inertia and over-predict rationality.</strong> Almost every wrong call has that shape.</p>
<h2 id="the-thing-about-writing-news">the thing about writing news<a class="anchor" href="#the-thing-about-writing-news" aria-label="link to this section">#</a></h2>
<p>Two hundred pieces, roughly half of them about things that happened in a specific week.</p>
<p>Reading them back, the ones that held up are almost never the ones that reported the event. They are the ones that used the event to explain a mechanism.</p>
<p>Nobody needs my summary of what a company announced. They can read the announcement. What is worth writing is: <em>why does this shape of thing keep happening</em>, and <em>what does it imply for what you should do on Monday</em>.</p>
<p>The news is a prompt. The mechanism is the article.</p>
<p>I did not know that when I started and it took about forty pieces to figure out.</p>
<h2 id="on-the-archive">on the archive<a class="anchor" href="#on-the-archive" aria-label="link to this section">#</a></h2>
<p>Everything published here is still at its original URL. Nothing has been quietly edited, renamed, or removed. Corrections are marked in place.</p>
<p>That is a deliberate choice and it costs something — there are pieces I would write differently now, and a few I think are wrong.</p>
<p>Leaving them up is the point. A publication that silently revises its history is not a record, it is a marketing surface. The value of an archive is that it shows what someone thought at the time, including when that was wrong, and you can only get that by not touching it.</p>
<p>If you want to know whether to trust a technical writer, check whether their old pieces still exist and whether the wrong ones were corrected in the open.</p>
<h2 id="on-the-format">on the format<a class="anchor" href="#on-the-format" aria-label="link to this section">#</a></h2>
<p>Gray background. Monospace headings. One column. No popups, no cookie banner, no newsletter modal, no autoplaying anything, no third-party JavaScript on any page.</p>
<p>This is not minimalism as an aesthetic. It is that every one of those things was added to a website to serve the publisher at the reader's expense, and the cumulative effect has made reading on the web genuinely unpleasant.</p>
<p>A page should render before you notice it loading. Text is the <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>. Nothing moves unless the reader moved it.</p>
<p>Those are not hard constraints to meet. Almost nobody meets them, and the reason is never technical.</p>
<h2 id="the-next-two-hundred">the next two hundred<a class="anchor" href="#the-next-two-hundred" aria-label="link to this section">#</a></h2>
<p>Same beat. Tech, developers, and the code underneath. News where the news teaches something, essays where the news does not.</p>
<p>More on the verification problem, because it is the defining engineering question of this period and it is nowhere near resolved. More on the craft, because the craft is what survives the tooling cycles. Fewer pieces about model launches, because they have stopped being informative.</p>
<p>Thanks for reading. Corrections are always welcome and get priority over everything else.</p>
<p>— Dom</p>]]></content:encoded></item><item><title>Two years of agentic coding: what stuck</title><link>https://readme.news/two-years-of-agentic-coding-what-stuck/</link><guid isPermaLink="true">https://readme.news/two-years-of-agentic-coding-what-stuck/</guid><pubDate>Fri, 31 Jul 2026 09:00:00 +0000</pubDate><description>The workflows that survived contact with real work, the ones that did not, and what the whole thing actually changed.</description><content:encoded><![CDATA[<p>Terminal coding agents went from <a class="xref" href="/operator-and-the-long-road-to-an-agent-that-can-click/" title="Operator, and the long road to an agent that can click">research preview</a> to standard tooling in about two years. Enough time has passed to separate what stuck from what was a phase.</p>
<h2 id="what-stuck">what stuck<a class="anchor" href="#what-stuck" aria-label="link to this section">#</a></h2>
<p><strong>Mechanical refactors at scale.</strong> The clearest win, by a wide margin. Renaming a concept across four hundred files, migrating a deprecated API, converting a pattern used everywhere. Verifiable, tedious, and exactly what the tools are good at.</p>
<p>The important second-order effect: <strong>refactors that were too expensive to do now happen.</strong> A codebase where cross-cutting cleanup is affordable is a meaningfully better codebase, and that is a permanent improvement rather than a productivity number.</p>
<p><strong>Working in unfamiliar territory.</strong> A language you do not know, a framework you have not used, an API you have never touched. The median output in an unfamiliar domain is better than your first attempt, and reading it teaches you the idioms.</p>
<p>This is the use case I would defend most strongly and it is discussed least.</p>
<p><strong>Test generation from a specification.</strong> Not "write tests for this function" — that produces tests that assert the implementation. But "here is the behavior, write tests that verify it" works well and it inverts the effort in the right direction.</p>
<p><strong>Investigation.</strong> Reading logs, bisecting history, tracing a call path, summarizing a large diff. Parallelizable, cheap, and it saves the expensive resource, which is your attention.</p>
<p><strong>Repository-level instruction files.</strong> <code>AGENTS.md</code> and its equivalents became standard practice, and the discipline of writing down how your project actually works improved documentation for humans as a side effect.</p>
<h2 id="what-did-not-stick">what did not stick<a class="anchor" href="#what-did-not-stick" aria-label="link to this section">#</a></h2>
<p><strong>Fully autonomous feature development.</strong> The demo works. The real version produces a plausible implementation of a subtly different feature, because the requirements that live in someone's head were never written down.</p>
<p><strong>Agent fleets at high concurrency.</strong> The generation scales; the review does not. Two to three concurrent agents with one reviewer turned out to be the practical limit, and the constraint is entirely on the human side.</p>
<p><strong>Orchestration frameworks.</strong> Absorbed into the models, as function-calling libraries and JSON-repair libraries were before them. The durable layer was never orchestration.</p>
<p><strong>"Just describe it and it builds."</strong> For anything with design decisions, the description that is precise enough to produce the right result is approximately as long as the code, and writing it is the same work.</p>
<h2 id="what-actually-changed-about-the-job">what actually changed about the job<a class="anchor" href="#what-actually-changed-about-the-job" aria-label="link to this section">#</a></h2>
<p><strong>Review is the bottleneck, permanently.</strong> Generation got roughly two orders of magnitude cheaper. Verification got no cheaper at all. Everything downstream follows from that asymmetry and nothing in two years has changed it.</p>
<p><strong>Tests became the primary artifact.</strong> If the implementation is cheap and verification is expensive, effort moves to specification. The teams getting the most out of these tools are the ones with strong test suites, and the correlation is not subtle.</p>
<p><strong><a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">Type systems</a> got a promotion.</strong> Every constraint the compiler checks is verification you do not perform by reading. Teams that were ambivalent about strict typing became evangelists, and the reason is always the same: it catches the class of error generated code produces most.</p>
<p><strong>Small diffs became non-negotiable.</strong> A machine can produce two thousand lines effortlessly. Accepting it is not a favor to anyone.</p>
<p><strong>The skill that separates people is judgment, not speed.</strong> It always was. It is now the only thing, and the gap between engineers who can tell when output is wrong and engineers who cannot is much more visible than it was.</p>
<h2 id="the-thing-still-unresolved">the thing still unresolved<a class="anchor" href="#the-thing-still-unresolved" aria-label="link to this section">#</a></h2>
<p>The apprenticeship problem.</p>
<p>The judgment that makes a senior engineer valuable was acquired by writing a lot of code badly and then debugging it. That work is being automated. Nobody has a replacement for how the next generation acquires it, and junior hiring contracted sharply during exactly the period when the training mechanism was being removed.</p>
<p>I have written about this several times and I still do not have an answer beyond: deliberately do the hard part yourself sometimes, review generated code carefully as a learning exercise, and hire juniors anyway.</p>
<p>That is a partial answer to a structural problem and I am not satisfied with it.</p>
<h2 id="the-honest-summary-two-years-in">the honest summary, two years in<a class="anchor" href="#the-honest-summary-two-years-in" aria-label="link to this section">#</a></h2>
<p>These tools are genuinely useful and the useful envelope is narrower and more specific than either the enthusiasts or the skeptics claimed.</p>
<p>They are excellent at bounded, verifiable, tedious work. They are unreliable at anything requiring judgment about what should be built. They multiply output and do not multiply throughput, because throughput is limited by review.</p>
<p>The engineers getting the most from them are the ones who were already good at specifying problems precisely and at telling when something is wrong. That is not a new skill and it was never evenly distributed.</p>
<p>Which is roughly what every previous tooling revolution did: raised the floor, moved the bottleneck, and made expertise more valuable rather than less.</p>]]></content:encoded></item><item><title>Deleting code is the highest-value work nobody schedules</title><link>https://readme.news/deleting-code-is-the-highest-value-work-nobody-schedules/</link><guid isPermaLink="true">https://readme.news/deleting-code-is-the-highest-value-work-nobody-schedules/</guid><pubDate>Wed, 29 Jul 2026 09:00:00 +0000</pubDate><description>Every line you remove is one that cannot break, cannot be misread, and does not need to be maintained.</description><content:encoded><![CDATA[<p>The most valuable pull request I have ever reviewed removed eleven thousand lines and added forty.</p>
<p>Deleting code is the only refactoring that is unambiguously good. It cannot introduce a bug in the deleted code, because there is no deleted code. It reduces build time, test time, cognitive load, security surface, and the probability that someone reads the wrong thing.</p>
<p>And nobody schedules it.</p>
<h2 id="what-accumulates">what accumulates<a class="anchor" href="#what-accumulates" aria-label="link to this section">#</a></h2>
<p><strong>Dead code.</strong> Never called. Nobody noticed, because nothing fails when unused code exists.</p>
<p><strong>Features nobody uses.</strong> Built for a customer who churned, an experiment that ended, a requirement that changed. Still there, still tested, still maintained, still appearing in every search result.</p>
<p><strong>Abandoned abstractions.</strong> Someone built a plugin system for the second plugin, which was never written. Now every call goes through a registry that has one entry.</p>
<p><strong>Configuration for conditions that no longer occur.</strong> A flag for a migration that completed in 2023.</p>
<p><strong>Compatibility shims.</strong> For a version nobody runs, an API that was removed, a browser that no longer exists.</p>
<p><strong>Tests for deleted behavior.</strong> Still running, still slow, testing something that cannot happen.</p>
<p><strong>Vendored copies.</strong> Of a library that is now a real dependency.</p>
<p><strong>Commented-out code.</strong> Always. Delete it. Git remembers.</p>
<h2 id="how-to-find-it">how to find it<a class="anchor" href="#how-to-find-it" aria-label="link to this section">#</a></h2>
<p><strong>Coverage, over a long window.</strong> Run coverage in production if your language supports it, or over your full integration suite. Anything at zero across a month is a candidate.</p>
<p>Be careful: zero coverage does not prove dead. It might be an error path, a rare branch, or something exercised only in a region you did not sample. Verify before deleting.</p>
<p><strong><a class="xref" href="/static-analysis-is-finally-worth-the-false-positives/" title="Static analysis is finally worth the false positives">Static analysis</a> for unreachable code.</strong> Most linters find unreferenced functions within a module. Cross-module dead code is harder and several tools do it.</p>
<p><strong><a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">Feature flags</a> at 100% for over a quarter.</strong> The disabled branch is dead code with a switch on it.</p>
<p><strong>Endpoints with no traffic.</strong> Log every route. Anything with zero requests in ninety days is a candidate. This is trivially easy to check and almost nobody does.</p>
<p><strong>Deprecation warnings nobody triggers.</strong> If you have been logging a deprecation warning for a year and it has never fired, the deprecated thing is unused.</p>
<p><strong><code>git log</code> on the file.</strong> Anything untouched for three years in an actively developed codebase is either perfect or forgotten.</p>
<h2 id="how-to-do-it-safely">how to do it safely<a class="anchor" href="#how-to-do-it-safely" aria-label="link to this section">#</a></h2>
<p><strong>Log before you delete.</strong> For anything you are not certain about, add logging and wait. A month of zero calls is strong evidence.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def old_thing(x):
    log.warning("old_thing called", stack=traceback.format_stack())
    return new_thing(x)</code></pre></div>
<p><strong>Delete in a separate commit</strong> from any other change. A deletion mixed with a refactor is unreviewable and un-revertable.</p>
<p><strong>Delete the tests too.</strong> Tests for deleted code are the most common thing left behind and they will confuse the next person enormously.</p>
<p><strong>Do not comment it out.</strong> Do not move it to an <code>old/</code> directory. Do not keep it "just in case." Git has it. If you genuinely might need it, note the commit hash in the deletion's commit message.</p>
<p><strong>Do it in batches, by area.</strong> One area, fully cleaned, is better than a thousand scattered deletions that are impossible to review.</p>
<h2 id="the-resistance-you-will-meet">the resistance you will meet<a class="anchor" href="#the-resistance-you-will-meet" aria-label="link to this section">#</a></h2>
<p><strong>"What if we need it?"</strong> You will not. And if you do, it is in git. In many years I have never seen a team need to recover deleted code that they could not recover.</p>
<p><strong>"Someone might be using it."</strong> Measure. That is what the logging is for. Guessing in either direction is worse than checking.</p>
<p><strong>"It works, why touch it?"</strong> Because it costs. Every line is read by every person who greps this file, is compiled on every build, is scanned by every security tool, and is a possible place for a future bug.</p>
<p><strong>"That is not a priority."</strong> Correct, and it never will be, which is why it has to be scheduled rather than prioritized. Ten percent of one sprint, quarterly, with a line count as the deliverable.</p>
<h2 id="the-framing-that-gets-it-done">the framing that gets it done<a class="anchor" href="#the-framing-that-gets-it-done" aria-label="link to this section">#</a></h2>
<p>Make it a competition. Track lines removed. Celebrate the largest deletion of the quarter.</p>
<p>This is slightly silly and it works, because it inverts the default incentive. Engineers are implicitly rewarded for adding — features shipped, code written — and never for removing, even though removing is frequently worth more.</p>
<p>Naming it, tracking it, and praising it is enough to change the behavior.</p>
<h2 id="the-number">the number<a class="anchor" href="#the-number" aria-label="link to this section">#</a></h2>
<p>A large fraction of most mature codebases is dead or effectively dead. I have never audited one where it was under 10%, and I have seen 40%.</p>
<p>Every line of that is being read, compiled, tested, scanned, and maintained, at a cost nobody has ever measured.</p>
<p>Go find some.</p>]]></content:encoded></item><item><title>The performance budget</title><link>https://readme.news/the-performance-budget/</link><guid isPermaLink="true">https://readme.news/the-performance-budget/</guid><pubDate>Mon, 27 Jul 2026 09:00:00 +0000</pubDate><description>A number, agreed in advance, that fails the build. The only mechanism that has ever kept a system fast.</description><content:encoded><![CDATA[<p>Every system starts fast and gets slower. Not because of one bad decision — because of a hundred small ones, each of which added two milliseconds, none of which anyone could reasonably object to.</p>
<p>The only mechanism I have seen reliably prevent this is a budget: a number, agreed in advance, that fails the build when exceeded.</p>
<h2 id="why-we-should-keep-it-fast-does-not-work">why "we should keep it fast" does not work<a class="anchor" href="#why-we-should-keep-it-fast-does-not-work" aria-label="link to this section">#</a></h2>
<p>Because "fast" has no threshold, so there is never a moment where a specific change is the problem.</p>
<p>Every individual addition is defensible. The library is 12 KB and saves a week. The query is 8 ms and enables a feature. The middleware is 3 ms and improves security.</p>
<p>Ten of those and your page is 400 ms slower, and no single change was wrong. There was no point at which anyone could say no, because the comparison was always "this change versus nothing" rather than "this change versus the budget."</p>
<p>A budget changes the comparison. Now the question is "what are you willing to remove to make room for this," which is a real conversation.</p>
<h2 id="setting-the-numbers">setting the numbers<a class="anchor" href="#setting-the-numbers" aria-label="link to this section">#</a></h2>
<p>Derive them from user-facing outcomes, not from what you currently have.</p>
<p><strong>For a web frontend:</strong></p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">metric</th><th style="text-align:left">budget</th></tr></thead><tbody><tr><td style="text-align:left">JavaScript, compressed</td><td style="text-align:left">170 KB</td></tr><tr><td style="text-align:left">CSS, compressed</td><td style="text-align:left">60 KB</td></tr><tr><td style="text-align:left">Largest Contentful Paint (p75, mobile)</td><td style="text-align:left">2.5 s</td></tr><tr><td style="text-align:left">Interaction to Next Paint (p75)</td><td style="text-align:left">200 ms</td></tr><tr><td style="text-align:left">Total requests, initial load</td><td style="text-align:left">40</td></tr></tbody></table></div>
<p><strong>For an API:</strong></p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">metric</th><th style="text-align:left">budget</th></tr></thead><tbody><tr><td style="text-align:left">p50 latency</td><td style="text-align:left">50 ms</td></tr><tr><td style="text-align:left">p99 latency</td><td style="text-align:left">500 ms</td></tr><tr><td style="text-align:left">database queries per request</td><td style="text-align:left">10</td></tr><tr><td style="text-align:left">memory per instance</td><td style="text-align:left">512 MB</td></tr></tbody></table></div>
<p><strong>Per-layer budgets are the underrated part.</strong> A single end-to-end number tells you something regressed. Per-layer budgets tell you <em>where</em>.</p>
<div class="code"><pre><code>total request budget: 200 ms
  auth:        10 ms
  validation:   5 ms
  database:    80 ms
  business:    40 ms
  serialize:   15 ms
  overhead:    50 ms</code></pre></div>
<p>When the database layer goes to 120 ms, that specific budget fails, and the person who changed the query knows immediately rather than in a month when someone investigates a general slowdown.</p>
<p>This is borrowed directly from game development, where per-subsystem frame budgets have been standard practice for decades.</p>
<h2 id="enforcement">enforcement<a class="anchor" href="#enforcement" aria-label="link to this section">#</a></h2>
<p>The budget must fail something, or it is a wish.</p>
<p><strong>In CI, on every pull request:</strong></p>
<div class="code"><span class="code-lang">yaml</span><pre><code class="lang-yaml">- name: bundle size
  run: npx size-limit          # fails if over the configured budget

- name: performance test
  run: k6 run --threshold 'http_req_duration{p(99)}&lt;500' load.js</code></pre></div>
<p><strong>Report the delta on the pull request.</strong> "This change adds 8 KB to the main bundle (142 KB → 150 KB, budget 170 KB)." Visible, in context, at the moment of decision.</p>
<p><strong>Allow overrides with a written reason.</strong> Not blocking forever — blocking until somebody says why. The record of overrides is itself useful; if you are overriding every week, the budget is wrong or the system is losing.</p>
<h2 id="the-budget-for-query-count">the budget for query count<a class="anchor" href="#the-budget-for-query-count" aria-label="link to this section">#</a></h2>
<p>The single most useful backend budget, and the least common.</p>
<p>Count database queries per request in tests. Assert a maximum.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">with assert_max_queries(10):
    client.get("/api/orders")</code></pre></div>
<p>This catches N+1 queries at the moment they are introduced, which is the only cheap time to catch them. An N+1 that reaches production is found weeks later by someone investigating a slow endpoint, and by then it is embedded in an ORM relationship that four other things depend on.</p>
<p>Almost every framework has a way to do this, and almost nobody does.</p>
<h2 id="what-happens-when-you-exceed-it">what happens when you exceed it<a class="anchor" href="#what-happens-when-you-exceed-it" aria-label="link to this section">#</a></h2>
<p>The conversation that a budget forces, in order:</p>
<ol><li><strong>Can we make the new thing cheaper?</strong> Lazy load it, defer it, make the query better.</li><li><strong>Can we remove something else?</strong> This is the valuable one. It surfaces the feature nobody uses and the library that is doing 5% of what it costs.</li><li><strong>Should the budget change?</strong> Sometimes yes, deliberately, with a reason recorded. A budget that never changes is a budget that will be ignored.</li><li><strong>Do we not ship this?</strong> Rare and it should be available.</li></ol>
<p>Any of those is better than the default, which is that the change lands and the system is permanently slower.</p>
<h2 id="the-thing-to-measure-first">the thing to measure first<a class="anchor" href="#the-thing-to-measure-first" aria-label="link to this section">#</a></h2>
<p>Before setting a budget, get the current numbers and the distribution. p50, p75, p95, p99. On real user hardware and real networks, not on a developer laptop on office wifi.</p>
<p>Then set the budget at roughly where you are, and ratchet it down over time rather than setting an aspirational number you fail immediately.</p>
<p>A budget you exceed on day one gets disabled on day two.</p>]]></content:encoded></item><item><title>What I look for in a codebase in the first hour</title><link>https://readme.news/what-i-look-for-in-a-codebase-in-the-first-hour/</link><guid isPermaLink="true">https://readme.news/what-i-look-for-in-a-codebase-in-the-first-hour/</guid><pubDate>Fri, 24 Jul 2026 09:00:00 +0000</pubDate><description>A checklist for assessing an unfamiliar codebase quickly — for a job, a due diligence, or a project you inherited.</description><content:encoded><![CDATA[<p>You have an hour with an unfamiliar codebase and you need to form a judgment: is this healthy, what will it cost to work in, what is the risk.</p>
<p>Here is the order I go in, and what each thing tells you.</p>
<h2 id="1-can-i-run-it-15-minutes">1. can I run it? (15 minutes)<a class="anchor" href="#1-can-i-run-it-15-minutes" aria-label="link to this section">#</a></h2>
<p>Clone it. Follow the README. Start a timer.</p>
<p>This is the single most informative test and most people skip it in favor of reading code.</p>
<ul><li><strong>Under 10 minutes to a running application:</strong> the team cares about developer experience, and probably about a lot of other things.</li><li><strong>An hour, with several undocumented steps:</strong> onboarding costs a week and every new hire pays it.</li><li><strong>You cannot get it running:</strong> the only people who can work on this are the ones who already have it working. This is a serious risk and it is invisible from the outside.</li></ul>
<p>Note every step that failed. That list <em>is</em> the health assessment.</p>
<h2 id="2-the-test-suite-10-minutes">2. the test suite (10 minutes)<a class="anchor" href="#2-the-test-suite-10-minutes" aria-label="link to this section">#</a></h2>
<p>Run it. Then look at it.</p>
<p><strong>Does it pass?</strong> On a clean checkout, first try. If not, that tells you the team has normalized a red build.</p>
<p><strong>How long?</strong> Under two minutes is excellent. Over ten and people have stopped running it locally.</p>
<p><strong>What is the ratio of assertions to setup?</strong> Read three test files. If setup dominates, the code is heavily coupled.</p>
<p><strong>Are there tests for the error paths?</strong> Almost nobody writes these and their presence is a strong positive signal.</p>
<p><strong>Is there a <code>skip</code> or <code>xfail</code> graveyard?</strong> Count them. A pile of disabled tests means the suite has been losing a slow argument with reality.</p>
<h2 id="3-the-shape-of-the-repository-5-minutes">3. the shape of the repository (5 minutes)<a class="anchor" href="#3-the-shape-of-the-repository-5-minutes" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">tokei .            # lines by language
git log --oneline | wc -l
git log --format='%an' | sort | uniq -c | sort -rn | head</code></pre></div>
<p><strong>Contributor concentration.</strong> If one person wrote 80% of it and they left, that is the largest risk in the codebase, larger than anything technical.</p>
<p><strong>Language sprawl.</strong> Four languages in a small project usually means four sets of tooling and nobody who understands all of it.</p>
<p><strong>Directory structure.</strong> Does it reflect the domain or the framework? <code>models/</code>, <code>views/</code>, <code>controllers/</code> tells you nothing about what the software does. <code>billing/</code>, <code>inventory/</code>, <code>shipping/</code> tells you everything.</p>
<h2 id="4-the-largest-files-5-minutes">4. the largest files (5 minutes)<a class="anchor" href="#4-the-largest-files-5-minutes" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">find . -name '*.py' -not -path '*/.venv/*' | xargs wc -l | sort -rn | head -20</code></pre></div>
<p>Every codebase has a few enormous files. Open the biggest one.</p>
<ul><li><strong>Is it generated?</strong> Fine, ignore it.</li><li><strong>Is it a god object?</strong> A 4,000-line service class is where all the complexity accumulated and where all the bugs live.</li><li><strong>When was it last changed?</strong> <code>git log -1</code> on it. If it is huge and changes weekly, that is the hot spot, and any work you do will touch it.</li></ul>
<h2 id="5-dependencies-5-minutes">5. dependencies (5 minutes)<a class="anchor" href="#5-dependencies-5-minutes" aria-label="link to this section">#</a></h2>
<p><strong>How many?</strong> Compare against similar projects. An unusual number in either direction is worth understanding.</p>
<p><strong>How old?</strong> Anything more than two major versions behind is an upgrade project waiting for you.</p>
<p><strong>Anything abandoned?</strong> Check the largest ones for last release date. A critical dependency with no release in three years is a fork you have not made yet.</p>
<p><strong>Anything surprising?</strong> A cryptography library nobody has heard of. A vendored copy of something. A dependency on a specific fork.</p>
<h2 id="6-the-commit-history-10-minutes">6. the commit history (10 minutes)<a class="anchor" href="#6-the-commit-history-10-minutes" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">git log --oneline -50</code></pre></div>
<p><strong>Message quality.</strong> "fix", "wip", "asdf" versus real descriptions. This is a direct readout of the engineering culture and it is remarkably predictive.</p>
<p><strong>Commit size.</strong> Are they atomic, or is every commit a 3,000-line dump?</p>
<p><strong>Is there review?</strong> Merge commits from pull requests, or direct pushes to the main branch?</p>
<p><strong>Cadence.</strong> Steady, or bursts separated by silence?</p>
<h2 id="7-the-things-that-are-missing-5-minutes">7. the things that are missing (5 minutes)<a class="anchor" href="#7-the-things-that-are-missing-5-minutes" aria-label="link to this section">#</a></h2>
<p>Frequently the most informative part.</p>
<ul><li><strong>No CI configuration.</strong> Nothing is checked automatically.</li><li><strong>No linter or formatter config.</strong> Every file is a different style and every review argues about it.</li><li><strong>No <code>CONTRIBUTING.md</code> or equivalent</strong> on a project with multiple contributors.</li><li><strong>No <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a>.</strong> Nobody knows what changed between versions.</li><li><strong>No architecture documentation of any kind.</strong> The system exists only in people's heads.</li><li><strong>No <code>.env.example</code>.</strong> You cannot configure it without asking someone.</li></ul>
<h2 id="8-one-real-feature-end-to-end-10-minutes">8. one real feature, end to end (10 minutes)<a class="anchor" href="#8-one-real-feature-end-to-end-10-minutes" aria-label="link to this section">#</a></h2>
<p>Pick a user-visible feature and trace it from the entry point to the database.</p>
<p>This is the highest-value ten minutes. You learn: how many layers, how much indirection, where the business logic lives, whether the abstractions are consistent, and whether you could add a similar feature without asking anyone.</p>
<p>If you cannot follow it in ten minutes, neither can anyone else, and every change will be expensive.</p>
<h2 id="the-summary-judgment">the summary judgment<a class="anchor" href="#the-summary-judgment" aria-label="link to this section">#</a></h2>
<p>After an hour I can usually say:</p>
<ul><li><strong>How long until a new engineer is productive.</strong> (From step 1 and 8.)</li><li><strong>Whether changes are safe.</strong> (From step 2.)</li><li><strong>Where the risk is concentrated.</strong> (From steps 3 and 4.)</li><li><strong>What the team values.</strong> (From steps 6 and 7.)</li></ul>
<p>None of that requires understanding what the software does. It is all structural, and structure is what determines the cost of working in something.</p>]]></content:encoded></item><item><title>The unreasonable effectiveness of a changelog</title><link>https://readme.news/the-unreasonable-effectiveness-of-a-changelog/</link><guid isPermaLink="true">https://readme.news/the-unreasonable-effectiveness-of-a-changelog/</guid><pubDate>Wed, 22 Jul 2026 09:00:00 +0000</pubDate><description>A file that takes ten minutes per release and answers most of the questions your users would otherwise ask you.</description><content:encoded><![CDATA[<p>Most projects do not have a changelog. Most projects have a commit log and a release page auto-generated from pull request titles, which is not the same thing and does not serve the same purpose.</p>
<p>A real changelog is a small amount of work with an outsized return.</p>
<h2 id="what-it-is-for">what it is for<a class="anchor" href="#what-it-is-for" aria-label="link to this section">#</a></h2>
<p><strong>Deciding whether to upgrade.</strong> The single most common reason someone reads a changelog. They are on version 3.2, version 3.7 exists, and they want to know whether it is worth the risk.</p>
<p><strong>Knowing what will break.</strong> The most important information you can provide, and the thing auto-generated release notes are worst at.</p>
<p><strong>Debugging.</strong> "This started failing after we upgraded" — a good changelog turns that into "here is the change that caused it" in thirty seconds.</p>
<p><strong>Finding out what exists.</strong> People discover features by reading changelogs. This is a real and underrated distribution channel for your own work.</p>
<h2 id="why-generated-release-notes-are-not-enough">why generated release notes are not enough<a class="anchor" href="#why-generated-release-notes-are-not-enough" aria-label="link to this section">#</a></h2>
<p>A list of merged pull request titles has three problems.</p>
<p><strong>It is written for the wrong audience.</strong> "Refactor connection handling" means something to the maintainer and nothing to the user. What changed <em>for them</em>?</p>
<p><strong>It has no hierarchy.</strong> A breaking change and a typo fix appear as sibling bullets of equal weight.</p>
<p><strong>It has no migration guidance.</strong> "Remove deprecated <code>parse()</code> method" tells you something broke. It does not tell you what to do about it.</p>
<h2 id="the-format">the format<a class="anchor" href="#the-format" aria-label="link to this section">#</a></h2>
<p>Keep a Changelog is the established convention and it is good. The structure:</p>
<div class="code"><span class="code-lang">markdown</span><pre><code class="lang-markdown">## [4.2.0] - 2026-07-22

### Breaking
- `Client.connect()` no longer accepts a positional timeout.
  Pass `timeout=` as a keyword.
      # before
      client.connect(host, 30)
      # after
      client.connect(host, timeout=30)

### Added
- `Client.ping()` for health checks without a full round trip (#412)
- Support for Unix domain sockets via `unix://` URLs (#398)

### Fixed
- Connections leaked when the handshake timed out (#405).
  If you saw file descriptor exhaustion under load, this was it.

### Deprecated
- `Client.legacy_mode` — will be removed in 5.0. Use `compatibility=`.

### Security
- Fixed a case where credentials could appear in debug logs (GHSA-xxxx-xxxx).
  Affects 4.0.0–4.1.3. Rotate credentials if debug logging was enabled.</code></pre></div>
<h2 id="the-rules-that-make-it-useful">the rules that make it useful<a class="anchor" href="#the-rules-that-make-it-useful" aria-label="link to this section">#</a></h2>
<p><strong>Breaking changes first, always.</strong> That is what people are scanning for. Do not bury them under twelve feature bullets.</p>
<p><strong>Include the migration.</strong> A breaking change without "do this instead" makes the reader open your source code. Two lines of before-and-after saves everyone time.</p>
<p><strong>Describe the user-visible effect, not the implementation.</strong> Not "refactored the retry logic." Rather: "<a class="xref" href="/retries-a-complete-guide-to-not-making-it-worse/" title="Retries: a complete guide to not making it worse">retries</a> now use exponential backoff with jitter; if you relied on the previous fixed 1-second interval, set <code>retry_delay=1.0</code>."</p>
<p><strong>Say who is affected.</strong> "If you use X, this changes for you. Otherwise nothing changes." Most readers can then stop reading, which is a service.</p>
<p><strong>Link to the issue or pull request</strong> for anyone who wants detail. The changelog is a summary, not a substitute.</p>
<p><strong>Date every release</strong>, in ISO format. Version numbers alone do not tell you whether you are two months or three years behind.</p>
<p><strong>Write it as you go</strong>, not at release time. An <code>Unreleased</code> section at the top that each pull request adds to. Reconstructing a changelog from git history at release time is miserable and it is why changelogs get skipped.</p>
<h2 id="the-security-section-specifically">the security section specifically<a class="anchor" href="#the-security-section-specifically" aria-label="link to this section">#</a></h2>
<p>If you fix a security issue, say so, with:</p>
<ul><li>Which versions are affected.</li><li>What the impact is.</li><li>Whether any action beyond upgrading is required.</li></ul>
<p>That last one is the part that gets omitted and it is critical. "Upgrade to 4.2.0" is insufficient if credentials may have been exposed — the user also needs to rotate them, and they will not know unless you say so.</p>
<h2 id="for-internal-projects">for internal projects<a class="anchor" href="#for-internal-projects" aria-label="link to this section">#</a></h2>
<p>The same file works for internal services, and the audience is your future self and the person who takes over the service.</p>
<p>The most valuable internal changelog entries are the ones that record a decision:</p>
<div class="code"><span class="code-lang">markdown</span><pre><code class="lang-markdown">## 2026-07-14
- Switched from polling to webhooks for order status. Polling was
  costing ~40 requests/second against the vendor's rate limit and
  we were getting throttled during peaks.</code></pre></div>
<p>Six months later, when someone asks why there is webhook infrastructure, the answer is in one place.</p>
<h2 id="the-return-on-investment">the return on investment<a class="anchor" href="#the-return-on-investment" aria-label="link to this section">#</a></h2>
<p>Ten minutes per release. In exchange:</p>
<ul><li>Fewer support questions.</li><li>Faster upgrades by your users, which means fewer people on old versions you have to support.</li><li>Fewer surprised users after a breaking change.</li><li>A record you can search when debugging.</li></ul>
<p>There are not many ten-minute tasks with that profile.</p>]]></content:encoded></item>
</channel>
</rss>
