{
 "version": "https://jsonfeed.org/version/1.1",
 "title": "README \u2014 tech and developer news",
 "home_page_url": "https://readme.news/",
 "feed_url": "https://readme.news/feed.json",
 "description": "README is a newsletter about software and the people who write it. Tech news, developer tooling, language releases, infrastructure, and long-form writing on the craft.",
 "icon": "https://readme.news/card.png",
 "favicon": "https://readme.news/favicon.svg",
 "authors": [
  {
   "name": "Dom the Developer",
   "url": "https://domthedeveloper.com"
  }
 ],
 "language": "en-US",
 "items": [
  {
   "id": "https://readme.news/two-hundred-pieces-in/",
   "url": "https://readme.news/two-hundred-pieces-in/",
   "title": "Two hundred pieces in",
   "summary": "What writing in public taught me about engineering, what I was most wrong about, and why the archive is the point.",
   "content_html": "<p>This is the two hundredth piece published here. That seems like a reasonable moment to write about the writing rather than about the subject.</p>\n<h2 id=\"what-it-did-to-my-engineering\">what it did to my engineering<a class=\"anchor\" href=\"#what-it-did-to-my-engineering\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>It forced me to actually understand things.</strong> You can hold a fuzzy model of a technology in your head indefinitely and it feels like knowledge. The moment you try to explain it in a paragraph, the fuzziness becomes visible.</p>\n<p>I have abandoned drafts because I discovered, four hundred words in, that I did not understand the thing well enough to write about it. Every one of those was more educational than the pieces I finished.</p>\n<p><strong>It made me check things.</strong> Writing \"X is faster than Y\" in public means someone will ask for numbers. Knowing that in advance changes how you form the belief in the first place.</p>\n<p><strong>It made me notice my own patterns.</strong> Reading two hundred pieces of my own writing back, the recurring themes are obvious to me now and were invisible while writing: verification over generation, boring over clever, measurement over intuition, and a persistent suspicion of anything that requires you to trust rather than check.</p>\n<p>I did not set out with a thesis. It assembled itself.</p>\n<p><strong>It taught me to be wrong in public</strong>, which is a skill and is uncomfortable and is the only way to find out you were wrong quickly.</p>\n<h2 id=\"what-i-have-been-most-wrong-about\">what I have been most wrong about<a class=\"anchor\" href=\"#what-i-have-been-most-wrong-about\" aria-label=\"link to this section\">#</a></h2>\n<p>I graded myself in December and the pattern held for the following eight months.</p>\n<p><strong>I am reliably right about technical trajectories and reliably wrong about adoption.</strong></p>\n<p>Local models got good; people did not switch, because hosted models got cheap faster than I expected. RAG got less necessary; the infrastructure repositioned instead of dying, which I have now failed to predict three separate times. Registry security controls were obviously needed; they arrived after the incident rather than before.</p>\n<p>The lesson I keep relearning: the technology is the easy part to forecast, and the technology was never the hard part. Human and organizational behavior is where the uncertainty lives and where the consequences land.</p>\n<p><strong>I under-predict inertia and over-predict rationality.</strong> Almost every wrong call has that shape.</p>\n<h2 id=\"the-thing-about-writing-news\">the thing about writing news<a class=\"anchor\" href=\"#the-thing-about-writing-news\" aria-label=\"link to this section\">#</a></h2>\n<p>Two hundred pieces, roughly half of them about things that happened in a specific week.</p>\n<p>Reading them back, the ones that held up are almost never the ones that reported the event. They are the ones that used the event to explain a mechanism.</p>\n<p>Nobody needs my summary of what a company announced. They can read the announcement. What is worth writing is: <em>why does this shape of thing keep happening</em>, and <em>what does it imply for what you should do on Monday</em>.</p>\n<p>The news is a prompt. The mechanism is the article.</p>\n<p>I did not know that when I started and it took about forty pieces to figure out.</p>\n<h2 id=\"on-the-archive\">on the archive<a class=\"anchor\" href=\"#on-the-archive\" aria-label=\"link to this section\">#</a></h2>\n<p>Everything published here is still at its original URL. Nothing has been quietly edited, renamed, or removed. Corrections are marked in place.</p>\n<p>That is a deliberate choice and it costs something \u2014 there are pieces I would write differently now, and a few I think are wrong.</p>\n<p>Leaving them up is the point. A publication that silently revises its history is not a record, it is a marketing surface. The value of an archive is that it shows what someone thought at the time, including when that was wrong, and you can only get that by not touching it.</p>\n<p>If you want to know whether to trust a technical writer, check whether their old pieces still exist and whether the wrong ones were corrected in the open.</p>\n<h2 id=\"on-the-format\">on the format<a class=\"anchor\" href=\"#on-the-format\" aria-label=\"link to this section\">#</a></h2>\n<p>Gray background. Monospace headings. One column. No popups, no cookie banner, no newsletter modal, no autoplaying anything, no third-party JavaScript on any page.</p>\n<p>This is not minimalism as an aesthetic. It is that every one of those things was added to a website to serve the publisher at the reader's expense, and the cumulative effect has made reading on the web genuinely unpleasant.</p>\n<p>A page should render before you notice it loading. Text is the interface. Nothing moves unless the reader moved it.</p>\n<p>Those are not hard constraints to meet. Almost nobody meets them, and the reason is never technical.</p>\n<h2 id=\"the-next-two-hundred\">the next two hundred<a class=\"anchor\" href=\"#the-next-two-hundred\" aria-label=\"link to this section\">#</a></h2>\n<p>Same beat. Tech, developers, and the code underneath. News where the news teaches something, essays where the news does not.</p>\n<p>More on the verification problem, because it is the defining engineering question of this period and it is nowhere near resolved. More on the craft, because the craft is what survives the tooling cycles. Fewer pieces about model launches, because they have stopped being informative.</p>\n<p>Thanks for reading. Corrections are always welcome and get priority over everything else.</p>\n<p>\u2014 Dom</p>",
   "image": "https://readme.news/cards/two-hundred-pieces-in.png",
   "date_published": "2026-08-03T09:00:00Z",
   "date_modified": "2026-08-03T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "retrospective",
    "craft",
    "opinion"
   ]
  },
  {
   "id": "https://readme.news/the-ai-acts-high-risk-obligations-take-effect-today/",
   "url": "https://readme.news/the-ai-acts-high-risk-obligations-take-effect-today/",
   "title": "The AI Act's high-risk obligations take effect today",
   "summary": "Two years after entry into force, the substantive requirements arrive. What changes, what was delayed, and what to do now.",
   "content_html": "<p>The EU AI Act's obligations for high-risk AI systems apply from today. This has been scheduled since the Act entered into force in August 2024, and it is the point at which the most substantive requirements become enforceable.</p>\n<h2 id=\"what-applies-now\">what applies now<a class=\"anchor\" href=\"#what-applies-now\" aria-label=\"link to this section\">#</a></h2>\n<p>For providers of high-risk systems \u2014 those used in employment, education access, credit, essential services, law enforcement, migration, justice, and as safety components in regulated products:</p>\n<ul><li><strong>Risk management system</strong>, documented and maintained across the lifecycle.</li><li><strong>Data governance</strong>, with documented training, validation, and test data and attention to bias.</li><li><strong>Technical documentation</strong> sufficient for conformity assessment.</li><li><strong>Automatic logging</strong> of the system's operation, retained.</li><li><strong>Transparency</strong> to deployers about capabilities, limitations, and intended use.</li><li><strong>Human oversight</strong> designed into the system.</li><li><strong>Accuracy, robustness, and cybersecurity</strong> proportionate to the purpose.</li><li><strong>Conformity assessment</strong> and CE marking before placing on the market.</li><li><strong>Registration</strong> in the EU database.</li><li><strong>Post-market monitoring</strong> and serious incident reporting.</li></ul>\n<p>Deployers \u2014 organizations using a high-risk system rather than supplying it \u2014 have a lighter but real set: use it according to instructions, ensure input data is relevant, monitor operation, retain logs, and assign competent human oversight.</p>\n<h2 id=\"the-parts-that-moved\">the parts that moved<a class=\"anchor\" href=\"#the-parts-that-moved\" aria-label=\"link to this section\">#</a></h2>\n<p>The implementation timeline has been adjusted more than once during the run-up, which is normal for regulation of this scope and which made planning genuinely difficult for anyone trying to comply.</p>\n<p>The important practical consequences:</p>\n<p><strong>Some obligations phase in later</strong> for systems already on the market, and for high-risk systems that are safety components of products covered by other EU legislation.</p>\n<p><strong>Harmonised standards are still being finalised.</strong> Conformity assessment is easier when you can demonstrate compliance against a standard; where standards are not yet available, providers must demonstrate compliance against the requirements directly, which is more work and more uncertain.</p>\n<p><strong>Enforcement capacity varies by member state.</strong> Market surveillance authorities are at different stages of readiness. That is not a reason to assume non-enforcement \u2014 it is a reason to expect inconsistency in the first period.</p>\n<p>Check the current state before making decisions. The details have moved and may move again.</p>\n<h2 id=\"what-to-do-if-you-are-in-scope\">what to do if you are in scope<a class=\"anchor\" href=\"#what-to-do-if-you-are-in-scope\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>If you have not classified your systems, do that today.</strong> It is the prerequisite for everything else and it is a legal question with technical inputs. A number of organizations discover they are in scope for one feature they had not considered \u2014 a CV screening tool, a proctoring feature, a creditworthiness proxy.</p>\n<p><strong>Documentation is the bulk of the work.</strong> Training data provenance, validation methodology, known failure modes, intended use and misuse. Most engineering teams do not have this written down and reconstructing it is slower than producing it as you go.</p>\n<p><strong>Logging is a technical requirement.</strong> Automatic recording of operation, retained, with enough detail to trace a decision. If your system does not do this, it is engineering work with a deadline that has passed.</p>\n<p><strong>Human oversight is a design constraint, not a checkbox.</strong> \"A person reviews the output\" is insufficient if the person cannot meaningfully understand, question, or override it. Surfacing confidence, explaining the inputs that drove a decision, and making override easy and recorded are interface requirements.</p>\n<h2 id=\"if-you-are-not-in-scope\">if you are not in scope<a class=\"anchor\" href=\"#if-you-are-not-in-scope\" aria-label=\"link to this section\">#</a></h2>\n<p>Most software is not. The transparency obligations \u2014 disclose that users are interacting with an AI system, label synthetic media \u2014 apply much more broadly and are much lighter.</p>\n<p>The thing worth doing regardless of scope: <strong>know which category you are in, and write down why.</strong> \"We assessed this and concluded it is not high-risk because X\" is a one-page document that saves an enormous amount of time when someone asks, and someone will ask.</p>\n<h2 id=\"the-honest-assessment\">the honest assessment<a class=\"anchor\" href=\"#the-honest-assessment\" aria-label=\"link to this section\">#</a></h2>\n<p>The Act is the most comprehensive AI regulation any jurisdiction has attempted. It has real criticisms \u2014 compliance cost falls hardest on small companies, the risk categories map imperfectly onto how systems are actually built, and the standards process has lagged the deadlines.</p>\n<p>It is also law, it has extraterritorial reach, and penalties scale with global turnover.</p>\n<p>The organizations that started classification work a year ago are in reasonable shape today. The ones that waited for clarity are discovering that regulatory clarity tends to arrive after the deadline rather than before it, which is a lesson that generalizes well beyond this particular Act.</p>",
   "image": "https://readme.news/cards/the-ai-acts-high-risk-obligations-take-effect-today.png",
   "date_published": "2026-08-02T09:00:00Z",
   "date_modified": "2026-08-02T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "policy",
    "ai"
   ]
  },
  {
   "id": "https://readme.news/two-years-of-agentic-coding-what-stuck/",
   "url": "https://readme.news/two-years-of-agentic-coding-what-stuck/",
   "title": "Two years of agentic coding: what stuck",
   "summary": "The workflows that survived contact with real work, the ones that did not, and what the whole thing actually changed.",
   "content_html": "<p>Terminal coding agents went from research preview to standard tooling in about two years. Enough time has passed to separate what stuck from what was a phase.</p>\n<h2 id=\"what-stuck\">what stuck<a class=\"anchor\" href=\"#what-stuck\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Mechanical refactors at scale.</strong> The clearest win, by a wide margin. Renaming a concept across four hundred files, migrating a deprecated API, converting a pattern used everywhere. Verifiable, tedious, and exactly what the tools are good at.</p>\n<p>The important second-order effect: <strong>refactors that were too expensive to do now happen.</strong> A codebase where cross-cutting cleanup is affordable is a meaningfully better codebase, and that is a permanent improvement rather than a productivity number.</p>\n<p><strong>Working in unfamiliar territory.</strong> A language you do not know, a framework you have not used, an API you have never touched. The median output in an unfamiliar domain is better than your first attempt, and reading it teaches you the idioms.</p>\n<p>This is the use case I would defend most strongly and it is discussed least.</p>\n<p><strong>Test generation from a specification.</strong> Not \"write tests for this function\" \u2014 that produces tests that assert the implementation. But \"here is the behavior, write tests that verify it\" works well and it inverts the effort in the right direction.</p>\n<p><strong>Investigation.</strong> Reading logs, bisecting history, tracing a call path, summarizing a large diff. Parallelizable, cheap, and it saves the expensive resource, which is your attention.</p>\n<p><strong>Repository-level instruction files.</strong> <code>AGENTS.md</code> and its equivalents became standard practice, and the discipline of writing down how your project actually works improved documentation for humans as a side effect.</p>\n<h2 id=\"what-did-not-stick\">what did not stick<a class=\"anchor\" href=\"#what-did-not-stick\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Fully autonomous feature development.</strong> The demo works. The real version produces a plausible implementation of a subtly different feature, because the requirements that live in someone's head were never written down.</p>\n<p><strong>Agent fleets at high concurrency.</strong> The generation scales; the review does not. Two to three concurrent agents with one reviewer turned out to be the practical limit, and the constraint is entirely on the human side.</p>\n<p><strong>Orchestration frameworks.</strong> Absorbed into the models, as function-calling libraries and JSON-repair libraries were before them. The durable layer was never orchestration.</p>\n<p><strong>\"Just describe it and it builds.\"</strong> For anything with design decisions, the description that is precise enough to produce the right result is approximately as long as the code, and writing it is the same work.</p>\n<h2 id=\"what-actually-changed-about-the-job\">what actually changed about the job<a class=\"anchor\" href=\"#what-actually-changed-about-the-job\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Review is the bottleneck, permanently.</strong> Generation got roughly two orders of magnitude cheaper. Verification got no cheaper at all. Everything downstream follows from that asymmetry and nothing in two years has changed it.</p>\n<p><strong>Tests became the primary artifact.</strong> If the implementation is cheap and verification is expensive, effort moves to specification. The teams getting the most out of these tools are the ones with strong test suites, and the correlation is not subtle.</p>\n<p><strong>Type systems got a promotion.</strong> Every constraint the compiler checks is verification you do not perform by reading. Teams that were ambivalent about strict typing became evangelists, and the reason is always the same: it catches the class of error generated code produces most.</p>\n<p><strong>Small diffs became non-negotiable.</strong> A machine can produce two thousand lines effortlessly. Accepting it is not a favor to anyone.</p>\n<p><strong>The skill that separates people is judgment, not speed.</strong> It always was. It is now the only thing, and the gap between engineers who can tell when output is wrong and engineers who cannot is much more visible than it was.</p>\n<h2 id=\"the-thing-still-unresolved\">the thing still unresolved<a class=\"anchor\" href=\"#the-thing-still-unresolved\" aria-label=\"link to this section\">#</a></h2>\n<p>The apprenticeship problem.</p>\n<p>The judgment that makes a senior engineer valuable was acquired by writing a lot of code badly and then debugging it. That work is being automated. Nobody has a replacement for how the next generation acquires it, and junior hiring contracted sharply during exactly the period when the training mechanism was being removed.</p>\n<p>I have written about this several times and I still do not have an answer beyond: deliberately do the hard part yourself sometimes, review generated code carefully as a learning exercise, and hire juniors anyway.</p>\n<p>That is a partial answer to a structural problem and I am not satisfied with it.</p>\n<h2 id=\"the-honest-summary-two-years-in\">the honest summary, two years in<a class=\"anchor\" href=\"#the-honest-summary-two-years-in\" aria-label=\"link to this section\">#</a></h2>\n<p>These tools are genuinely useful and the useful envelope is narrower and more specific than either the enthusiasts or the skeptics claimed.</p>\n<p>They are excellent at bounded, verifiable, tedious work. They are unreliable at anything requiring judgment about what should be built. They multiply output and do not multiply throughput, because throughput is limited by review.</p>\n<p>The engineers getting the most from them are the ones who were already good at specifying problems precisely and at telling when something is wrong. That is not a new skill and it was never evenly distributed.</p>\n<p>Which is roughly what every previous tooling revolution did: raised the floor, moved the bottleneck, and made expertise more valuable rather than less.</p>",
   "image": "https://readme.news/cards/two-years-of-agentic-coding-what-stuck.png",
   "date_published": "2026-07-31T09:00:00Z",
   "date_modified": "2026-07-31T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "ai",
    "coding-agents",
    "craft",
    "retrospective"
   ]
  },
  {
   "id": "https://readme.news/deleting-code-is-the-highest-value-work-nobody-schedules/",
   "url": "https://readme.news/deleting-code-is-the-highest-value-work-nobody-schedules/",
   "title": "Deleting code is the highest-value work nobody schedules",
   "summary": "Every line you remove is one that cannot break, cannot be misread, and does not need to be maintained.",
   "content_html": "<p>The most valuable pull request I have ever reviewed removed eleven thousand lines and added forty.</p>\n<p>Deleting code is the only refactoring that is unambiguously good. It cannot introduce a bug in the deleted code, because there is no deleted code. It reduces build time, test time, cognitive load, security surface, and the probability that someone reads the wrong thing.</p>\n<p>And nobody schedules it.</p>\n<h2 id=\"what-accumulates\">what accumulates<a class=\"anchor\" href=\"#what-accumulates\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Dead code.</strong> Never called. Nobody noticed, because nothing fails when unused code exists.</p>\n<p><strong>Features nobody uses.</strong> Built for a customer who churned, an experiment that ended, a requirement that changed. Still there, still tested, still maintained, still appearing in every search result.</p>\n<p><strong>Abandoned abstractions.</strong> Someone built a plugin system for the second plugin, which was never written. Now every call goes through a registry that has one entry.</p>\n<p><strong>Configuration for conditions that no longer occur.</strong> A flag for a migration that completed in 2023.</p>\n<p><strong>Compatibility shims.</strong> For a version nobody runs, an API that was removed, a browser that no longer exists.</p>\n<p><strong>Tests for deleted behavior.</strong> Still running, still slow, testing something that cannot happen.</p>\n<p><strong>Vendored copies.</strong> Of a library that is now a real dependency.</p>\n<p><strong>Commented-out code.</strong> Always. Delete it. Git remembers.</p>\n<h2 id=\"how-to-find-it\">how to find it<a class=\"anchor\" href=\"#how-to-find-it\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Coverage, over a long window.</strong> Run coverage in production if your language supports it, or over your full integration suite. Anything at zero across a month is a candidate.</p>\n<p>Be careful: zero coverage does not prove dead. It might be an error path, a rare branch, or something exercised only in a region you did not sample. Verify before deleting.</p>\n<p><strong>Static analysis for unreachable code.</strong> Most linters find unreferenced functions within a module. Cross-module dead code is harder and several tools do it.</p>\n<p><strong>Feature flags at 100% for over a quarter.</strong> The disabled branch is dead code with a switch on it.</p>\n<p><strong>Endpoints with no traffic.</strong> Log every route. Anything with zero requests in ninety days is a candidate. This is trivially easy to check and almost nobody does.</p>\n<p><strong>Deprecation warnings nobody triggers.</strong> If you have been logging a deprecation warning for a year and it has never fired, the deprecated thing is unused.</p>\n<p><strong><code>git log</code> on the file.</strong> Anything untouched for three years in an actively developed codebase is either perfect or forgotten.</p>\n<h2 id=\"how-to-do-it-safely\">how to do it safely<a class=\"anchor\" href=\"#how-to-do-it-safely\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Log before you delete.</strong> For anything you are not certain about, add logging and wait. A month of zero calls is strong evidence.</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">def old_thing(x):\n    log.warning(\"old_thing called\", stack=traceback.format_stack())\n    return new_thing(x)</code></pre></div>\n<p><strong>Delete in a separate commit</strong> from any other change. A deletion mixed with a refactor is unreviewable and un-revertable.</p>\n<p><strong>Delete the tests too.</strong> Tests for deleted code are the most common thing left behind and they will confuse the next person enormously.</p>\n<p><strong>Do not comment it out.</strong> Do not move it to an <code>old/</code> directory. Do not keep it \"just in case.\" Git has it. If you genuinely might need it, note the commit hash in the deletion's commit message.</p>\n<p><strong>Do it in batches, by area.</strong> One area, fully cleaned, is better than a thousand scattered deletions that are impossible to review.</p>\n<h2 id=\"the-resistance-you-will-meet\">the resistance you will meet<a class=\"anchor\" href=\"#the-resistance-you-will-meet\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>\"What if we need it?\"</strong> You will not. And if you do, it is in git. In many years I have never seen a team need to recover deleted code that they could not recover.</p>\n<p><strong>\"Someone might be using it.\"</strong> Measure. That is what the logging is for. Guessing in either direction is worse than checking.</p>\n<p><strong>\"It works, why touch it?\"</strong> Because it costs. Every line is read by every person who greps this file, is compiled on every build, is scanned by every security tool, and is a possible place for a future bug.</p>\n<p><strong>\"That is not a priority.\"</strong> Correct, and it never will be, which is why it has to be scheduled rather than prioritized. Ten percent of one sprint, quarterly, with a line count as the deliverable.</p>\n<h2 id=\"the-framing-that-gets-it-done\">the framing that gets it done<a class=\"anchor\" href=\"#the-framing-that-gets-it-done\" aria-label=\"link to this section\">#</a></h2>\n<p>Make it a competition. Track lines removed. Celebrate the largest deletion of the quarter.</p>\n<p>This is slightly silly and it works, because it inverts the default incentive. Engineers are implicitly rewarded for adding \u2014 features shipped, code written \u2014 and never for removing, even though removing is frequently worth more.</p>\n<p>Naming it, tracking it, and praising it is enough to change the behavior.</p>\n<h2 id=\"the-number\">the number<a class=\"anchor\" href=\"#the-number\" aria-label=\"link to this section\">#</a></h2>\n<p>A large fraction of most mature codebases is dead or effectively dead. I have never audited one where it was under 10%, and I have seen 40%.</p>\n<p>Every line of that is being read, compiled, tested, scanned, and maintained, at a cost nobody has ever measured.</p>\n<p>Go find some.</p>",
   "image": "https://readme.news/cards/deleting-code-is-the-highest-value-work-nobody-schedules.png",
   "date_published": "2026-07-29T09:00:00Z",
   "date_modified": "2026-07-29T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "craft",
    "engineering",
    "opinion"
   ]
  },
  {
   "id": "https://readme.news/the-performance-budget/",
   "url": "https://readme.news/the-performance-budget/",
   "title": "The performance budget",
   "summary": "A number, agreed in advance, that fails the build. The only mechanism that has ever kept a system fast.",
   "content_html": "<p>Every system starts fast and gets slower. Not because of one bad decision \u2014 because of a hundred small ones, each of which added two milliseconds, none of which anyone could reasonably object to.</p>\n<p>The only mechanism I have seen reliably prevent this is a budget: a number, agreed in advance, that fails the build when exceeded.</p>\n<h2 id=\"why-we-should-keep-it-fast-does-not-work\">why \"we should keep it fast\" does not work<a class=\"anchor\" href=\"#why-we-should-keep-it-fast-does-not-work\" aria-label=\"link to this section\">#</a></h2>\n<p>Because \"fast\" has no threshold, so there is never a moment where a specific change is the problem.</p>\n<p>Every individual addition is defensible. The library is 12 KB and saves a week. The query is 8 ms and enables a feature. The middleware is 3 ms and improves security.</p>\n<p>Ten of those and your page is 400 ms slower, and no single change was wrong. There was no point at which anyone could say no, because the comparison was always \"this change versus nothing\" rather than \"this change versus the budget.\"</p>\n<p>A budget changes the comparison. Now the question is \"what are you willing to remove to make room for this,\" which is a real conversation.</p>\n<h2 id=\"setting-the-numbers\">setting the numbers<a class=\"anchor\" href=\"#setting-the-numbers\" aria-label=\"link to this section\">#</a></h2>\n<p>Derive them from user-facing outcomes, not from what you currently have.</p>\n<p><strong>For a web frontend:</strong></p>\n<div class=\"table-wrap\"><table><thead><tr><th style=\"text-align:left\">metric</th><th style=\"text-align:left\">budget</th></tr></thead><tbody><tr><td style=\"text-align:left\">JavaScript, compressed</td><td style=\"text-align:left\">170 KB</td></tr><tr><td style=\"text-align:left\">CSS, compressed</td><td style=\"text-align:left\">60 KB</td></tr><tr><td style=\"text-align:left\">Largest Contentful Paint (p75, mobile)</td><td style=\"text-align:left\">2.5 s</td></tr><tr><td style=\"text-align:left\">Interaction to Next Paint (p75)</td><td style=\"text-align:left\">200 ms</td></tr><tr><td style=\"text-align:left\">Total requests, initial load</td><td style=\"text-align:left\">40</td></tr></tbody></table></div>\n<p><strong>For an API:</strong></p>\n<div class=\"table-wrap\"><table><thead><tr><th style=\"text-align:left\">metric</th><th style=\"text-align:left\">budget</th></tr></thead><tbody><tr><td style=\"text-align:left\">p50 latency</td><td style=\"text-align:left\">50 ms</td></tr><tr><td style=\"text-align:left\">p99 latency</td><td style=\"text-align:left\">500 ms</td></tr><tr><td style=\"text-align:left\">database queries per request</td><td style=\"text-align:left\">10</td></tr><tr><td style=\"text-align:left\">memory per instance</td><td style=\"text-align:left\">512 MB</td></tr></tbody></table></div>\n<p><strong>Per-layer budgets are the underrated part.</strong> A single end-to-end number tells you something regressed. Per-layer budgets tell you <em>where</em>.</p>\n<div class=\"code\"><pre><code>total request budget: 200 ms\n  auth:        10 ms\n  validation:   5 ms\n  database:    80 ms\n  business:    40 ms\n  serialize:   15 ms\n  overhead:    50 ms</code></pre></div>\n<p>When the database layer goes to 120 ms, that specific budget fails, and the person who changed the query knows immediately rather than in a month when someone investigates a general slowdown.</p>\n<p>This is borrowed directly from game development, where per-subsystem frame budgets have been standard practice for decades.</p>\n<h2 id=\"enforcement\">enforcement<a class=\"anchor\" href=\"#enforcement\" aria-label=\"link to this section\">#</a></h2>\n<p>The budget must fail something, or it is a wish.</p>\n<p><strong>In CI, on every pull request:</strong></p>\n<div class=\"code\"><span class=\"code-lang\">yaml</span><pre><code class=\"lang-yaml\">- name: bundle size\n  run: npx size-limit          # fails if over the configured budget\n\n- name: performance test\n  run: k6 run --threshold 'http_req_duration{p(99)}&lt;500' load.js</code></pre></div>\n<p><strong>Report the delta on the pull request.</strong> \"This change adds 8 KB to the main bundle (142 KB \u2192 150 KB, budget 170 KB).\" Visible, in context, at the moment of decision.</p>\n<p><strong>Allow overrides with a written reason.</strong> Not blocking forever \u2014 blocking until somebody says why. The record of overrides is itself useful; if you are overriding every week, the budget is wrong or the system is losing.</p>\n<h2 id=\"the-budget-for-query-count\">the budget for query count<a class=\"anchor\" href=\"#the-budget-for-query-count\" aria-label=\"link to this section\">#</a></h2>\n<p>The single most useful backend budget, and the least common.</p>\n<p>Count database queries per request in tests. Assert a maximum.</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">with assert_max_queries(10):\n    client.get(\"/api/orders\")</code></pre></div>\n<p>This catches N+1 queries at the moment they are introduced, which is the only cheap time to catch them. An N+1 that reaches production is found weeks later by someone investigating a slow endpoint, and by then it is embedded in an ORM relationship that four other things depend on.</p>\n<p>Almost every framework has a way to do this, and almost nobody does.</p>\n<h2 id=\"what-happens-when-you-exceed-it\">what happens when you exceed it<a class=\"anchor\" href=\"#what-happens-when-you-exceed-it\" aria-label=\"link to this section\">#</a></h2>\n<p>The conversation that a budget forces, in order:</p>\n<ol><li><strong>Can we make the new thing cheaper?</strong> Lazy load it, defer it, make the query better.</li><li><strong>Can we remove something else?</strong> This is the valuable one. It surfaces the feature nobody uses and the library that is doing 5% of what it costs.</li><li><strong>Should the budget change?</strong> Sometimes yes, deliberately, with a reason recorded. A budget that never changes is a budget that will be ignored.</li><li><strong>Do we not ship this?</strong> Rare and it should be available.</li></ol>\n<p>Any of those is better than the default, which is that the change lands and the system is permanently slower.</p>\n<h2 id=\"the-thing-to-measure-first\">the thing to measure first<a class=\"anchor\" href=\"#the-thing-to-measure-first\" aria-label=\"link to this section\">#</a></h2>\n<p>Before setting a budget, get the current numbers and the distribution. p50, p75, p95, p99. On real user hardware and real networks, not on a developer laptop on office wifi.</p>\n<p>Then set the budget at roughly where you are, and ratchet it down over time rather than setting an aspirational number you fail immediately.</p>\n<p>A budget you exceed on day one gets disabled on day two.</p>",
   "image": "https://readme.news/cards/the-performance-budget.png",
   "date_published": "2026-07-27T09:00:00Z",
   "date_modified": "2026-07-27T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "performance",
    "engineering",
    "craft",
    "frontend"
   ]
  },
  {
   "id": "https://readme.news/what-i-look-for-in-a-codebase-in-the-first-hour/",
   "url": "https://readme.news/what-i-look-for-in-a-codebase-in-the-first-hour/",
   "title": "What I look for in a codebase in the first hour",
   "summary": "A checklist for assessing an unfamiliar codebase quickly \u2014 for a job, a due diligence, or a project you inherited.",
   "content_html": "<p>You have an hour with an unfamiliar codebase and you need to form a judgment: is this healthy, what will it cost to work in, what is the risk.</p>\n<p>Here is the order I go in, and what each thing tells you.</p>\n<h2 id=\"1-can-i-run-it-15-minutes\">1. can I run it? (15 minutes)<a class=\"anchor\" href=\"#1-can-i-run-it-15-minutes\" aria-label=\"link to this section\">#</a></h2>\n<p>Clone it. Follow the README. Start a timer.</p>\n<p>This is the single most informative test and most people skip it in favor of reading code.</p>\n<ul><li><strong>Under 10 minutes to a running application:</strong> the team cares about developer experience, and probably about a lot of other things.</li><li><strong>An hour, with several undocumented steps:</strong> onboarding costs a week and every new hire pays it.</li><li><strong>You cannot get it running:</strong> the only people who can work on this are the ones who already have it working. This is a serious risk and it is invisible from the outside.</li></ul>\n<p>Note every step that failed. That list <em>is</em> the health assessment.</p>\n<h2 id=\"2-the-test-suite-10-minutes\">2. the test suite (10 minutes)<a class=\"anchor\" href=\"#2-the-test-suite-10-minutes\" aria-label=\"link to this section\">#</a></h2>\n<p>Run it. Then look at it.</p>\n<p><strong>Does it pass?</strong> On a clean checkout, first try. If not, that tells you the team has normalized a red build.</p>\n<p><strong>How long?</strong> Under two minutes is excellent. Over ten and people have stopped running it locally.</p>\n<p><strong>What is the ratio of assertions to setup?</strong> Read three test files. If setup dominates, the code is heavily coupled.</p>\n<p><strong>Are there tests for the error paths?</strong> Almost nobody writes these and their presence is a strong positive signal.</p>\n<p><strong>Is there a <code>skip</code> or <code>xfail</code> graveyard?</strong> Count them. A pile of disabled tests means the suite has been losing a slow argument with reality.</p>\n<h2 id=\"3-the-shape-of-the-repository-5-minutes\">3. the shape of the repository (5 minutes)<a class=\"anchor\" href=\"#3-the-shape-of-the-repository-5-minutes\" aria-label=\"link to this section\">#</a></h2>\n<div class=\"code\"><span class=\"code-lang\">bash</span><pre><code class=\"lang-bash\">tokei .            # lines by language\ngit log --oneline | wc -l\ngit log --format='%an' | sort | uniq -c | sort -rn | head</code></pre></div>\n<p><strong>Contributor concentration.</strong> If one person wrote 80% of it and they left, that is the largest risk in the codebase, larger than anything technical.</p>\n<p><strong>Language sprawl.</strong> Four languages in a small project usually means four sets of tooling and nobody who understands all of it.</p>\n<p><strong>Directory structure.</strong> Does it reflect the domain or the framework? <code>models/</code>, <code>views/</code>, <code>controllers/</code> tells you nothing about what the software does. <code>billing/</code>, <code>inventory/</code>, <code>shipping/</code> tells you everything.</p>\n<h2 id=\"4-the-largest-files-5-minutes\">4. the largest files (5 minutes)<a class=\"anchor\" href=\"#4-the-largest-files-5-minutes\" aria-label=\"link to this section\">#</a></h2>\n<div class=\"code\"><span class=\"code-lang\">bash</span><pre><code class=\"lang-bash\">find . -name '*.py' -not -path '*/.venv/*' | xargs wc -l | sort -rn | head -20</code></pre></div>\n<p>Every codebase has a few enormous files. Open the biggest one.</p>\n<ul><li><strong>Is it generated?</strong> Fine, ignore it.</li><li><strong>Is it a god object?</strong> A 4,000-line service class is where all the complexity accumulated and where all the bugs live.</li><li><strong>When was it last changed?</strong> <code>git log -1</code> on it. If it is huge and changes weekly, that is the hot spot, and any work you do will touch it.</li></ul>\n<h2 id=\"5-dependencies-5-minutes\">5. dependencies (5 minutes)<a class=\"anchor\" href=\"#5-dependencies-5-minutes\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>How many?</strong> Compare against similar projects. An unusual number in either direction is worth understanding.</p>\n<p><strong>How old?</strong> Anything more than two major versions behind is an upgrade project waiting for you.</p>\n<p><strong>Anything abandoned?</strong> Check the largest ones for last release date. A critical dependency with no release in three years is a fork you have not made yet.</p>\n<p><strong>Anything surprising?</strong> A cryptography library nobody has heard of. A vendored copy of something. A dependency on a specific fork.</p>\n<h2 id=\"6-the-commit-history-10-minutes\">6. the commit history (10 minutes)<a class=\"anchor\" href=\"#6-the-commit-history-10-minutes\" aria-label=\"link to this section\">#</a></h2>\n<div class=\"code\"><span class=\"code-lang\">bash</span><pre><code class=\"lang-bash\">git log --oneline -50</code></pre></div>\n<p><strong>Message quality.</strong> \"fix\", \"wip\", \"asdf\" versus real descriptions. This is a direct readout of the engineering culture and it is remarkably predictive.</p>\n<p><strong>Commit size.</strong> Are they atomic, or is every commit a 3,000-line dump?</p>\n<p><strong>Is there review?</strong> Merge commits from pull requests, or direct pushes to the main branch?</p>\n<p><strong>Cadence.</strong> Steady, or bursts separated by silence?</p>\n<h2 id=\"7-the-things-that-are-missing-5-minutes\">7. the things that are missing (5 minutes)<a class=\"anchor\" href=\"#7-the-things-that-are-missing-5-minutes\" aria-label=\"link to this section\">#</a></h2>\n<p>Frequently the most informative part.</p>\n<ul><li><strong>No CI configuration.</strong> Nothing is checked automatically.</li><li><strong>No linter or formatter config.</strong> Every file is a different style and every review argues about it.</li><li><strong>No <code>CONTRIBUTING.md</code> or equivalent</strong> on a project with multiple contributors.</li><li><strong>No changelog.</strong> Nobody knows what changed between versions.</li><li><strong>No architecture documentation of any kind.</strong> The system exists only in people's heads.</li><li><strong>No <code>.env.example</code>.</strong> You cannot configure it without asking someone.</li></ul>\n<h2 id=\"8-one-real-feature-end-to-end-10-minutes\">8. one real feature, end to end (10 minutes)<a class=\"anchor\" href=\"#8-one-real-feature-end-to-end-10-minutes\" aria-label=\"link to this section\">#</a></h2>\n<p>Pick a user-visible feature and trace it from the entry point to the database.</p>\n<p>This is the highest-value ten minutes. You learn: how many layers, how much indirection, where the business logic lives, whether the abstractions are consistent, and whether you could add a similar feature without asking anyone.</p>\n<p>If you cannot follow it in ten minutes, neither can anyone else, and every change will be expensive.</p>\n<h2 id=\"the-summary-judgment\">the summary judgment<a class=\"anchor\" href=\"#the-summary-judgment\" aria-label=\"link to this section\">#</a></h2>\n<p>After an hour I can usually say:</p>\n<ul><li><strong>How long until a new engineer is productive.</strong> (From step 1 and 8.)</li><li><strong>Whether changes are safe.</strong> (From step 2.)</li><li><strong>Where the risk is concentrated.</strong> (From steps 3 and 4.)</li><li><strong>What the team values.</strong> (From steps 6 and 7.)</li></ul>\n<p>None of that requires understanding what the software does. It is all structural, and structure is what determines the cost of working in something.</p>",
   "image": "https://readme.news/cards/what-i-look-for-in-a-codebase-in-the-first-hour.png",
   "date_published": "2026-07-24T09:00:00Z",
   "date_modified": "2026-07-24T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "craft",
    "engineering",
    "code-review",
    "careers"
   ]
  },
  {
   "id": "https://readme.news/the-unreasonable-effectiveness-of-a-changelog/",
   "url": "https://readme.news/the-unreasonable-effectiveness-of-a-changelog/",
   "title": "The unreasonable effectiveness of a changelog",
   "summary": "A file that takes ten minutes per release and answers most of the questions your users would otherwise ask you.",
   "content_html": "<p>Most projects do not have a changelog. Most projects have a commit log and a release page auto-generated from pull request titles, which is not the same thing and does not serve the same purpose.</p>\n<p>A real changelog is a small amount of work with an outsized return.</p>\n<h2 id=\"what-it-is-for\">what it is for<a class=\"anchor\" href=\"#what-it-is-for\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Deciding whether to upgrade.</strong> The single most common reason someone reads a changelog. They are on version 3.2, version 3.7 exists, and they want to know whether it is worth the risk.</p>\n<p><strong>Knowing what will break.</strong> The most important information you can provide, and the thing auto-generated release notes are worst at.</p>\n<p><strong>Debugging.</strong> \"This started failing after we upgraded\" \u2014 a good changelog turns that into \"here is the change that caused it\" in thirty seconds.</p>\n<p><strong>Finding out what exists.</strong> People discover features by reading changelogs. This is a real and underrated distribution channel for your own work.</p>\n<h2 id=\"why-generated-release-notes-are-not-enough\">why generated release notes are not enough<a class=\"anchor\" href=\"#why-generated-release-notes-are-not-enough\" aria-label=\"link to this section\">#</a></h2>\n<p>A list of merged pull request titles has three problems.</p>\n<p><strong>It is written for the wrong audience.</strong> \"Refactor connection handling\" means something to the maintainer and nothing to the user. What changed <em>for them</em>?</p>\n<p><strong>It has no hierarchy.</strong> A breaking change and a typo fix appear as sibling bullets of equal weight.</p>\n<p><strong>It has no migration guidance.</strong> \"Remove deprecated <code>parse()</code> method\" tells you something broke. It does not tell you what to do about it.</p>\n<h2 id=\"the-format\">the format<a class=\"anchor\" href=\"#the-format\" aria-label=\"link to this section\">#</a></h2>\n<p>Keep a Changelog is the established convention and it is good. The structure:</p>\n<div class=\"code\"><span class=\"code-lang\">markdown</span><pre><code class=\"lang-markdown\">## [4.2.0] - 2026-07-22\n\n### Breaking\n- `Client.connect()` no longer accepts a positional timeout.\n  Pass `timeout=` as a keyword.\n      # before\n      client.connect(host, 30)\n      # after\n      client.connect(host, timeout=30)\n\n### Added\n- `Client.ping()` for health checks without a full round trip (#412)\n- Support for Unix domain sockets via `unix://` URLs (#398)\n\n### Fixed\n- Connections leaked when the handshake timed out (#405).\n  If you saw file descriptor exhaustion under load, this was it.\n\n### Deprecated\n- `Client.legacy_mode` \u2014 will be removed in 5.0. Use `compatibility=`.\n\n### Security\n- Fixed a case where credentials could appear in debug logs (GHSA-xxxx-xxxx).\n  Affects 4.0.0\u20134.1.3. Rotate credentials if debug logging was enabled.</code></pre></div>\n<h2 id=\"the-rules-that-make-it-useful\">the rules that make it useful<a class=\"anchor\" href=\"#the-rules-that-make-it-useful\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Breaking changes first, always.</strong> That is what people are scanning for. Do not bury them under twelve feature bullets.</p>\n<p><strong>Include the migration.</strong> A breaking change without \"do this instead\" makes the reader open your source code. Two lines of before-and-after saves everyone time.</p>\n<p><strong>Describe the user-visible effect, not the implementation.</strong> Not \"refactored the retry logic.\" Rather: \"retries now use exponential backoff with jitter; if you relied on the previous fixed 1-second interval, set <code>retry_delay=1.0</code>.\"</p>\n<p><strong>Say who is affected.</strong> \"If you use X, this changes for you. Otherwise nothing changes.\" Most readers can then stop reading, which is a service.</p>\n<p><strong>Link to the issue or pull request</strong> for anyone who wants detail. The changelog is a summary, not a substitute.</p>\n<p><strong>Date every release</strong>, in ISO format. Version numbers alone do not tell you whether you are two months or three years behind.</p>\n<p><strong>Write it as you go</strong>, not at release time. An <code>Unreleased</code> section at the top that each pull request adds to. Reconstructing a changelog from git history at release time is miserable and it is why changelogs get skipped.</p>\n<h2 id=\"the-security-section-specifically\">the security section specifically<a class=\"anchor\" href=\"#the-security-section-specifically\" aria-label=\"link to this section\">#</a></h2>\n<p>If you fix a security issue, say so, with:</p>\n<ul><li>Which versions are affected.</li><li>What the impact is.</li><li>Whether any action beyond upgrading is required.</li></ul>\n<p>That last one is the part that gets omitted and it is critical. \"Upgrade to 4.2.0\" is insufficient if credentials may have been exposed \u2014 the user also needs to rotate them, and they will not know unless you say so.</p>\n<h2 id=\"for-internal-projects\">for internal projects<a class=\"anchor\" href=\"#for-internal-projects\" aria-label=\"link to this section\">#</a></h2>\n<p>The same file works for internal services, and the audience is your future self and the person who takes over the service.</p>\n<p>The most valuable internal changelog entries are the ones that record a decision:</p>\n<div class=\"code\"><span class=\"code-lang\">markdown</span><pre><code class=\"lang-markdown\">## 2026-07-14\n- Switched from polling to webhooks for order status. Polling was\n  costing ~40 requests/second against the vendor's rate limit and\n  we were getting throttled during peaks.</code></pre></div>\n<p>Six months later, when someone asks why there is webhook infrastructure, the answer is in one place.</p>\n<h2 id=\"the-return-on-investment\">the return on investment<a class=\"anchor\" href=\"#the-return-on-investment\" aria-label=\"link to this section\">#</a></h2>\n<p>Ten minutes per release. In exchange:</p>\n<ul><li>Fewer support questions.</li><li>Faster upgrades by your users, which means fewer people on old versions you have to support.</li><li>Fewer surprised users after a breaking change.</li><li>A record you can search when debugging.</li></ul>\n<p>There are not many ten-minute tasks with that profile.</p>",
   "image": "https://readme.news/cards/the-unreasonable-effectiveness-of-a-changelog.png",
   "date_published": "2026-07-22T09:00:00Z",
   "date_modified": "2026-07-22T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "documentation",
    "open-source",
    "craft",
    "communication"
   ]
  },
  {
   "id": "https://readme.news/property-based-testing-deserves-its-moment/",
   "url": "https://readme.news/property-based-testing-deserves-its-moment/",
   "title": "Property-based testing deserves its moment",
   "summary": "State an invariant, let the machine find the counterexample. A twenty-year-old technique that fits this moment exactly.",
   "content_html": "<p>Example-based tests check the cases you thought of. Property-based tests check the cases you did not.</p>\n<p>The technique has been around since QuickCheck in 1999 and has stayed niche. It fits the current moment unusually well, and it is worth another look.</p>\n<h2 id=\"the-idea\">the idea<a class=\"anchor\" href=\"#the-idea\" aria-label=\"link to this section\">#</a></h2>\n<p>Instead of \"for this input, expect this output,\" you state a property that should hold for <em>all</em> inputs, and the framework generates hundreds of inputs trying to break it.</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">from hypothesis import given, strategies as st\n\n@given(st.lists(st.integers()))\ndef test_sort_is_idempotent(xs):\n    assert sorted(sorted(xs)) == sorted(xs)\n\n@given(st.lists(st.integers()))\ndef test_sort_preserves_elements(xs):\n    assert sorted(xs).count == xs.count   # same multiset\n\n@given(st.text())\ndef test_roundtrip(s):\n    assert decode(encode(s)) == s</code></pre></div>\n<p>The framework generates empty lists, single elements, lists with duplicates, huge values, negative numbers, and \u2014 critically \u2014 when it finds a failure, it <em>shrinks</em> the input to the minimal case that still fails.</p>\n<p>That shrinking is what makes the technique practical. A failure on a 400-element list is not actionable. The same failure shrunk to <code>[0, 0]</code> tells you exactly what is wrong.</p>\n<h2 id=\"the-properties-that-are-actually-useful\">the properties that are actually useful<a class=\"anchor\" href=\"#the-properties-that-are-actually-useful\" aria-label=\"link to this section\">#</a></h2>\n<p>The hard part is not the tooling. It is identifying properties, and there is a small catalogue that covers most cases.</p>\n<p><strong>Round trip.</strong> <code>decode(encode(x)) == x</code>. Serialization, compression, parsing, encryption. This one property catches an enormous number of real bugs and applies almost everywhere.</p>\n<p><strong>Invariants.</strong> Something that is always true regardless of operations. A balanced tree stays balanced. A sorted list stays sorted. Account balances sum to the same total across a transfer. A cache never returns stale data past its TTL.</p>\n<p><strong>Oracle.</strong> Compare against a simpler, slower, obviously-correct implementation. You have an optimized version; write the naive one, and assert they agree. This is extremely powerful and underused \u2014 it is how you test the fast path against the version you can reason about.</p>\n<p><strong>Idempotence.</strong> <code>f(f(x)) == f(x)</code>. Normalization, deduplication, and \u2014 importantly \u2014 any API operation that claims to be idempotent. If you have idempotency keys, this is how you actually verify them.</p>\n<p><strong>Commutativity and associativity.</strong> Order should not matter. Merge operations, set operations, CRDT merges.</p>\n<p><strong>Metamorphic relations.</strong> When you cannot state the correct output, state how the output should <em>change</em>. Adding an item to a cart should increase the total by that item's price. Searching for a more specific query should return a subset. Sorting descending should be the reverse of sorting ascending.</p>\n<p>This last category is the one that unlocks property testing for business logic, where there is no obvious oracle.</p>\n<h2 id=\"where-it-fits-best\">where it fits best<a class=\"anchor\" href=\"#where-it-fits-best\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Parsers and serializers.</strong> Round trip, always.</p>\n<p><strong>Data structures.</strong> Invariants after every operation sequence.</p>\n<p><strong>State machines.</strong> Generate random valid operation sequences, assert the invariants hold throughout. This finds ordering bugs that no hand-written test would.</p>\n<p><strong>Anything with an obvious naive implementation.</strong> Oracle testing.</p>\n<p><strong>Financial and unit arithmetic.</strong> Rounding, currency, conversions. The edge cases are numerous and boring, which is exactly what a generator is for.</p>\n<p><strong>Concurrent code</strong>, with a framework that generates interleavings. This is the hardest category to test any other way.</p>\n<h2 id=\"where-it-does-not-fit\">where it does not fit<a class=\"anchor\" href=\"#where-it-does-not-fit\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>UI.</strong> The properties are aesthetic.</p>\n<p><strong>Anything where the correct output requires human judgment.</strong></p>\n<p><strong>Integration tests against real systems.</strong> Generation implies many runs, and many runs against a real database is slow.</p>\n<p><strong>Code where you cannot state a property.</strong> If you genuinely cannot, that may itself be information about the design \u2014 code with no statable invariants is code with no contract.</p>\n<h2 id=\"the-practical-advice\">the practical advice<a class=\"anchor\" href=\"#the-practical-advice\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Do not replace example tests.</strong> Keep them. They document intent and they are the fastest way to communicate what a function is for. Add properties alongside.</p>\n<p><strong>Start with round-trip properties.</strong> Easiest to state, highest hit rate for real bugs.</p>\n<p><strong>Save the failing seed.</strong> When a property fails, the framework gives you the counterexample. Add it as a regression example test so it is checked deterministically forever.</p>\n<p><strong>Bound the input space.</strong> Unbounded generation produces absurd inputs and slow tests. Constrain to realistic ranges: strings up to a reasonable length, numbers in a plausible domain.</p>\n<p><strong>Run more cases in CI than locally.</strong> A hundred examples locally for fast feedback, a thousand in CI, ten thousand in a nightly run.</p>\n<h2 id=\"why-now\">why now<a class=\"anchor\" href=\"#why-now\" aria-label=\"link to this section\">#</a></h2>\n<p>Two reasons this fits the current moment.</p>\n<p><strong>Verification is the bottleneck.</strong> Property testing is verification you write once and that checks a large input space forever. That is exactly the leverage the moment calls for.</p>\n<p><strong>Stating a property is a task worth doing carefully by hand</strong>, and generating the implementation is not. Writing \"the total must always equal the sum of line items\" requires understanding the domain; nothing else in the test does.</p>\n<p>That is a good division of labor, and it points at where human effort should concentrate: on specifying what must be true, and letting everything else be derived.</p>",
   "image": "https://readme.news/cards/property-based-testing-deserves-its-moment.png",
   "date_published": "2026-07-20T09:00:00Z",
   "date_modified": "2026-07-20T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "testing",
    "craft",
    "engineering",
    "type-systems"
   ]
  },
  {
   "id": "https://readme.news/vendor-lock-in-an-honest-cost-model/",
   "url": "https://readme.news/vendor-lock-in-an-honest-cost-model/",
   "title": "Vendor lock-in: an honest cost model",
   "summary": "Portability is not free and neither is dependence. A framework for deciding how much abstraction to buy.",
   "content_html": "<p>\"Avoid vendor lock-in\" is treated as self-evidently good advice. It is not advice, it is a preference, and following it uncritically produces systems that are worse in exchange for optionality nobody will ever exercise.</p>\n<p>Here is a way to actually decide.</p>\n<h2 id=\"the-cost-of-lock-in\">the cost of lock-in<a class=\"anchor\" href=\"#the-cost-of-lock-in\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Switching cost.</strong> How much engineering time to move off, if you had to.</p>\n<p><strong>Pricing power.</strong> A vendor who knows you cannot leave prices accordingly. This is real and it is usually the largest ongoing cost.</p>\n<p><strong>Capability ceiling.</strong> You are limited to what they support, on their timeline.</p>\n<p><strong>Correlated risk.</strong> They have an outage, you have an outage. They change their terms, you comply. They get acquired and sunset the product, you migrate on their schedule.</p>\n<h2 id=\"the-cost-of-avoiding-lock-in\">the cost of avoiding lock-in<a class=\"anchor\" href=\"#the-cost-of-avoiding-lock-in\" aria-label=\"link to this section\">#</a></h2>\n<p>This is the half that gets ignored, and it is frequently larger.</p>\n<p><strong>The abstraction layer itself.</strong> Code to write, maintain, test, and debug through. It is a permanent tax and it makes stack traces longer.</p>\n<p><strong>Lowest common denominator.</strong> Your abstraction can only expose what all candidate providers support. You give up the features that made the good option good.</p>\n<p><strong>The abstraction is usually wrong anyway.</strong> It was designed against one provider's model. When you actually try to swap, you discover the abstraction encoded assumptions that do not hold, and you rewrite it.</p>\n<p><strong>Delayed value.</strong> Time spent on portability is time not spent on the product.</p>\n<h2 id=\"the-framework\">the framework<a class=\"anchor\" href=\"#the-framework\" aria-label=\"link to this section\">#</a></h2>\n<p>For each dependency, estimate:</p>\n<ol><li><strong>Switching cost</strong> \u2014 engineer-weeks to move.</li><li><strong>Probability you switch</strong> in the next three years.</li><li><strong>Cost of the abstraction</strong> \u2014 engineer-weeks now, plus ongoing drag.</li></ol>\n<p>Then: if <code>switching_cost \u00d7 probability &lt; abstraction_cost</code>, do not abstract.</p>\n<p>The numbers are rough. The exercise still clarifies, because it forces you to state the probability out loud, and stated probabilities are usually much lower than the implied ones people are acting on.</p>\n<h2 id=\"the-categories-worked-through\">the categories, worked through<a class=\"anchor\" href=\"#the-categories-worked-through\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Object storage.</strong> Switching cost: low. The S3 API is a de facto standard and every provider implements it. Probability: moderate \u2014 people do move for pricing.</p>\n<p><strong>Verdict: use the S3 API, do not abstract further.</strong> The API is already the abstraction.</p>\n<p><strong>Compute.</strong> Switching cost: moderate if containerized, high if you use provider-specific serverless. Probability: low.</p>\n<p><strong>Verdict: containerize</strong> \u2014 which is good practice anyway \u2014 <strong>and use whatever managed service you want.</strong> Do not build a compute abstraction layer.</p>\n<p><strong>Relational database.</strong> Switching cost: high. Probability: low.</p>\n<p><strong>Verdict: use the database's features.</strong> Teams that avoid stored procedures, database-specific types, and advanced indexing to stay portable are giving up real capability for an event that will not happen. Postgres-specific SQL is fine. You are not going to migrate to a different engine, and if you do, the SQL dialect will be the smallest part of the pain.</p>\n<p><strong>Managed queues, streams, and similar.</strong> Switching cost: moderate. The semantics differ enough between providers that a thin abstraction genuinely helps.</p>\n<p><strong>Verdict: a thin interface \u2014 publish, subscribe, ack \u2014 is worth it.</strong> Not a full abstraction; a boundary.</p>\n<p><strong>Authentication.</strong> Switching cost: very high \u2014 you have to migrate user credentials, sessions, and integrations. Probability: low, but the consequences of being stuck are severe.</p>\n<p><strong>Verdict: use standard protocols.</strong> OIDC and SAML are the abstraction. A provider that supports them is replaceable in principle; one with a proprietary SDK is not.</p>\n<p><strong>AI model providers.</strong> Switching cost: low if you kept the interface thin. Probability: <strong>high</strong> \u2014 the market is moving fast and the right choice changes quarterly.</p>\n<p><strong>Verdict: definitely abstract.</strong> This is the clearest case on the list. A thin interface plus an eval harness in your own repository makes model changes an afternoon. Teams that did this in 2024 have been switching providers casually ever since.</p>\n<p><strong>Observability.</strong> Switching cost: moderate to high \u2014 instrumentation is everywhere in your code. Probability: moderate, usually driven by cost.</p>\n<p><strong>Verdict: OpenTelemetry.</strong> The instrumentation is vendor-neutral, the collector handles routing, and swapping backends is a configuration change.</p>\n<h2 id=\"the-pattern\">the pattern<a class=\"anchor\" href=\"#the-pattern\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Abstract where the switching probability is high and the abstraction is cheap.</strong> AI providers, observability backends, object storage.</p>\n<p><strong>Do not abstract where switching is unlikely and the abstraction costs you real capability.</strong> Databases, compute platforms, managed services you chose for their specific features.</p>\n<p><strong>Use standard protocols wherever they exist.</strong> OIDC, S3, OpenTelemetry, SQL. A standard is an abstraction someone else maintains, and it is always cheaper than yours.</p>\n<h2 id=\"the-thing-that-actually-protects-you\">the thing that actually protects you<a class=\"anchor\" href=\"#the-thing-that-actually-protects-you\" aria-label=\"link to this section\">#</a></h2>\n<p>Not an abstraction layer. <strong>Your data, in a format you can export, and a documented process for leaving.</strong></p>\n<p>Ask, before adopting anything: can I get all my data out, in a usable format, without their cooperation? If yes, you have real optionality regardless of how coupled your code is, because the expensive part of a migration is never the code \u2014 it is the data.</p>\n<p>If the answer is no, that is a much bigger red flag than any API coupling, and it is the question almost nobody asks during procurement.</p>",
   "image": "https://readme.news/cards/vendor-lock-in-an-honest-cost-model.png",
   "date_published": "2026-07-17T09:00:00Z",
   "date_modified": "2026-07-17T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "architecture",
    "devops",
    "engineering",
    "opinion"
   ]
  },
  {
   "id": "https://readme.news/connection-pooling-explained-properly/",
   "url": "https://readme.news/connection-pooling-explained-properly/",
   "title": "Connection pooling, explained properly",
   "summary": "Why your database has 400 connections, why that is bad, and how to size a pool without guessing.",
   "content_html": "<p>Connection pool sizing is done by copying a number from a blog post, and the number is usually wrong in a specific and expensive way.</p>\n<h2 id=\"why-connections-are-expensive\">why connections are expensive<a class=\"anchor\" href=\"#why-connections-are-expensive\" aria-label=\"link to this section\">#</a></h2>\n<p>In Postgres specifically, each connection is a separate operating system process with its own memory. A few megabytes of baseline, plus work memory for sorting and hashing, plus its share of shared buffer access.</p>\n<p>Four hundred connections means four hundred processes. The scheduler is context switching between them, they are contending for the same locks and buffers, and the memory is largely wasted because most of them are idle.</p>\n<p>The counterintuitive result, which has been measured many times: <strong>throughput frequently goes down as connection count goes up, past a fairly low threshold.</strong></p>\n<p>More connections does not mean more concurrency. It means more contention.</p>\n<h2 id=\"the-actual-number\">the actual number<a class=\"anchor\" href=\"#the-actual-number\" aria-label=\"link to this section\">#</a></h2>\n<p>A widely used starting formula:</p>\n<div class=\"code\"><pre><code>connections = (core_count \u00d7 2) + effective_spindle_count</code></pre></div>\n<p>For an 8-core machine with SSD storage, that is somewhere around 16 to 20.</p>\n<p>That number seems shockingly low to people running pools of 100 or more. It is correct, and the reasoning is straightforward: a query is either using CPU or waiting on I/O. You need enough connections to keep the cores busy and to have some work queued behind I/O waits. Past that, additional connections are queued at the database instead of queued in your pool, and queueing at the database is worse because it consumes resources.</p>\n<p><strong>Test it.</strong> Take your load test, run it at pool sizes of 10, 20, 40, 80, and 160, and plot throughput and p99 latency. The curve rises, flattens, and then degrades. Most people are on the degrading side and have never looked.</p>\n<h2 id=\"the-pooler-layers\">the pooler layers<a class=\"anchor\" href=\"#the-pooler-layers\" aria-label=\"link to this section\">#</a></h2>\n<p>Three, and they do different things.</p>\n<p><strong>Application-side pool.</strong> In-process, reuses connections across requests. Every ORM and database driver has one. This is the minimum.</p>\n<p><strong>External pooler</strong> \u2014 PgBouncer, pgcat, or a cloud provider's equivalent. Sits between your application and the database and multiplexes many client connections onto few server connections.</p>\n<p>This is what you need when you have many application instances. Twenty instances with a pool of 20 each is 400 connections to the database, even if each instance is mostly idle. A pooler collapses that to the number the database actually wants.</p>\n<p><strong>The database's own limit.</strong> <code>max_connections</code>. Set it lower than you think \u2014 it is a safety valve, and setting it high does not make the database faster, it makes the failure mode worse.</p>\n<h2 id=\"pooling-modes-and-the-one-that-bites\">pooling modes, and the one that bites<a class=\"anchor\" href=\"#pooling-modes-and-the-one-that-bites\" aria-label=\"link to this section\">#</a></h2>\n<p>External poolers have modes and choosing wrong causes subtle correctness bugs.</p>\n<p><strong>Session pooling.</strong> A client gets a server connection for the duration of its session. Safe, and provides little multiplexing benefit.</p>\n<p><strong>Transaction pooling.</strong> A server connection is assigned per transaction and returned after commit. This is where the big multiplexing win is, and it is what most people want.</p>\n<p><strong>The catch:</strong> anything that depends on session state breaks.</p>\n<ul><li>Prepared statements (unless the pooler supports them explicitly, which newer ones do)</li><li><code>SET</code> at the session level</li><li>Session-level advisory locks</li><li><code>LISTEN</code>/<code>NOTIFY</code></li><li>Temporary tables</li><li><code>WITH HOLD</code> cursors</li></ul>\n<p>If your ORM uses server-side prepared statements by default \u2014 many do \u2014 you must either disable them or use a pooler that handles them. This is the single most common transaction-pooling problem and it manifests as intermittent errors under load, which is a miserable thing to debug.</p>\n<p><strong>Statement pooling.</strong> A connection per statement. Maximum multiplexing, breaks multi-statement transactions. Almost never what you want.</p>\n<h2 id=\"the-settings-that-matter\">the settings that matter<a class=\"anchor\" href=\"#the-settings-that-matter\" aria-label=\"link to this section\">#</a></h2>\n<p>Beyond size:</p>\n<p><strong>Connection timeout.</strong> How long a request waits for a connection before failing. This should be short \u2014 a few seconds. A request waiting thirty seconds for a connection has already been abandoned by its caller.</p>\n<p><strong>Idle timeout.</strong> Return connections to the database when not in use. Important with many application instances.</p>\n<p><strong>Max lifetime.</strong> Recycle connections periodically. This handles the case where a database failover happens and your pool is holding connections to the old primary \u2014 without a max lifetime, some pools hold those indefinitely.</p>\n<p><strong>Validation.</strong> Test a connection before handing it out, or handle the failure on first use. Something has to deal with connections that died while idle.</p>\n<h2 id=\"the-diagnostic\">the diagnostic<a class=\"anchor\" href=\"#the-diagnostic\" aria-label=\"link to this section\">#</a></h2>\n<div class=\"code\"><span class=\"code-lang\">sql</span><pre><code class=\"lang-sql\">SELECT state, count(*), max(now() - state_change) AS longest\nFROM pg_stat_activity\nWHERE backend_type = 'client backend'\nGROUP BY state;</code></pre></div>\n<p>What you are looking for:</p>\n<ul><li><strong>Many <code>idle in transaction</code></strong> \u2014 the worst state. A transaction is open and doing nothing, holding locks and preventing vacuum. This is an application bug: a transaction opened and not committed, usually because of an early return or an exception path.</li><li><strong>Many <code>idle</code></strong> \u2014 pool is oversized. Harmless but wasteful.</li><li><strong>Many <code>active</code> with long durations</strong> \u2014 queries are slow; the pool is not the problem.</li></ul>\n<p><code>idle in transaction</code> is the one to alert on. It is always a bug and it causes outages that look like database problems and are application problems.</p>\n<h2 id=\"the-summary\">the summary<a class=\"anchor\" href=\"#the-summary\" aria-label=\"link to this section\">#</a></h2>\n<p>Size your pool from a load test, not from a blog post. It is smaller than you think. Use an external pooler in transaction mode if you have many application instances, and check your prepared statement behavior when you do.</p>",
   "image": "https://readme.news/cards/connection-pooling-explained-properly.png",
   "date_published": "2026-07-15T09:00:00Z",
   "date_modified": "2026-07-15T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "databases",
    "performance",
    "architecture"
   ]
  },
  {
   "id": "https://readme.news/the-onboarding-document-that-actually-works/",
   "url": "https://readme.news/the-onboarding-document-that-actually-works/",
   "title": "The onboarding document that actually works",
   "summary": "Most onboarding docs are written by people who already know. Here is what a new engineer needs on day one, in order.",
   "content_html": "<p>Onboarding documentation is written by someone who already knows the system, which means it is written from the wrong side of the knowledge gap.</p>\n<p>The result is consistently: an architecture overview that is meaningless without context, a list of tools with no explanation of why, and no answer to any question a new person actually has.</p>\n<h2 id=\"what-a-new-engineer-actually-needs-in-order\">what a new engineer actually needs, in order<a class=\"anchor\" href=\"#what-a-new-engineer-actually-needs-in-order\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Day one: get something running.</strong></p>\n<p>Not understanding. Running. A new engineer who has the application running locally on day one is in a completely different position from one who is still fighting a dependency on day three.</p>\n<p>This section should be a numbered list of commands that work. Tested. On a clean machine. Recently.</p>\n<div class=\"code\"><span class=\"code-lang\">markdown</span><pre><code class=\"lang-markdown\">## get it running\n\n1. Install prerequisites:  `brew bundle`   (or the equivalent \u2014 see below)\n2. Copy the environment:   `cp .env.example .env`\n3. Get secrets:            `./scripts/fetch-dev-secrets` (needs VPN)\n4. Start dependencies:     `docker compose up -d`\n5. Migrate:                `./scripts/db-reset`\n6. Run:                    `pnpm dev`\n\nYou should see the app at http://localhost:5173 with seeded data.\nLog in with dev@example.com / password.\n\nIf step 3 fails with \"unauthorized\", you need to be added to the\ndev-secrets group \u2014 ask in #eng-onboarding.</code></pre></div>\n<p>That last paragraph \u2014 the anticipated failure with its resolution \u2014 is the part that separates a document that works from one that does not. Every step that has ever failed for anyone should have a note.</p>\n<p><strong>Day two: make a change and see it.</strong></p>\n<p>A guided first change. Something trivial and real: change a label, add a field, fix a typo in a template. All the way through: edit, test, review, merge, deploy.</p>\n<p>The point is not the change. It is that they have now executed the entire delivery pipeline once, so every subsequent change is a variation on something they have done.</p>\n<p><strong>Day three to five: the map.</strong></p>\n<p><em>Now</em> the architecture overview, and now it means something, because they have seen the system run.</p>\n<p>Keep it to a page. What are the major pieces, what does each do, how do they talk. A diagram. Where the code for each piece lives.</p>\n<p>Not a complete description. A map, at the resolution of \"which building do I go to.\"</p>\n<p><strong>Week two: the why.</strong></p>\n<p>The decisions that are not obvious from the code. Why the database is structured that way. Why there is a queue between those two services. Why that module is frozen. Why the obvious approach to X does not work.</p>\n<p>This is the highest-value and least-written documentation in any organization, because it exists only in the heads of people who were there. When they leave, it is gone, and the next person spends a year rediscovering it \u2014 usually by proposing the obvious approach and being told no.</p>\n<h2 id=\"the-sections-everyone-forgets\">the sections everyone forgets<a class=\"anchor\" href=\"#the-sections-everyone-forgets\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>The glossary.</strong> Every organization has jargon: internal product names, acronyms, words used with a specific local meaning. A new person hears twenty of these in their first week and cannot ask about all of them without feeling stupid.</p>\n<p>Write them down. This is a thirty-minute task with an outsized payoff.</p>\n<p><strong>Who to ask about what.</strong> Not the org chart. \"Payments: ask Priya. Deploy pipeline: #platform-help. Anything about the legacy importer: Marcus, and be warned it is complicated.\"</p>\n<p><strong>The things that will surprise you.</strong> The test suite that fails on the first run until you seed a fixture. The service that takes four minutes to start. The one flaky test everyone knows about. The staging environment that resets on Sundays.</p>\n<p>Every codebase has these. Writing them down converts \"this is broken and I do not want to admit I cannot fix it\" into \"oh, that is expected.\"</p>\n<p><strong>What not to touch.</strong> Frozen modules, generated files, anything requiring a specific review.</p>\n<h2 id=\"the-maintenance-mechanism\">the maintenance mechanism<a class=\"anchor\" href=\"#the-maintenance-mechanism\" aria-label=\"link to this section\">#</a></h2>\n<p>Onboarding docs rot faster than any other documentation, because the people who would notice the errors are the ones who no longer read it.</p>\n<p><strong>The fix: the last person onboarded owns it.</strong></p>\n<p>Every new engineer's first task is to follow the document, fix everything that was wrong, and add the things they had to ask about. Then they own it until the next person arrives.</p>\n<p>This works because their frustration is fresh and their perspective is exactly the target reader's. It is the only mechanism I have seen that keeps these documents accurate.</p>\n<h2 id=\"the-test\">the test<a class=\"anchor\" href=\"#the-test\" aria-label=\"link to this section\">#</a></h2>\n<p>Hand it to a new engineer and do not help them.</p>\n<p>Time how long until they have the application running. Note every question they had to ask. Each question is a gap, with a measured cost.</p>\n<p>Then fix them, and repeat with the next person.</p>\n<p>Most teams have never done this and would be surprised by the result.</p>",
   "image": "https://readme.news/cards/the-onboarding-document-that-actually-works.png",
   "date_published": "2026-07-13T09:00:00Z",
   "date_modified": "2026-07-13T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "documentation",
    "engineering-management",
    "craft",
    "careers"
   ]
  },
  {
   "id": "https://readme.news/why-your-container-image-is-14-gigabytes/",
   "url": "https://readme.news/why-your-container-image-is-14-gigabytes/",
   "title": "Why your container image is 1.4 gigabytes",
   "summary": "It should be forty megabytes. Here is where the rest of it came from and how to get it back.",
   "content_html": "<p>A container image for a compiled service should be tens of megabytes. For an interpreted one, low hundreds. If yours is over a gigabyte, something specific went wrong and it is usually one of six things.</p>\n<p>Size matters for real reasons: pull time on cold start, registry cost, deployment speed when you are scaling out under load, and attack surface \u2014 every package in the image is something that can have a CVE you have to answer for.</p>\n<h2 id=\"the-six-causes\">the six causes<a class=\"anchor\" href=\"#the-six-causes\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>1. You shipped the build toolchain.</strong></p>\n<p>The compiler, the headers, the package manager cache, the source tree, the test fixtures. All of it needed to build, none of it needed to run.</p>\n<p>Multi-stage builds fix this completely:</p>\n<div class=\"code\"><span class=\"code-lang\">dockerfile</span><pre><code class=\"lang-dockerfile\">FROM golang:1.24 AS build\nWORKDIR /src\nCOPY go.mod go.sum ./\nRUN go mod download\nCOPY . .\nRUN CGO_ENABLED=0 go build -ldflags=\"-s -w\" -o /app ./cmd/server\n\nFROM gcr.io/distroless/static-debian12\nCOPY --from=build /app /app\nENTRYPOINT [\"/app\"]</code></pre></div>\n<p>Final image: the binary, plus CA certificates and timezone data. Tens of megabytes.</p>\n<p><strong>2. You started from a full distribution image.</strong></p>\n<p><code>FROM ubuntu</code> is roughly 80 MB before you install anything, and it includes a package manager, a shell, and a hundred utilities you will never invoke.</p>\n<p>The ladder, from largest to smallest:</p>\n<ul><li>Full distribution \u2014 80 MB+</li><li><code>-slim</code> variants \u2014 30\u201380 MB</li><li>Alpine \u2014 5\u201310 MB, with musl libc, which will occasionally surprise you</li><li>Distroless \u2014 just the runtime, no shell, no package manager</li><li><code>scratch</code> \u2014 nothing at all, for static binaries</li></ul>\n<p><strong>The Alpine caveat</strong>, since it bites people: musl's allocator and DNS resolver behave differently from glibc's. Python performance in particular can be significantly worse, and some binary wheels do not exist for musl. Test rather than assume.</p>\n<p><strong>3. Your layers are ordered wrong.</strong></p>\n<p>Each instruction creates a layer. A layer is invalidated when it or anything before it changes.</p>\n<div class=\"code\"><span class=\"code-lang\">dockerfile</span><pre><code class=\"lang-dockerfile\"># bad \u2014 any source change reinstalls every dependency\nCOPY . .\nRUN npm ci\n\n# good \u2014 dependencies are cached until the lockfile changes\nCOPY package.json package-lock.json ./\nRUN npm ci\nCOPY . .</code></pre></div>\n<p>This does not shrink the final image but it dramatically speeds up builds, which is usually what people actually care about.</p>\n<p><strong>4. You deleted things in a later layer.</strong></p>\n<div class=\"code\"><span class=\"code-lang\">dockerfile</span><pre><code class=\"lang-dockerfile\">RUN apt-get install -y build-essential   # layer 1: +400 MB\nRUN apt-get remove -y build-essential    # layer 2: marks deleted, image unchanged</code></pre></div>\n<p>Layers are additive. Deleting a file in a later layer hides it and does not remove it. The bytes are still in the image and still transferred on pull.</p>\n<p>Everything must happen in one <code>RUN</code>:</p>\n<div class=\"code\"><span class=\"code-lang\">dockerfile</span><pre><code class=\"lang-dockerfile\">RUN apt-get update \\\n &amp;&amp; apt-get install -y --no-install-recommends build-essential \\\n &amp;&amp; make \\\n &amp;&amp; apt-get purge -y build-essential \\\n &amp;&amp; apt-get autoremove -y \\\n &amp;&amp; rm -rf /var/lib/apt/lists/*</code></pre></div>\n<p>Better: use a multi-stage build and do not install the toolchain in the final image at all.</p>\n<p><strong>5. You have no <code>.dockerignore</code>.</strong></p>\n<p><code>COPY . .</code> copies <code>.git</code>, <code>node_modules</code>, build artifacts, test fixtures, and your local <code>.env</code>.</p>\n<div class=\"code\"><pre><code>.git\nnode_modules\ndist\n*.log\n.env*\n**/__pycache__\ncoverage</code></pre></div>\n<p>The <code>.git</code> directory alone is frequently hundreds of megabytes on a mature repository, and it is in a lot of images.</p>\n<p><strong>6. Your dependencies are enormous.</strong></p>\n<p>Sometimes it is genuinely the dependencies \u2014 machine learning stacks with CUDA libraries are legitimately multiple gigabytes.</p>\n<p>Check whether you need the GPU variant. <code>torch</code> with CUDA is roughly 2.5 GB; the CPU build is a fraction of that. If you are serving on CPU, you are shipping GPU libraries for nothing.</p>\n<h2 id=\"finding-out-where-it-went\">finding out where it went<a class=\"anchor\" href=\"#finding-out-where-it-went\" aria-label=\"link to this section\">#</a></h2>\n<div class=\"code\"><span class=\"code-lang\">bash</span><pre><code class=\"lang-bash\">docker history --no-trunc &lt;image&gt;       # size per layer</code></pre></div>\n<p>Or use a layer inspection tool that shows you which files are in which layer and how much space is wasted. Ten minutes with one of those tells you exactly what to fix.</p>\n<h2 id=\"the-security-dimension\">the security dimension<a class=\"anchor\" href=\"#the-security-dimension\" aria-label=\"link to this section\">#</a></h2>\n<p>Every package in the image is potential CVE surface, and your scanner will report all of them regardless of whether the code is reachable.</p>\n<p>A distroless image has almost nothing to report, which means the reports you do get are signal rather than noise. That is worth more than the size reduction \u2014 a vulnerability report with three entries gets read; one with four hundred does not.</p>\n<p>The trade-off: no shell means you cannot <code>docker exec</code> in to debug. Use ephemeral debug containers that attach to the running pod's namespaces instead, which is a better practice anyway because it means your production image is not a debugging toolkit.</p>\n<h2 id=\"the-target\">the target<a class=\"anchor\" href=\"#the-target\" aria-label=\"link to this section\">#</a></h2>\n<ul><li>Compiled language, static binary: <strong>under 30 MB.</strong></li><li>Interpreted with dependencies: <strong>under 200 MB.</strong></li><li>Anything over a gigabyte without a machine learning stack: something is wrong and it is one of the six above.</li></ul>",
   "image": "https://readme.news/cards/why-your-container-image-is-14-gigabytes.png",
   "date_published": "2026-07-10T09:00:00Z",
   "date_modified": "2026-07-10T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "devops",
    "performance",
    "tooling"
   ]
  },
  {
   "id": "https://readme.news/secrets-management-practically/",
   "url": "https://readme.news/secrets-management-practically/",
   "title": "Secrets management, practically",
   "summary": "Not a survey of vaults. The specific practices that actually reduce risk, in order of what to do first.",
   "content_html": "<p>Most secret management advice is a product comparison. Here is the practice instead, ordered by what to do first.</p>\n<h2 id=\"1-stop-long-lived-credentials-existing\">1. stop long-lived credentials existing<a class=\"anchor\" href=\"#1-stop-long-lived-credentials-existing\" aria-label=\"link to this section\">#</a></h2>\n<p>The single highest-value change, and it is architectural rather than a tool.</p>\n<p>A static credential can be stolen, leaked, committed, logged, or exfiltrated from a developer laptop. A credential that lives for fifteen minutes and is scoped to one operation cannot be usefully stolen.</p>\n<p><strong>Workload identity.</strong> Your CI job, your container, your function proves its identity to the cloud provider and receives a short-lived token. No stored secret at all.</p>\n<div class=\"code\"><span class=\"code-lang\">yaml</span><pre><code class=\"lang-yaml\"># GitHub Actions with OIDC \u2014 no stored cloud credentials\npermissions:\n  id-token: write\nsteps:\n  - uses: aws-actions/configure-aws-credentials@v4\n    with:\n      role-to-assume: arn:aws:iam::123456789012:role/deploy\n      aws-region: us-east-1</code></pre></div>\n<p>That eliminates the most commonly stolen class of credential entirely. If you do one thing from this article, do this one.</p>\n<p><strong>Database credentials</strong> can work the same way \u2014 IAM authentication, or dynamically generated credentials from a secret manager with a short lease.</p>\n<h2 id=\"2-keep-them-out-of-the-places-they-end-up\">2. keep them out of the places they end up<a class=\"anchor\" href=\"#2-keep-them-out-of-the-places-they-end-up\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Not in version control.</strong> Obvious, universally violated. Run a scanner in pre-commit and in CI. Both \u2014 pre-commit catches it before it happens, CI catches it when someone skipped the hook.</p>\n<p><strong>Once committed, assume compromised.</strong> Rewriting history does not help; the object is in every clone and in every fork. Rotate, then clean up.</p>\n<p><strong>Not in the container image.</strong> Layers are inspectable. <code>docker history</code> and a layer extraction tool will find it.</p>\n<p><strong>Not in environment variables, ideally.</strong> This is more controversial. Environment variables leak: into crash dumps, into child processes, into <code>/proc</code>, into logs when someone prints the environment for debugging, into error tracking services that capture context.</p>\n<p>Files with restrictive permissions, mounted at runtime, are better. Environment variables are convenient and are the pragmatic choice for many systems \u2014 just know what you are accepting.</p>\n<p><strong>Not in logs.</strong> Redact at the logger, with a deny-list of key names, not at each call site. Somebody will forget at a call site. The logger never forgets.</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">REDACT = {\"password\", \"token\", \"secret\", \"api_key\", \"authorization\", \"cookie\"}\n\ndef scrub(d):\n    return {k: (\"***\" if k.lower() in REDACT else v) for k, v in d.items()}</code></pre></div>\n<p><strong>Not in error tracking.</strong> Most error trackers capture local variables and request headers by default. Configure the redaction, then verify it by triggering a test error and reading what arrived.</p>\n<h2 id=\"3-rotate-and-test-the-rotation\">3. rotate, and test the rotation<a class=\"anchor\" href=\"#3-rotate-and-test-the-rotation\" aria-label=\"link to this section\">#</a></h2>\n<p>Rotation is only real if it has been executed. A rotation procedure that has never run is a document, not a control.</p>\n<p><strong>The property that makes rotation painless: support two valid credentials at once.</strong> Add the new one, deploy, verify, remove the old one. Without overlap, rotation requires downtime, so it does not happen.</p>\n<p>Design for this when you create the credential, not when you need to rotate it under incident pressure.</p>\n<h2 id=\"4-scope-narrowly\">4. scope narrowly<a class=\"anchor\" href=\"#4-scope-narrowly\" aria-label=\"link to this section\">#</a></h2>\n<p>A credential should permit exactly what its holder needs.</p>\n<ul><li>Read-only where writes are not needed.</li><li>One bucket, one prefix, not the account.</li><li>One database, one schema, one set of tables.</li><li>Time-bounded where possible.</li></ul>\n<p>The test: if this credential leaked, what could an attacker do? If the answer is \"anything,\" the scope is wrong regardless of how well you protect it.</p>\n<h2 id=\"5-know-what-you-have\">5. know what you have<a class=\"anchor\" href=\"#5-know-what-you-have\" aria-label=\"link to this section\">#</a></h2>\n<p>An inventory of every secret: what it is, where it lives, what it grants, who owns it, when it was last rotated.</p>\n<p>Most organizations cannot produce this, which means they cannot answer \"what do we rotate\" during an incident, which turns a two-hour response into a two-day one.</p>\n<h2 id=\"the-incident-procedure\">the incident procedure<a class=\"anchor\" href=\"#the-incident-procedure\" aria-label=\"link to this section\">#</a></h2>\n<p>When a secret is exposed, in this order:</p>\n<ol><li><strong>Rotate first.</strong> Before investigating, before cleanup, before the postmortem. Every minute the credential is valid is a minute of exposure.</li><li><strong>Then assess what it could reach</strong>, and check logs for use.</li><li><strong>Then clean up</strong> the exposure.</li><li><strong>Then work out how it happened.</strong></li></ol>\n<p>The common mistake is investigating first. While you investigate, the credential is live.</p>\n<h2 id=\"the-tooling-note\">the tooling note<a class=\"anchor\" href=\"#the-tooling-note\" aria-label=\"link to this section\">#</a></h2>\n<p>Every major cloud has a secret manager. They are all adequate. The dedicated tools add dynamic credential generation, fine-grained policy, and audit \u2014 genuinely useful at scale and not where to start.</p>\n<p><strong>Start with: OIDC workload identity for machine-to-machine, a cloud secret manager for what remains, scanning in CI, redaction in the logger, and a tested rotation procedure.</strong></p>\n<p>That covers the overwhelming majority of real-world credential compromise, and it is a week of work rather than a platform migration.</p>",
   "image": "https://readme.news/cards/secrets-management-practically.png",
   "date_published": "2026-07-08T09:00:00Z",
   "date_modified": "2026-07-08T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "security",
    "infrastructure",
    "engineering",
    "craft"
   ]
  },
  {
   "id": "https://readme.news/code-review-comments-that-change-things/",
   "url": "https://readme.news/code-review-comments-that-change-things/",
   "title": "Code review comments that change things",
   "summary": "Most review comments are noise or nitpicks. A small taxonomy of the ones that are worth writing.",
   "content_html": "<p>Most code review comments do not change the code, and of the ones that do, most change something that did not matter.</p>\n<p>Here is a taxonomy of comments worth writing, roughly in descending order of value.</p>\n<h2 id=\"the-ones-worth-writing\">the ones worth writing<a class=\"anchor\" href=\"#the-ones-worth-writing\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>\"This will break when X.\"</strong> The highest-value comment there is. A specific failure scenario the author did not consider.</p>\n<blockquote><p>\"If two requests hit this concurrently, both will pass the existence check and both will insert. We saw this exact bug in the invoicing path last year.\"</p></blockquote>\n<p>Concrete, falsifiable, and it comes with evidence.</p>\n<p><strong>\"This contradicts how we do it elsewhere.\"</strong> Consistency has real value and the author frequently does not know the precedent exists.</p>\n<blockquote><p>\"<code>orders/</code> uses the repository pattern for this. Worth matching, or is there a reason to differ here?\"</p></blockquote>\n<p>Note the question at the end. Sometimes there is a reason and you have just learned something.</p>\n<p><strong>\"I don't understand this.\"</strong> Underrated. If a reviewer with context cannot follow it, a stranger in two years will not either.</p>\n<p>This is not an admission of inadequacy. It is a measurement of the code's clarity, and it is a measurement only a reader can take.</p>\n<p><strong>\"What happens if this fails?\"</strong> The most consistently productive question in code review. Error paths are the least-considered part of most changes and the most-exercised part in production.</p>\n<p><strong>\"Is this the right layer?\"</strong> Business logic in a controller, presentation logic in a model, a database call in a template. Structural, cheap to fix now, expensive later.</p>\n<p><strong>\"This is good and here is why.\"</strong> Genuinely valuable and almost never written. Naming what worked teaches the pattern and it makes the critical comments land better, because they arrive from someone who is paying attention rather than someone who is looking for problems.</p>\n<h2 id=\"the-ones-not-worth-writing\">the ones not worth writing<a class=\"anchor\" href=\"#the-ones-not-worth-writing\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Anything a formatter or linter handles.</strong> If you are commenting on spacing, quotes, or import order, fix your tooling instead. A human enforcing mechanical rules is a broken process.</p>\n<p><strong>Style preferences without a reason.</strong> \"I would have used a map here\" is not a review comment, it is a preference. If there is a reason \u2014 clarity, performance, consistency with a convention \u2014 say the reason. If there is not, do not send it.</p>\n<p><strong>Speculative generality.</strong> \"What if we later need to support multiple currencies?\" Usually they will not, and building for it costs now. If you genuinely believe it, say what makes you believe it.</p>\n<p><strong>A redesign in a review comment.</strong> If the approach is fundamentally wrong, that is a conversation, not a comment thread. Comments are for improving an approach; a different approach needs a discussion, and doing it in review comments after the work is done is the most expensive possible time.</p>\n<p>That failure is on the process, not the reviewer: the design should have been discussed before implementation.</p>\n<h2 id=\"how-to-phrase-things\">how to phrase things<a class=\"anchor\" href=\"#how-to-phrase-things\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Distinguish blocking from non-blocking.</strong> A convention that removes an enormous amount of ambiguity:</p>\n<div class=\"code\"><pre><code>blocking: this deletes rows without the tenant filter\nsuggestion: this could use the existing helper in utils/dates\nquestion: is the retry here intentional given the caller already retries?\nnit: typo in the comment\npraise: nice \u2014 this handles the empty case correctly, which the old one did not</code></pre></div>\n<p>The author knows exactly what to act on. The reviewer can leave a thought without implying it must be addressed. Both people save time.</p>\n<p><strong>Ask rather than assert when you are unsure.</strong> \"Why does this need a lock?\" is better than \"this does not need a lock\" when you are not certain, and it is better even when you are, because the answer might teach you something about the system.</p>\n<p><strong>Explain the why, not just the what.</strong> \"Use a set here\" is an instruction. \"Use a set here \u2014 this is O(n\u00b2) on a list that can have thousands of entries\" is teaching, and the author will apply it next time without being told.</p>\n<p><strong>Comment on the code, not the person.</strong> \"This is confusing\" rather than \"you wrote this confusingly.\" Small difference, and it consistently changes how the comment is received.</p>\n<h2 id=\"the-process-problems-that-comments-cannot-fix\">the process problems that comments cannot fix<a class=\"anchor\" href=\"#the-process-problems-that-comments-cannot-fix\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>The diff is too large.</strong> Past a few hundred lines, review quality collapses. The comment \"this PR is too large, can you split it\" is the most valuable one available and it is socially costly to write, which is why nobody does.</p>\n<p>Make it a policy so it is not a personal judgment.</p>\n<p><strong>Review is too late.</strong> If a fundamental problem is found in review, the process failed earlier. That belongs in design.</p>\n<p><strong>Only one person reviews.</strong> Different reviewers see different things. For anything significant, two, with different backgrounds.</p>\n<p><strong>Comments arrive over three days.</strong> The author has moved on and has to reload the entire context. Batch your review into one pass, and do it within a day.</p>\n<h2 id=\"the-goal\">the goal<a class=\"anchor\" href=\"#the-goal\" aria-label=\"link to this section\">#</a></h2>\n<p>The purpose of code review is not to find bugs \u2014 tests find bugs more reliably and more cheaply.</p>\n<p>It is to spread understanding, maintain coherence, and catch the class of problem that automation cannot see: wrong abstraction, wrong layer, missing case, code that will confuse the next reader.</p>\n<p>Comments that serve those are worth writing. Everything else is noise with a notification attached.</p>",
   "image": "https://readme.news/cards/code-review-comments-that-change-things.png",
   "date_published": "2026-07-06T09:00:00Z",
   "date_modified": "2026-07-06T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "code-review",
    "craft",
    "communication",
    "engineering"
   ]
  },
  {
   "id": "https://readme.news/backpressure-is-the-concept-your-system-is-missing/",
   "url": "https://readme.news/backpressure-is-the-concept-your-system-is-missing/",
   "title": "Backpressure is the concept your system is missing",
   "summary": "When a fast producer meets a slow consumer, something has to give. Deciding what, in advance, is the whole discipline.",
   "content_html": "<p>Every system with a producer and a consumer eventually has a moment where the producer is faster. What happens next is either a design decision you made or an emergent behavior you discover during an incident.</p>\n<h2 id=\"the-four-options\">the four options<a class=\"anchor\" href=\"#the-four-options\" aria-label=\"link to this section\">#</a></h2>\n<p>There are only four. Every system picks one, explicitly or by accident.</p>\n<p><strong>1. Buffer.</strong> Queue the excess.</p>\n<p>Works for bursts. Fails for sustained overload, because a buffer is a delay, and an unbounded buffer is a memory leak with a friendly name. The failure mode is that memory grows until the process dies, taking the buffer with it \u2014 so you lose everything, at the worst possible moment.</p>\n<p><strong>2. Drop.</strong> Discard the excess.</p>\n<p>Correct more often than people are comfortable with. Metrics, logs, telemetry, non-critical events \u2014 dropping 5% of samples under load is fine and dying is not.</p>\n<p>The requirement: <strong>know that you dropped, and how much.</strong> Silent drops are how you get a dashboard that looks healthy while data is missing.</p>\n<p><strong>3. Block.</strong> Make the producer wait.</p>\n<p>This is real backpressure. The consumer's slowness propagates upstream, the producer slows down, and the system reaches equilibrium at the consumer's rate.</p>\n<p>Correct for internal pipelines where the producer can wait. Dangerous when the producer is a user-facing request handler, because now user requests are blocked on a background process.</p>\n<p><strong>4. Reject.</strong> Tell the producer no.</p>\n<p>The right answer at a service boundary. A 429 or 503 returned in one millisecond is much better than a request that waits thirty seconds and then times out, because the caller can make a decision \u2014 retry later, degrade, or tell the user.</p>\n<h2 id=\"the-wrong-default\">the wrong default<a class=\"anchor\" href=\"#the-wrong-default\" aria-label=\"link to this section\">#</a></h2>\n<p>Most systems buffer by default, unboundedly, without anyone deciding.</p>\n<ul><li>An in-memory list that grows.</li><li>A queue with no maximum length.</li><li>A connection pool that queues waiters forever.</li><li>A channel with a very large capacity, which is unbounded in practice.</li></ul>\n<p>The failure mode is always the same: latency grows, memory grows, and then the process dies. And the requests you were holding were abandoned by their callers ten seconds earlier, so all that work was for nothing.</p>\n<p><strong>A queue that is always full is not a buffer. It is a delay you cannot see.</strong></p>\n<h2 id=\"the-practical-rules\">the practical rules<a class=\"anchor\" href=\"#the-practical-rules\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Every queue has a maximum size.</strong> Every one. Choose the number by asking: how long should a request be willing to wait? Multiply by the consumer's rate. That is your queue depth.</p>\n<p>If you cannot answer that question, the queue is not designed.</p>\n<p><strong>Prefer rejecting to queueing at the edge.</strong> When a request arrives and the system is saturated, reject fast. The caller has a timeout; use it as your budget.</p>\n<p><strong>Propagate deadlines.</strong> If the caller has 200 ms left, every downstream operation should know that. Work performed after the caller has given up is pure waste, and under overload it is the majority of the work being done.</p>\n<div class=\"code\"><span class=\"code-lang\">go</span><pre><code class=\"lang-go\">ctx, cancel := context.WithTimeout(ctx, remaining)\ndefer cancel()</code></pre></div>\n<p><strong>Shed by priority, not randomly.</strong> Under load, serve health checks, serve authenticated users, serve the critical path. Shed background work, analytics, and prefetches. Random shedding means your health checks fail and your orchestrator kills healthy instances, which is a self-inflicted outage.</p>\n<p>**Measure queue <em>age</em>, not depth.** Depth tells you how many. Age tells you how far behind. \"The oldest item has been waiting four minutes\" is actionable in a way that \"there are 30,000 items\" is not.</p>\n<h2 id=\"the-pattern-that-ties-it-together\">the pattern that ties it together<a class=\"anchor\" href=\"#the-pattern-that-ties-it-together\" aria-label=\"link to this section\">#</a></h2>\n<p>Little's Law: <code>L = \u03bbW</code>. Items in the system equals arrival rate times time in system.</p>\n<p>Rearranged: <strong>wait time equals queue length divided by service rate.</strong></p>\n<p>If your queue holds 10,000 items and you process 100 per second, the newest item waits 100 seconds. That is not a hypothetical \u2014 it is arithmetic, and it means your queue length choice <em>is</em> your latency choice, whether or not you framed it that way.</p>\n<p>Pick the latency you can accept, multiply by the service rate, and that is your maximum queue length. Everything past it gets rejected.</p>\n<h2 id=\"the-test\">the test<a class=\"anchor\" href=\"#the-test\" aria-label=\"link to this section\">#</a></h2>\n<p>Point a load generator at your service at three times its capacity. Watch:</p>\n<ul><li>Does memory grow without bound?</li><li>Does latency grow without bound?</li><li>Does anything get rejected, or does everything just get slower?</li><li>After you stop the load, how long until it recovers?</li></ul>\n<p>That last one is the important one. A system that takes twenty minutes to recover from a two-minute overload has a backpressure problem, and it will turn a small incident into a large one on a day you did not choose.</p>",
   "image": "https://readme.news/cards/backpressure-is-the-concept-your-system-is-missing.png",
   "date_published": "2026-07-03T09:00:00Z",
   "date_modified": "2026-07-03T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "architecture",
    "reliability",
    "engineering"
   ]
  },
  {
   "id": "https://readme.news/the-crawler-tolls-one-year-on/",
   "url": "https://readme.news/the-crawler-tolls-one-year-on/",
   "title": "The crawler tolls, one year on",
   "summary": "Default blocking and pay-per-crawl changed who can read the web. An assessment of what actually happened.",
   "content_html": "<p>A year ago today a CDN sitting in front of a large fraction of the web flipped its default: AI crawlers blocked unless explicitly allowed, with a marketplace for charging per crawl.</p>\n<p>Enough time has passed to say what actually happened rather than what everyone predicted.</p>\n<h2 id=\"what-changed\">what changed<a class=\"anchor\" href=\"#what-changed\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>The norm inverted.</strong> Before, crawling was permitted by default and <code>robots.txt</code> was a request. Now, for a large share of the web, crawling is denied by default and access is a negotiation.</p>\n<p>That is a genuine structural change to how the web works and it happened through one company's configuration default rather than through any standards process, legislation, or public deliberation.</p>\n<p><strong>Licensing deals concentrated.</strong> Large AI companies negotiated bulk access with large publishers. That was always the likely outcome: the parties with lawyers and leverage made arrangements, and the arrangements are private.</p>\n<p><strong>Small publishers got very little.</strong> The pay-per-crawl mechanism works, technically. The revenue for a site with modest traffic is negligible \u2014 the arithmetic never supported anything else. The publishers who most needed a new economic model got the one that pays least.</p>\n<p><strong>Non-commercial crawling got harder.</strong> Academic researchers, archivists, and independent search projects have no licensing budget and no negotiating position. Carve-outs exist and are discretionary, which means the ability to study the web is now something you apply for.</p>\n<p>That is the outcome I was most worried about and it is the one that materialized most clearly.</p>\n<h2 id=\"what-did-not-change\">what did not change<a class=\"anchor\" href=\"#what-did-not-change\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Training data supply.</strong> The frontier labs have enormous existing corpora, licensed sources, and synthetic generation. The marginal value of newly crawled web text was already declining. Restricting it did not create the leverage publishers hoped for.</p>\n<p><strong>Traffic.</strong> Referral traffic from search to publishers continued its decline, driven by AI answers in search results, which is a completely separate mechanism from training crawlers and which blocking crawlers does nothing about.</p>\n<p>This is the part that was most misunderstood at the time. The traffic problem and the training problem have different causes and blocking crawlers only addresses one of them \u2014 the one with less economic impact.</p>\n<h2 id=\"the-thing-to-actually-take-from-it\">the thing to actually take from it<a class=\"anchor\" href=\"#the-thing-to-actually-take-from-it\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Infrastructure defaults are policy.</strong> A configuration change at a chokepoint reshaped access to a large fraction of the web, with no process and no appeal.</p>\n<p>That is not a criticism of the specific decision, which was popular and defensible. It is an observation about where power actually sits, and it generalizes: the entities that can change the web's behavior are the ones with concentration at a layer everyone depends on, and there are about five of them.</p>\n<p><strong>For your own site</strong>, the decision remains yours and it is worth making deliberately rather than accepting a default:</p>\n<ul><li><strong>Documentation sites</strong> frequently want to be in the training data. Being the thing the model knows about is worth more than the pageview you did not get.</li><li><strong>Original reporting and analysis</strong> has a stronger case for restriction.</li><li><strong>Anything you want found</strong> should still permit search crawlers, which are a different category and are frequently blocked by accident when people configure this.</li></ul>\n<p>Check what you are actually blocking. A meaningful number of sites blocked their own search indexing in the first months of this and did not notice for weeks.</p>\n<h2 id=\"the-unresolved-thing\">the unresolved thing<a class=\"anchor\" href=\"#the-unresolved-thing\" aria-label=\"link to this section\">#</a></h2>\n<p>The web's economic model \u2014 publish freely, get traffic, monetize traffic \u2014 is breaking, and nothing has replaced it.</p>\n<p>Crawler tolls are not the replacement; the arithmetic does not work at the scale of the actual web, where most content is made by people with no ability to negotiate anything.</p>\n<p>Licensing deals are not the replacement either; they work for a few hundred large publishers and for nobody else.</p>\n<p>I do not know what the replacement is. I am increasingly convinced that nobody does, and that the interval between the old model failing and a new one existing is going to be long and is going to be bad for the open web.</p>\n<p>That is a genuinely pessimistic conclusion and I have not found a way around it in a year of thinking about it.</p>",
   "image": "https://readme.news/cards/the-crawler-tolls-one-year-on.png",
   "date_published": "2026-07-01T09:00:00Z",
   "date_modified": "2026-07-01T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "frontend",
    "ai",
    "policy",
    "infrastructure",
    "review"
   ]
  },
  {
   "id": "https://readme.news/cross-platform-is-a-promise-you-make-to-your-budget/",
   "url": "https://readme.news/cross-platform-is-a-promise-you-make-to-your-budget/",
   "title": "Cross-platform is a promise you make to your budget",
   "summary": "Write once, run anywhere, debug everywhere. The honest accounting of what each approach actually costs.",
   "content_html": "<p>The cross-platform question \u2014 one codebase or several \u2014 gets argued as a technical matter and is mostly an organizational one. Here is the honest accounting.</p>\n<h2 id=\"the-actual-trade\">the actual trade<a class=\"anchor\" href=\"#the-actual-trade\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Native</strong> gives you: full platform capability, best performance, platform-idiomatic interface, immediate access to new OS features, and the best debugging tools.</p>\n<p>It costs you: two or three codebases, two or three teams, features implemented multiple times, and behavior that diverges over years in ways nobody tracks.</p>\n<p><strong>Cross-platform</strong> gives you: one codebase, one team, features implemented once, consistent behavior.</p>\n<p>It costs you: a framework layer between you and the platform, a lag before new OS features are available, worse debugging when the problem is in the bridge, and an interface that is either non-idiomatic on every platform or requires platform-specific work anyway.</p>\n<h2 id=\"the-thing-the-arguments-miss\">the thing the arguments miss<a class=\"anchor\" href=\"#the-thing-the-arguments-miss\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>The largest cost of cross-platform is not performance. It is the escape hatch.</strong></p>\n<p>Everything is fine until you need something the framework does not support. Then you are writing platform-specific native code, plus a bridge, plus a fallback, plus tests for all three \u2014 and you have the complexity of native development <em>plus</em> the framework.</p>\n<p>This happens. It always happens. The question is how often, and that depends entirely on what your app does.</p>\n<p><strong>Low escape-hatch pressure:</strong> content, commerce, forms, dashboards, CRUD, most business applications. These use the standard widget set and standard capabilities. Cross-platform works well.</p>\n<p><strong>High escape-hatch pressure:</strong> camera and media processing, background execution, Bluetooth and hardware peripherals, complex custom rendering, deep platform integration (widgets, shortcuts, app extensions), anything real-time.</p>\n<p>For high-pressure apps, cross-platform frequently costs more than native, because you pay for the framework and then write the native code anyway.</p>\n<p><strong>Be honest about which one you are</strong> before choosing. Most teams that regret their choice were high-pressure and assessed themselves as low.</p>\n<h2 id=\"the-options-honestly\">the options, honestly<a class=\"anchor\" href=\"#the-options-honestly\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>React Native.</strong> Mature, large ecosystem, native widgets. The new architecture removed the old asynchronous bridge, which was the main performance complaint. Best choice if your team is already React.</p>\n<p><strong>Flutter.</strong> Renders its own widgets, which means true visual consistency and a non-native feel that some users notice and most do not. Excellent performance for custom interfaces. Dart is a real adoption cost for a team that does not know it.</p>\n<p><strong>Kotlin Multiplatform.</strong> Share business logic, write native UI. This is the approach I find most defensible: the logic layer \u2014 networking, models, validation, persistence \u2014 is where duplication is most wasteful and least visible to users, and the UI layer is where platform idiom matters most.</p>\n<p>Requires two UI implementations, which is the point rather than a limitation.</p>\n<p><strong>Web technologies in a wrapper.</strong> Fastest to build if you have web engineers, and the platform feel is the weakest. Fine for content-heavy applications, poor for anything interaction-heavy.</p>\n<p><strong>Native.</strong> Still correct for a large category, and the category is smaller than native advocates believe.</p>\n<h2 id=\"the-organizational-question-that-actually-decides-it\">the organizational question that actually decides it<a class=\"anchor\" href=\"#the-organizational-question-that-actually-decides-it\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Do you have or can you hire two platform teams?</strong></p>\n<p>If yes, native is viable and gives you the best result.</p>\n<p>If no \u2014 and for most companies below a certain size the answer is no \u2014 the choice is between cross-platform and shipping on one platform. Framed that way, the decision is usually easy.</p>\n<p><strong>What is your feature velocity?</strong></p>\n<p>If you ship a large feature monthly, implementing it twice is a permanent 2\u00d7 cost on your most expensive activity. If you ship quarterly, the duplication matters less.</p>\n<p><strong>How much does platform idiom matter to your users?</strong></p>\n<p>For a consumer app competing on polish: a lot. For an internal tool: nothing. For a B2B product where the buyer is not the user: less than you think.</p>\n<h2 id=\"the-hybrid-that-most-people-should-consider\">the hybrid that most people should consider<a class=\"anchor\" href=\"#the-hybrid-that-most-people-should-consider\" aria-label=\"link to this section\">#</a></h2>\n<p>Share the logic, write the UI natively.</p>\n<p>The business logic \u2014 API clients, data models, validation, offline storage, sync, analytics \u2014 is genuinely identical across platforms and duplicating it produces bugs that exist on one platform and not the other, which are the worst bugs to diagnose.</p>\n<p>The UI is where platform conventions matter, where users notice, and where the framework abstraction costs the most.</p>\n<p>This is more work than full cross-platform and less than full native, and it puts the sharing where the value is.</p>\n<h2 id=\"the-thing-i-would-tell-someone-deciding\">the thing I would tell someone deciding<a class=\"anchor\" href=\"#the-thing-i-would-tell-someone-deciding\" aria-label=\"link to this section\">#</a></h2>\n<p>Prototype the hardest thing first.</p>\n<p>Not the login screen. The thing you are worried about \u2014 the camera flow, the background sync, the complex list, the offline behavior. Build that on your candidate stack, in a week.</p>\n<p>You will learn more from that week than from any amount of comparison, and you will learn it while changing your mind is still cheap.</p>\n<p>The teams that regret their choice almost always chose based on a comparison article, built the easy part first, and discovered the hard part in month five.</p>",
   "image": "https://readme.news/cards/cross-platform-is-a-promise-you-make-to-your-budget.png",
   "date_published": "2026-06-30T09:00:00Z",
   "date_modified": "2026-06-30T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "platforms",
    "architecture",
    "engineering",
    "opinion"
   ]
  },
  {
   "id": "https://readme.news/retries-a-complete-guide-to-not-making-it-worse/",
   "url": "https://readme.news/retries-a-complete-guide-to-not-making-it-worse/",
   "title": "Retries: a complete guide to not making it worse",
   "summary": "The most common way a small incident becomes a large one is a retry policy written without thinking about aggregate behavior.",
   "content_html": "<p>Retries are the most commonly implemented and most commonly wrong piece of resilience engineering. A retry policy that seems obviously correct in isolation is frequently the mechanism that turns a brief degradation into a full outage.</p>\n<h2 id=\"the-failure-mode\">the failure mode<a class=\"anchor\" href=\"#the-failure-mode\" aria-label=\"link to this section\">#</a></h2>\n<p>A downstream service slows down. Every caller times out. Every caller retries.</p>\n<p>The downstream now receives double its normal traffic while already struggling. More requests time out. More retries. The load multiplies with every round.</p>\n<p>The original problem might have been a five-second blip. The retry storm keeps the service down for twenty minutes, and it stays down after the original cause is resolved because the queued retries are still arriving.</p>\n<p>This is metastable failure and retries are its most common cause.</p>\n<h2 id=\"the-rules\">the rules<a class=\"anchor\" href=\"#the-rules\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>1. Only retry idempotent operations.</strong></p>\n<p>A retried non-idempotent operation charges the card twice. If you need to retry a write, make it idempotent first with an idempotency key.</p>\n<p><strong>2. Only retry retryable errors.</strong></p>\n<p>A 400 will be a 400 next time. Retrying it wastes a request and delays the error the caller needs to see.</p>\n<div class=\"table-wrap\"><table><thead><tr><th style=\"text-align:left\">retry</th><th style=\"text-align:left\">do not retry</th></tr></thead><tbody><tr><td style=\"text-align:left\">connection refused / reset</td><td style=\"text-align:left\">400, 401, 403, 404</td></tr><tr><td style=\"text-align:left\">timeout</td><td style=\"text-align:left\">422</td></tr><tr><td style=\"text-align:left\">502, 503, 504</td><td style=\"text-align:left\">any deterministic validation failure</td></tr><tr><td style=\"text-align:left\">429 (respect <code>Retry-After</code>)</td><td style=\"text-align:left\">501</td></tr></tbody></table></div>\n<p>The one people get wrong: <strong>500 is ambiguous.</strong> It might be transient, it might be a deterministic bug. Retrying it once is usually reasonable; retrying it five times is usually pointless.</p>\n<p><strong>3. Exponential backoff with jitter. Always jitter.</strong></p>\n<p>Without jitter, all your clients retry at the same moments, and you have built a synchronized load generator.</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">def delay(attempt, base=0.1, cap=30):\n    return random.uniform(0, min(cap, base * (2 ** attempt)))</code></pre></div>\n<p>That is full jitter, and it is the recommended default. It spreads retries across the whole interval, which is what you want. Half jitter \u2014 <code>d/2 + random(0, d/2)</code> \u2014 is a reasonable alternative when you want a guaranteed minimum delay.</p>\n<p>The version without jitter is the one everyone writes first and it is the one that causes the storm.</p>\n<p><strong>4. Cap the attempts and cap the total time.</strong></p>\n<p>Three attempts, usually. And a total deadline \u2014 if the caller is going to give up after two seconds, retrying at three seconds is pure waste and it is load on a struggling service.</p>\n<p>Propagate the deadline. If your caller has 500 ms left, your retry budget is 500 ms, not your configured default.</p>\n<p><strong>5. Do not retry at every layer.</strong></p>\n<p>This is the one that produces the shocking numbers. Three attempts at the HTTP client, three in the service wrapper, three in the caller, three at the gateway: 3\u2074 = 81 requests for one logical call.</p>\n<p><strong>Retry at exactly one layer.</strong> Usually the outermost one that has the context to decide. Every other layer fails fast and propagates.</p>\n<p>Audit this. Most systems that have grown organically retry at three or four layers and nobody knows.</p>\n<p><strong>6. Use a retry budget.</strong></p>\n<p>The refinement that actually prevents storms: cap retries as a <em>fraction of total traffic</em>, not per request.</p>\n<div class=\"code\"><pre><code>if retries_in_window / requests_in_window &gt; 0.1:\n    do_not_retry()</code></pre></div>\n<p>When things are healthy, occasional retries are well under the budget and everything works. When things are broken, the budget is exhausted immediately and retries stop entirely \u2014 exactly when they would do the most harm.</p>\n<p>This single mechanism converts retries from a failure amplifier into a bounded safety net, and it is not widely implemented.</p>\n<p><strong>7. Circuit break.</strong></p>\n<p>When a downstream is clearly failing, stop calling it. Fail fast, return a cached or degraded response, and probe occasionally to see if it has recovered.</p>\n<p>Three states: closed (normal), open (failing fast), half-open (probing). The half-open state must allow only a trickle \u2014 if you send full traffic at a recovering service you will knock it over again.</p>\n<h2 id=\"the-server-side\">the server side<a class=\"anchor\" href=\"#the-server-side\" aria-label=\"link to this section\">#</a></h2>\n<p>The other half, which is usually forgotten.</p>\n<p><strong>Send <code>Retry-After</code> on 429 and 503.</strong> Then clients that respect it retry at a time you chose rather than a time they chose.</p>\n<p><strong>Shed load rather than queueing it.</strong> A request that will time out anyway should be rejected immediately, not queued. Queueing under overload increases latency for everything without increasing throughput, and the requests you eventually serve have often already been abandoned.</p>\n<p><strong>Prioritize.</strong> Under load, serve health checks and critical paths, shed the rest. An unprioritized overload sheds randomly, which means your health checks fail and your orchestrator kills healthy instances.</p>\n<h2 id=\"the-test\">the test<a class=\"anchor\" href=\"#the-test\" aria-label=\"link to this section\">#</a></h2>\n<p>Take your service. Make a downstream dependency return 503 for everything. Watch the request rate at the downstream.</p>\n<p>If it goes up by more than a small factor, your retry configuration will cause an outage. You have just not had the trigger yet.</p>",
   "image": "https://readme.news/cards/retries-a-complete-guide-to-not-making-it-worse.png",
   "date_published": "2026-06-29T09:00:00Z",
   "date_modified": "2026-06-29T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "reliability",
    "architecture",
    "engineering"
   ]
  },
  {
   "id": "https://readme.news/half-of-2026-in-one-page/",
   "url": "https://readme.news/half-of-2026-in-one-page/",
   "title": "Half of 2026, in one page",
   "summary": "Six months of regulation, agent tooling that mostly did not fix the bottleneck, and a web platform that quietly finished.",
   "content_html": "<p>The first half, compressed.</p>\n<h2 id=\"the-three-things-that-actually-happened\">the three things that actually happened<a class=\"anchor\" href=\"#the-three-things-that-actually-happened\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Regulation became an engineering concern.</strong> The EU AI Act's high-risk obligations arrive in August; the Cyber Resilience Act's reporting obligations in September; post-quantum migration guidance firmed up across national agencies. For the first time in a while, compliance is showing up on engineering roadmaps rather than only in legal reviews.</p>\n<p><strong>Power stayed the binding constraint on compute.</strong> Interconnection queues, transformer lead times, and utility rate cases determine how fast capacity arrives. Every roadmap presented this year is downstream of a construction schedule.</p>\n<p><strong>The verification bottleneck did not move.</strong> Generation capability kept improving. Review capacity did not. The tooling investment continued to go into producing more code rather than into checking it, which is the wrong end of the problem and which I predicted would correct by now. It has not.</p>\n<h2 id=\"the-thing-i-keep-coming-back-to\">the thing I keep coming back to<a class=\"anchor\" href=\"#the-thing-i-keep-coming-back-to\" aria-label=\"link to this section\">#</a></h2>\n<p>Generation got roughly two orders of magnitude cheaper. Verification got no cheaper at all.</p>\n<p>Everything else \u2014 the review queue, the junior hiring contraction, the sudden importance of tests and type systems, the unease people cannot name about agent-written code \u2014 is a consequence of that asymmetry, and six months of watching it has not changed my read.</p>\n<p>The skill that matters is the ability to look at plausible output and tell whether it is right. That skill is built by doing the work, and the work is being automated.</p>\n<p>Nobody has solved this. Most of the discussion still treats it as either a non-problem or an apocalypse, and it is neither. It is an engineering and training problem with no established answer yet.</p>\n<h2 id=\"pieces-i-would-re-read\">pieces I would re-read<a class=\"anchor\" href=\"#pieces-i-would-re-read\" aria-label=\"link to this section\">#</a></h2>\n<ul><li><strong>\"Verification is the whole job now\"</strong> (8 Jan) \u2014 the argument, stated once, properly.</li><li><strong>\"Prompt injection is SQL injection without the fix\"</strong> (9 Feb) \u2014 why the analogy everyone makes stops before the part that matters.</li><li><strong>\"Schema design is the only design that lasts\"</strong> (6 Mar) \u2014 the thing you will still be living with in 2036.</li><li><strong>\"Agent fleets in production\"</strong> (27 Mar) \u2014 the honest numbers, which are much lower than the demos.</li><li><strong>\"Idempotency is the only distributed systems concept you need\"</strong> (26 May) \u2014 if you internalize one thing.</li></ul>\n<h2 id=\"the-grades-so-far\">the grades so far<a class=\"anchor\" href=\"#the-grades-so-far\" aria-label=\"link to this section\">#</a></h2>\n<p>In January I made five predictions. Three months in I graded one as wrong. At the halfway point:</p>\n<p><strong>Review tooling as a major funded category</strong> \u2014 still not happening at the scale I predicted. Funding continues to flow to generation. I was wrong about the market, and possibly wrong about the timing rather than the direction.</p>\n<p><strong>Repository-level agent configuration as standard practice</strong> \u2014 correct, and it happened faster than I expected. This is now table stakes for a well-maintained repository.</p>\n<p><strong>Power as a developer-visible line item</strong> \u2014 correct and accelerating. Regional compute price differentials are now large enough to affect architecture.</p>\n<p><strong>Supply chain controls with real teeth</strong> \u2014 in progress. The registry-level changes are landing and the migration is as painful as expected.</p>\n<p><strong>\"Senior engineer\" redefining around judgment</strong> \u2014 correct, uncomfortable, and nobody has a good answer for how the next generation acquires it.</p>\n<h2 id=\"what-i-am-watching-for-the-second-half\">what I am watching for the second half<a class=\"anchor\" href=\"#what-i-am-watching-for-the-second-half\" aria-label=\"link to this section\">#</a></h2>\n<ul><li>The August AI Act deadline: holds, slips, or gets phased.</li><li>Whether anyone ships a genuinely better code review interface.</li><li>Whether the small-model tier keeps compressing the frontier's addressable workload.</li><li>Whether junior hiring recovers, and what happens to the pipeline if it does not.</li></ul>\n<p>Back on Monday.</p>",
   "image": "https://readme.news/cards/half-of-2026-in-one-page.png",
   "date_published": "2026-06-26T09:00:00Z",
   "date_modified": "2026-06-26T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "retrospective"
   ]
  },
  {
   "id": "https://readme.news/configuration-is-the-most-under-designed-part-of-your-system/",
   "url": "https://readme.news/configuration-is-the-most-under-designed-part-of-your-system/",
   "title": "Configuration is the most under-designed part of your system",
   "summary": "It has no type system, no tests, no review, and it causes a disproportionate share of outages.",
   "content_html": "<p>Look at the last ten significant outages you can remember reading about. A disproportionate number were caused by configuration, not code.</p>\n<p>A config file with a wrong value. A feature flag flipped. A generated file that doubled in size. A DNS record. A permission change.</p>\n<p>Configuration receives a fraction of the engineering rigor that code does, and it has comparable power to break things.</p>\n<h2 id=\"why-it-goes-wrong\">why it goes wrong<a class=\"anchor\" href=\"#why-it-goes-wrong\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>No type system.</strong> A YAML file will happily contain <code>timeout: 30s</code> where the code expects an integer, or <code>enabled: \"false\"</code> which is a truthy string.</p>\n<p><strong>No tests.</strong> Nobody writes a test for their config.</p>\n<p><strong>No review, or perfunctory review.</strong> A config change is \"just a value\" and gets approved in ten seconds.</p>\n<p><strong>Deployed differently from code.</strong> Frequently faster, frequently without staging, frequently without a canary. The safety mechanisms built for code deployment routinely do not cover config.</p>\n<p><strong>Environment drift.</strong> Staging and production differ in ways nobody has enumerated, so staging validates a configuration that is not production's.</p>\n<p><strong>No rollback story.</strong> \"What was this value before\" is often unanswerable.</p>\n<h2 id=\"the-fixes\">the fixes<a class=\"anchor\" href=\"#the-fixes\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>1. Parse and validate at startup, not at use.</strong></p>\n<p>The worst failure mode is a config error that manifests three hours later when a rarely-used code path reads a malformed value.</p>\n<p>Validate everything at boot. Fail loudly and immediately if anything is wrong.</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">class Settings(BaseModel):\n    database_url: PostgresDsn\n    timeout_ms: int = Field(gt=0, le=60_000)\n    max_connections: int = Field(ge=1, le=1000)\n    feature_new_checkout: bool = False\n\nsettings = Settings(**load_config())   # raises at startup, with a clear message</code></pre></div>\n<p>An application that will not start is far better than one that starts and behaves wrongly.</p>\n<p><strong>2. Types, with a schema.</strong></p>\n<p>Whatever your language, there is a library that turns untyped config into a validated typed object. Use it. This eliminates the entire category of string-that-should-be-a-number bugs.</p>\n<p><strong>3. Config changes go through the same pipeline as code.</strong></p>\n<p>Version control. Review. Staging. Canary. Rollback.</p>\n<p>The argument against is that config changes need to be fast, especially for incident response. That is a real need and the answer is a small explicitly-defined set of emergency levers \u2014 kill switches, rate limits \u2014 with a fast path, and everything else on the normal pipeline.</p>\n<p>Not \"all config is fast\" because that is how config takes down your service.</p>\n<p><strong>4. Validate generated configuration.</strong></p>\n<p>If a config file is produced by a program, that program can be wrong. A size check against the previous version, a schema check, a sanity check on record count.</p>\n<p>This is the specific failure that has caused several high-profile outages: a generated file changed unexpectedly and propagated globally before anyone looked at it.</p>\n<p><strong>5. Make the environment explicit and diff it.</strong></p>\n<p>You should be able to answer \"how does staging differ from production\" with a command. If you cannot, staging is not validating production.</p>\n<div class=\"code\"><span class=\"code-lang\">bash</span><pre><code class=\"lang-bash\">config-diff staging production</code></pre></div>\n<p><strong>6. Log the effective configuration at startup.</strong></p>\n<p>Not the file \u2014 the resolved values after defaults, overrides, and environment variables are applied. Redact secrets. This is the single most useful thing for debugging \"it works on my machine,\" because the effective config is frequently not what anyone thinks it is.</p>\n<p><strong>7. Feature flags need the same discipline as code.</strong></p>\n<p>Owner. Expiry. Test both branches. Log the flag state on every event. A flag flip is a production change and should be treated as one.</p>\n<h2 id=\"the-hierarchy-that-works\">the hierarchy that works<a class=\"anchor\" href=\"#the-hierarchy-that-works\" aria-label=\"link to this section\">#</a></h2>\n<p>Most systems end up with layered configuration and the layering should be explicit and simple:</p>\n<div class=\"code\"><pre><code>defaults in code\n  \u2190 config file\n    \u2190 environment variables\n      \u2190 command line flags</code></pre></div>\n<p>Later overrides earlier. Log which layer each effective value came from when debugging.</p>\n<p><strong>Keep the layers few.</strong> Systems with six overlapping sources of configuration \u2014 defaults, file, environment, a service, a database table, a flag system \u2014 produce values nobody can trace. Each layer you add makes \"why is this value what it is\" harder to answer.</p>\n<h2 id=\"secrets\">secrets<a class=\"anchor\" href=\"#secrets\" aria-label=\"link to this section\">#</a></h2>\n<p>Not the same thing as configuration and should not be in the same place.</p>\n<ul><li>Never in the config file. Never in version control. Never in the image.</li><li>A secret manager, injected at runtime.</li><li>Rotated on a schedule, and the rotation must be tested.</li><li>Never logged. Redact by default at the logger, not at each call site \u2014 because somebody will forget.</li></ul>\n<h2 id=\"the-general-point\">the general point<a class=\"anchor\" href=\"#the-general-point\" aria-label=\"link to this section\">#</a></h2>\n<p>Configuration is the input to your program that is most likely to be wrong and least likely to be checked.</p>\n<p>Every rigor you apply to code \u2014 types, validation, tests, review, staged rollout, rollback \u2014 applies to configuration, and applying it costs a day of setup.</p>\n<p>The reason it does not happen is that config does not feel like code. It is data, and data feels safe.</p>\n<p>It is not data. It is the arguments to your program, and passing wrong arguments to a program is how programs go wrong.</p>",
   "image": "https://readme.news/cards/configuration-is-the-most-under-designed-part-of-your-system.png",
   "date_published": "2026-06-24T09:00:00Z",
   "date_modified": "2026-06-24T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "engineering",
    "reliability",
    "architecture",
    "craft"
   ]
  },
  {
   "id": "https://readme.news/randomness-and-the-bugs-you-cannot-reproduce/",
   "url": "https://readme.news/randomness-and-the-bugs-you-cannot-reproduce/",
   "title": "Randomness, and the bugs you cannot reproduce",
   "summary": "The bug that happens once a week and never in staging. A systematic approach to the class of problem everyone handles badly.",
   "content_html": "<p>The worst bugs are the ones that happen sometimes. They resist the standard debugging loop entirely, because the loop requires reproduction and reproduction is exactly what you do not have.</p>\n<p>Most engineers approach these by staring at code and hoping. There is a better method.</p>\n<h2 id=\"the-sources-of-nondeterminism\">the sources of nondeterminism<a class=\"anchor\" href=\"#the-sources-of-nondeterminism\" aria-label=\"link to this section\">#</a></h2>\n<p>There is a finite list. Work through it.</p>\n<p><strong>Concurrency.</strong> Two things running at once with insufficient ordering. The largest category by far. Includes: unsynchronized shared state, check-then-act races, lost updates, and the classic where two requests both check \"does this exist\" and both create it.</p>\n<p><strong>Time.</strong> Anything that depends on wall clock: timeouts, expiry, scheduled work, date boundaries. These fail at midnight, at month end, on the DST transition, on leap day, and when NTP adjusts the clock backward.</p>\n<p><strong>Ordering.</strong> Hash map iteration order, filesystem directory order, unordered message delivery, parallel test execution. Code that accidentally depends on an order that is not guaranteed works until it does not.</p>\n<p><strong>External state.</strong> A cache that is sometimes warm. A connection that is sometimes pooled. A DNS response that sometimes returns a different IP. A downstream that is sometimes slow.</p>\n<p><strong>Resource exhaustion.</strong> Memory pressure changes GC timing, which changes interleaving. Connection pool exhaustion changes behavior under load. File descriptor limits. These make other latent bugs appear.</p>\n<p><strong>Actual randomness.</strong> UUIDs, load balancing, sampling, retries with jitter, partitioning by hash.</p>\n<p><strong>Uninitialized memory or undefined behavior</strong>, in languages that permit it.</p>\n<h2 id=\"the-method\">the method<a class=\"anchor\" href=\"#the-method\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>1. Instrument before you theorize.</strong></p>\n<p>You cannot reproduce it, so you must capture it. Add logging around the suspicious area \u2014 not \"entering function,\" but the actual values, the timing, the thread or task identity, the state.</p>\n<p>The instinct is to avoid adding logging to production. Add it. A bug you cannot reproduce is a bug you must observe in the wild, and that requires observation.</p>\n<p><strong>2. Find the correlation.</strong></p>\n<p>You have occurrences. What do they have in common?</p>\n<ul><li>Time of day. (Points at scheduled work, or peak load.)</li><li>Specific tenant or user. (Points at data-dependent behavior.)</li><li>Specific host or region.</li><li>Specific client version.</li><li>Load level.</li><li>Whether some other event happened first.</li></ul>\n<p>This is why wide structured events matter. If every request logs its context, the correlation is a query. Without it, you are guessing.</p>\n<p><strong>3. Make it more likely.</strong></p>\n<p>Once you have a hypothesis, try to increase the failure rate:</p>\n<ul><li><strong>Suspect a race?</strong> Add a sleep in the window you think is unprotected. If the failure rate goes from 0.1% to 90%, you found it.</li><li><strong>Suspect ordering?</strong> Randomize the order deliberately. Run tests with randomized seeds and shuffled execution.</li><li><strong>Suspect load-related?</strong> Load test with the specific pattern.</li><li><strong>Suspect time?</strong> Set the clock. Run at 23:59:59. Run on 29 February.</li></ul>\n<p>Making a rare bug common is the single most effective technique in this whole category, and it is underused because it feels like making things worse.</p>\n<p><strong>4. Add an assertion.</strong></p>\n<p>If you believe an invariant holds, assert it. In production, with an alert.</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">assert order.total == sum(i.price * i.qty for i in order.items), \\\n    f\"order {order.id} total mismatch: {order.total}\"</code></pre></div>\n<p>You will find out that it does not hold, and you will find out with the context attached rather than downstream where the symptom appears.</p>\n<p>The window between \"the invariant broke\" and \"someone noticed the symptom\" is where the information lives, and assertions collapse it to zero.</p>\n<p><strong>5. Bisect it.</strong></p>\n<p>If it started recently, <code>git bisect</code> still works on statistical failures \u2014 you just need a test script that runs the operation many times and reports failure if the rate exceeds a threshold. Slower than a deterministic bisect and far better than reading diffs.</p>\n<h2 id=\"the-prevention\">the prevention<a class=\"anchor\" href=\"#the-prevention\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Make things deterministic where you can.</strong> Inject the clock rather than calling it. Inject the random source with a seed. Sort collections before iterating where order matters. Deterministic systems have bugs you can reproduce, which means bugs you can fix.</p>\n<p><strong>Run tests in randomized order, with a printed seed.</strong> If a test only passes in a specific order, it has a hidden dependency, and that dependency is a bug in your production code more often than people assume.</p>\n<p><strong>Use your language's race detector.</strong> Go's <code>-race</code>, thread sanitizer for C and C++, Java's concurrency tooling. These find real races that have never yet manifested. Run them in CI.</p>\n<p><strong>Prefer immutability.</strong> A value that cannot change cannot be changed concurrently. This eliminates the largest category of nondeterminism structurally.</p>\n<p><strong>Log the identifiers.</strong> Trace ID, request ID, tenant. Nine tenths of the difficulty in these investigations is being unable to correlate events you already recorded.</p>\n<h2 id=\"the-hard-truth\">the hard truth<a class=\"anchor\" href=\"#the-hard-truth\" aria-label=\"link to this section\">#</a></h2>\n<p>Some of these take weeks. A concurrency bug that occurs once per million requests in a system with a subtle ordering dependency is genuinely difficult, and no method makes it easy.</p>\n<p>The method makes it <em>tractable</em>: instead of staring and hoping, you have a list of candidate sources, a way to gather evidence, and a way to raise the failure rate until it is reproducible.</p>\n<p>That is the difference between a bug you eventually fix and one that sits in the backlog for two years labeled \"cannot reproduce.\"</p>",
   "image": "https://readme.news/cards/randomness-and-the-bugs-you-cannot-reproduce.png",
   "date_published": "2026-06-22T09:00:00Z",
   "date_modified": "2026-06-22T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "craft",
    "engineering",
    "reliability"
   ]
  },
  {
   "id": "https://readme.news/what-a-staff-engineer-does-all-day/",
   "url": "https://readme.news/what-a-staff-engineer-does-all-day/",
   "title": "What a staff engineer does all day",
   "summary": "The role is genuinely ambiguous and that ambiguity is load-bearing. An attempt at a concrete description.",
   "content_html": "<p>\"Staff engineer\" is the most poorly-defined common title in the industry. It means different things at different companies, and even within one company two staff engineers may do almost nothing in common.</p>\n<p>Here is an attempt at what the role actually is, based on watching people do it well.</p>\n<h2 id=\"what-it-is-not\">what it is not<a class=\"anchor\" href=\"#what-it-is-not\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Not \"a senior engineer who has been there longer.\"</strong> Time in seat produces a senior engineer with more context, which is valuable and is not this.</p>\n<p><strong>Not \"the best coder.\"</strong> The highest-output individual contributor is a valuable role and it is a different one. Staff engineers frequently write less code than seniors.</p>\n<p><strong>Not \"a manager who did not want to manage.\"</strong> The scope is comparable to a manager's; the mechanism is entirely different.</p>\n<h2 id=\"what-it-is\">what it is<a class=\"anchor\" href=\"#what-it-is\" aria-label=\"link to this section\">#</a></h2>\n<p>The concise version: <strong>a staff engineer is responsible for the technical success of work that spans more than one team, without having authority over those teams.</strong></p>\n<p>Everything distinctive about the role follows from \"without authority.\" You cannot assign work. You cannot approve headcount. You cannot make a decision stick by deciding it. Every outcome has to be achieved through information, credibility, and persuasion.</p>\n<h2 id=\"the-actual-activities\">the actual activities<a class=\"anchor\" href=\"#the-actual-activities\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Finding the problem nobody owns.</strong> The most common form of staff-level impact. Every organization has problems that fall between teams: the shared library nobody maintains, the integration that fails and each side thinks is the other's, the performance issue that is caused by the interaction of three services.</p>\n<p>Nobody owns these, so nobody fixes them, so they persist for years. Identifying one, proving it matters, and getting it fixed is a substantial contribution and it usually requires a person who is not on any of the teams involved.</p>\n<p><strong>Making a decision that spans teams.</strong> Which of three approaches, when each team has a preference and none of them can see the whole picture. This is where the \"no authority\" constraint bites hardest: the decision has to be made in a way that the people affected accept, which means the reasoning has to be visible and the objections have to be genuinely addressed.</p>\n<p><strong>Writing the document that ends the argument.</strong> A recurring disagreement that resurfaces every quarter because nobody wrote down the resolution. One good document \u2014 with the options, the trade-offs, the decision, and the reasoning \u2014 can end a multi-year debate.</p>\n<p><strong>Being the person who read the whole system.</strong> Most engineers know their team's code. Somebody needs to know how it all fits together, including the parts nobody has touched in three years. That knowledge is what makes the cross-cutting problems visible.</p>\n<p><strong>Raising the floor.</strong> Not by writing better code. By making the good pattern easy: a library that removes a class of bug, a template that encodes the right defaults, a lint rule that prevents the mistake, documentation that means nobody has to ask. Impact through leverage rather than output.</p>\n<p><strong>Mentoring, specifically on judgment.</strong> Not \"how do I use this API.\" \"Should we build this at all,\" \"how do I disagree with my manager about a technical decision,\" \"how do I tell if this design will be a problem in a year.\"</p>\n<p><strong>Saying no with a reason, at a level where it lands.</strong> Frequently the most valuable thing a staff engineer does, and the reason it requires seniority is that the no has to come with an alternative and with credibility behind it.</p>\n<h2 id=\"the-day-concretely\">the day, concretely<a class=\"anchor\" href=\"#the-day-concretely\" aria-label=\"link to this section\">#</a></h2>\n<p>A real week looks roughly like:</p>\n<ul><li>20% writing code, usually the hard or risky part of something, or a prototype that settles an argument.</li><li>25% writing documents \u2014 designs, decisions, analyses, postmortems.</li><li>25% in conversations \u2014 reviews, one-on-ones, arguing about designs, being asked \"does this seem right to you.\"</li><li>15% reading \u2014 code, incident reports, other people's designs, the thing everyone is complaining about.</li><li>15% on whatever is currently on fire.</li></ul>\n<p>The proportion of coding is the part that surprises people moving into the role, and the discomfort of not shipping visible code is the most common reason people bounce out of it.</p>\n<h2 id=\"how-to-tell-if-someone-is-good-at-it\">how to tell if someone is good at it<a class=\"anchor\" href=\"#how-to-tell-if-someone-is-good-at-it\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Do things get decided?</strong> Not \"do they have opinions.\" Do arguments they are involved in reach a resolution that holds.</p>\n<p><strong>Do other engineers get better?</strong> Look at the people around them over a year.</p>\n<p><strong>Do they work on things nobody asked them to?</strong> The highest-value staff work is usually self-directed, because if it were obvious and assigned it would already have an owner.</p>\n<p><strong>Are they trusted by people who disagree with them?</strong> This is the real test. Someone who is only trusted by people who already agree has influence, not credibility.</p>\n<h2 id=\"the-failure-modes\">the failure modes<a class=\"anchor\" href=\"#the-failure-modes\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Becoming an architecture astronaut.</strong> Designs, diagrams, opinions, no contact with running code. The credibility runs out within about a year, and it is very hard to get back.</p>\n<p><strong>Becoming a very expensive senior engineer.</strong> Doing excellent work with single-team scope. Comfortable, valuable, and not the job.</p>\n<p><strong>Spreading too thin.</strong> Involved in twelve things, effective in none. The scope is tempting and the constraint is real: two or three significant efforts at a time is the realistic maximum.</p>\n<p><strong>Losing the ability to build.</strong> The role requires enough hands-on work to stay credible and calibrated. An engineer who has not shipped in a year is guessing.</p>",
   "image": "https://readme.news/cards/what-a-staff-engineer-does-all-day.png",
   "date_published": "2026-06-19T09:00:00Z",
   "date_modified": "2026-06-19T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "engineering-management",
    "careers",
    "craft",
    "opinion"
   ]
  },
  {
   "id": "https://readme.news/licenses-a-practical-guide-for-people-who-ship/",
   "url": "https://readme.news/licenses-a-practical-guide-for-people-who-ship/",
   "title": "Licenses: a practical guide for people who ship",
   "summary": "Not legal advice. A working engineer's map of what the common licenses actually require you to do.",
   "content_html": "<p>Most engineers can name three licenses and could not tell you what any of them require. That is a problem when your product ships hundreds of dependencies and somebody in legal eventually asks.</p>\n<p>This is not legal advice. It is the working map.</p>\n<h2 id=\"the-permissive-family\">the permissive family<a class=\"anchor\" href=\"#the-permissive-family\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>MIT, BSD (2- and 3-clause), ISC, Apache 2.0.</strong></p>\n<p><strong>What you must do:</strong> include the license text and the copyright notice with your distribution. That is essentially it.</p>\n<p><strong>What you may do:</strong> everything. Use it commercially, modify it, ship it in a closed product, sublicense it.</p>\n<p><strong>Apache 2.0 additionally:</strong> grants patent rights explicitly, and terminates those rights if you sue a contributor for patent infringement over the software. It also requires you to state significant changes you made.</p>\n<p>That patent grant is why Apache 2.0 is preferred over MIT by legal departments at larger companies. MIT is silent on patents, which is not the same as safe.</p>\n<p><strong>In practice:</strong> include a <code>NOTICES</code> file or an attribution page listing your dependencies and their license texts. Generate it from your dependency tree. This is the entire compliance obligation for the permissive family and most products do not do it.</p>\n<h2 id=\"the-weak-copyleft-family\">the weak copyleft family<a class=\"anchor\" href=\"#the-weak-copyleft-family\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>LGPL, MPL 2.0, EPL.</strong></p>\n<p><strong>The rule:</strong> modifications to the licensed files must be released under the same license. Your own code that merely uses the library does not have to be.</p>\n<p><strong>MPL 2.0 is file-based</strong>, which is the cleanest formulation: if you modify a file that came under MPL, that file stays MPL. Everything else is yours.</p>\n<p><strong>LGPL is linking-based</strong> and the details matter. Dynamic linking is generally considered fine \u2014 the user must be able to replace the library. Static linking requires either providing object files so the user can relink, or releasing under compatible terms.</p>\n<p>For most modern language ecosystems, \"dynamic linking\" does not map cleanly onto how code is actually distributed, and this is a genuine gray area. If you statically link LGPL code into a distributed binary, that is worth a real legal conversation.</p>\n<p><strong>In practice:</strong> MPL and EPL are unproblematic for most commercial use. LGPL is usually fine and deserves attention if you are shipping a compiled binary rather than running a service.</p>\n<h2 id=\"the-strong-copyleft-family\">the strong copyleft family<a class=\"anchor\" href=\"#the-strong-copyleft-family\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>GPL v2, GPL v3.</strong></p>\n<p><strong>The rule:</strong> if you distribute a work based on GPL code, the whole work must be under the GPL, and you must provide source.</p>\n<p><strong>The key question is what \"distribute\" means.</strong> Running GPL software on your server and letting users access it over a network is not distribution. This is why a very large amount of GPL software runs inside commercial SaaS products without obligation \u2014 Linux being the obvious example.</p>\n<p><strong>GPL v3 additionally:</strong> anti-tivoization (you must let users install modified versions on hardware you ship), and explicit patent provisions.</p>\n<p><strong>In practice:</strong> GPL dependencies in a hosted service are usually fine. GPL dependencies in software you ship to customers make your software GPL, which is usually not what you intended.</p>\n<p>Know which of your dependencies are GPL. Most people do not.</p>\n<h2 id=\"the-network-copyleft-family\">the network copyleft family<a class=\"anchor\" href=\"#the-network-copyleft-family\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>AGPL v3.</strong></p>\n<p><strong>The rule:</strong> GPL, plus \u2014 if users interact with the software over a network, they must be able to get the source, including your modifications.</p>\n<p>This closes the \"SaaS loophole\" deliberately. It is why many companies have a blanket policy prohibiting AGPL dependencies: the obligation is triggered by normal SaaS operation, and determining exactly how much of your system counts as \"the software\" is a question nobody wants to litigate.</p>\n<p><strong>In practice:</strong> if your company has a license policy, AGPL is almost certainly on the prohibited list. Check before you add the dependency, not after.</p>\n<h2 id=\"the-source-available-family\">the source-available family<a class=\"anchor\" href=\"#the-source-available-family\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>BUSL, SSPL, Elastic License, various \"fair source\" licenses.</strong></p>\n<p><strong>These are not open source licenses.</strong> They restrict commercial use, usually prohibiting offering the software as a service that competes with the licensor.</p>\n<p>They exist because hyperscalers built managed services on open source projects without contributing back, and the projects had no recourse. That grievance is real.</p>\n<p><strong>In practice:</strong> read the actual text, every time. They differ substantially. Many convert to an open source license after a delay \u2014 BUSL typically becomes Apache 2.0 after four years \u2014 which means the version you are using today may become permissive before you care.</p>\n<p>The specific question to answer: does your use case compete with the licensor's commercial offering? Usually the answer is obviously no and you are fine.</p>\n<h2 id=\"the-practical-program\">the practical program<a class=\"anchor\" href=\"#the-practical-program\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>1. Generate an SBOM in CI.</strong> Every build, listing every dependency and its license. Tooling for this is mature in every major ecosystem.</p>\n<p><strong>2. Fail the build on prohibited licenses.</strong> Have a policy \u2014 usually: permissive fine, weak copyleft fine, GPL depends on whether you distribute, AGPL and source-available require review.</p>\n<p><strong>3. Generate the attribution file automatically.</strong> This is the actual compliance obligation for the licenses you are most likely to use, and generating it is a build step, not a project.</p>\n<p><strong>4. Check when you add, not when you ship.</strong> The cost of removing a dependency after it is embedded is enormous. The cost of checking at <code>npm add</code> time is zero.</p>\n<h2 id=\"the-thing-that-trips-people-up\">the thing that trips people up<a class=\"anchor\" href=\"#the-thing-that-trips-people-up\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>License compatibility is not transitive intuition.</strong> You can use GPL code in a GPL project. You cannot use GPL code in an Apache-licensed library that others will embed in proprietary software \u2014 you have made your library effectively GPL, and your users will find out later.</p>\n<p>If you publish a library, its license must be compatible with every dependency it pulls in. Check that specifically, because it is the mistake that is expensive to discover after adoption.</p>",
   "image": "https://readme.news/cards/licenses-a-practical-guide-for-people-who-ship.png",
   "date_published": "2026-06-17T09:00:00Z",
   "date_modified": "2026-06-17T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "open-source",
    "policy",
    "engineering"
   ]
  },
  {
   "id": "https://readme.news/distributed-tracing-that-people-actually-use/",
   "url": "https://readme.news/distributed-tracing-that-people-actually-use/",
   "title": "Distributed tracing that people actually use",
   "summary": "Most tracing deployments produce beautiful waterfalls nobody opens. The difference is three implementation details.",
   "content_html": "<p>Distributed tracing is the right answer to \"what happened to this request across nine services.\" Most implementations produce a system that is technically correct and that nobody opens during an incident.</p>\n<p>Three details separate the two outcomes.</p>\n<h2 id=\"detail-one-the-trace-id-must-be-everywhere\">detail one: the trace ID must be everywhere<a class=\"anchor\" href=\"#detail-one-the-trace-id-must-be-everywhere\" aria-label=\"link to this section\">#</a></h2>\n<p>A trace is only useful if you can find it. Which means the trace ID must appear:</p>\n<ul><li><strong>In every log line</strong>, so you can pivot from a log to the trace.</li><li><strong>In the response headers</strong>, so a client can report it.</li><li><strong>On the error page</strong>, so a user can paste it into a support ticket.</li><li><strong>In your error tracking</strong>, so an exception links to its trace.</li><li><strong>In the support tool</strong>, so a support engineer can hand it to an engineer.</li></ul>\n<p>The workflow that makes tracing valuable: customer reports a problem \u2192 support gets the trace ID \u2192 engineer opens exactly that request and sees everything.</p>\n<p>Without the ID in those places, the workflow is: customer reports a problem \u2192 engineer tries to guess which of eight million traces it was.</p>\n<p>That single difference determines whether the investment pays off, and it is a plumbing problem rather than a tracing problem.</p>\n<h2 id=\"detail-two-spans-need-attributes-not-just-timing\">detail two: spans need attributes, not just timing<a class=\"anchor\" href=\"#detail-two-spans-need-attributes-not-just-timing\" aria-label=\"link to this section\">#</a></h2>\n<p>A span that says \"database query, 45 ms\" is nearly useless. Which query? Which table? How many rows? Was it cached?</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">with tracer.start_as_current_span(\"db.query\") as span:\n    span.set_attribute(\"db.system\", \"postgresql\")\n    span.set_attribute(\"db.operation\", \"select\")\n    span.set_attribute(\"db.table\", \"orders\")\n    span.set_attribute(\"db.rows_returned\", len(rows))\n    span.set_attribute(\"app.tenant_id\", tenant)\n    span.set_attribute(\"app.cache_hit\", cached)</code></pre></div>\n<p>The attributes are what let you ask questions the waterfall cannot answer: is this slow for one tenant, is it slow when the cache misses, is it slow when the result set is large.</p>\n<p><strong>Put your business identifiers on the root span.</strong> Tenant, user, plan tier, feature flags, client version. Then you can filter traces by them, which is how you find the pattern rather than the instance.</p>\n<p>High cardinality is correct here. This is not a metrics system.</p>\n<h2 id=\"detail-three-sample-intelligently-keep-the-interesting-ones\">detail three: sample intelligently, keep the interesting ones<a class=\"anchor\" href=\"#detail-three-sample-intelligently-keep-the-interesting-ones\" aria-label=\"link to this section\">#</a></h2>\n<p>Full sampling is expensive and mostly wasteful \u2014 the successful, fast, boring requests are identical to each other.</p>\n<p>Tail-based sampling makes the decision after the trace completes, when you know whether it was interesting:</p>\n<ul><li><strong>100% of traces with an error.</strong></li><li><strong>100% of traces above a latency threshold.</strong></li><li><strong>A small percentage of everything else</strong>, for baseline comparison.</li><li><strong>100% of traces for a specific customer</strong>, when you are debugging that customer.</li></ul>\n<p>That last one is worth building explicitly. A flag that says \"capture everything for this tenant for the next hour\" turns an unreproducible customer report into a solvable problem, and it is the single most useful debugging feature you can add to a multi-tenant system.</p>\n<h2 id=\"the-things-that-waste-effort\">the things that waste effort<a class=\"anchor\" href=\"#the-things-that-waste-effort\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Instrumenting everything.</strong> Auto-instrumentation gives you spans for every framework operation, most of which are noise. A trace with four hundred spans is not more informative than one with twenty; it is less, because the signal is buried.</p>\n<p>Instrument the boundaries \u2014 service calls, database queries, external APIs, queue operations \u2014 and add manual spans only for genuinely expensive internal operations.</p>\n<p><strong>Perfect propagation.</strong> You will have services that drop the context: a message queue without header support, a third-party integration, a legacy component. Do not block the rollout on 100% coverage. A trace with a gap is still much better than no trace.</p>\n<p><strong>Building dashboards from traces.</strong> Traces answer \"what happened to this request.\" Metrics answer \"what is happening to all requests.\" Using traces for aggregate views is expensive and slow. Use both, for what each is good at.</p>\n<h2 id=\"the-question-tracing-answers-that-nothing-else-does\">the question tracing answers that nothing else does<a class=\"anchor\" href=\"#the-question-tracing-answers-that-nothing-else-does\" aria-label=\"link to this section\">#</a></h2>\n<p>Not \"is the system slow\" \u2014 metrics tell you that, cheaper.</p>\n<p>Not \"what error happened\" \u2014 logs tell you that.</p>\n<p>**\"Why was <em>this specific request</em> slow, and what was different about it?\"**</p>\n<p>That is the question. It is the question you have during an incident, when one customer is affected and the aggregate metrics look fine. And it is unanswerable without tracing.</p>\n<p>If your tracing setup does not make that question easy to answer in under a minute, the setup is the problem, not the concept.</p>\n<h2 id=\"the-minimum-viable-version\">the minimum viable version<a class=\"anchor\" href=\"#the-minimum-viable-version\" aria-label=\"link to this section\">#</a></h2>\n<p>If you are starting from nothing:</p>\n<ol><li>Adopt OpenTelemetry. It is the standard, the instrumentation libraries are broad, and it keeps you portable across backends.</li><li>Propagate context across every service boundary.</li><li>Put the trace ID in every log line and every response header.</li><li>Add business identifiers to the root span.</li><li>Tail-sample: all errors, all slow, 1% of the rest.</li></ol>\n<p>That is a week of work and it covers most of the value. Everything beyond it is refinement.</p>",
   "image": "https://readme.news/cards/distributed-tracing-that-people-actually-use.png",
   "date_published": "2026-06-15T09:00:00Z",
   "date_modified": "2026-06-15T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "reliability",
    "engineering",
    "craft"
   ]
  },
  {
   "id": "https://readme.news/refactoring-under-an-agent/",
   "url": "https://readme.news/refactoring-under-an-agent/",
   "title": "Refactoring under an agent",
   "summary": "Large mechanical refactors are the clearest win available from coding agents. Here is the process that keeps them safe.",
   "content_html": "<p>Large mechanical refactors \u2014 rename this concept across four hundred files, migrate every call site to a new API, convert a pattern used everywhere \u2014 are the single clearest win available from coding agents.</p>\n<p>They are also where an unattended agent can do the most damage quietly. Here is the process that has worked.</p>\n<h2 id=\"why-this-is-the-sweet-spot\">why this is the sweet spot<a class=\"anchor\" href=\"#why-this-is-the-sweet-spot\" aria-label=\"link to this section\">#</a></h2>\n<p>Mechanical refactors have exactly the properties agents are good at:</p>\n<ul><li><strong>Verifiable.</strong> The tests either pass or they do not.</li><li><strong>Repetitive.</strong> The same transformation, many times, which is where humans make mistakes from fatigue.</li><li><strong>Well-specified.</strong> You can state the transformation precisely.</li><li><strong>Boring.</strong> Nobody enjoys this work and nobody does it carefully after hour two.</li></ul>\n<p>And they have the property that makes human refactoring risky: <strong>the scale defeats attention.</strong> A person converting four hundred call sites will be careful for the first fifty.</p>\n<h2 id=\"the-process\">the process<a class=\"anchor\" href=\"#the-process\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>1. Do ten by hand first.</strong></p>\n<p>Before writing any instruction, do a representative sample yourself. You will discover:</p>\n<ul><li>The cases where the mechanical transformation is wrong.</li><li>The variations you did not know existed.</li><li>What the actual rule is, as opposed to what you thought it was.</li></ul>\n<p>This is the step people skip and it is the one that determines whether the whole thing works. You cannot specify a transformation you have not performed.</p>\n<p><strong>2. Write the transformation down precisely.</strong></p>\n<p>Not \"modernize the error handling.\" The exact before and after, with the exceptions named:</p>\n<div class=\"code\"><span class=\"code-lang\">markdown</span><pre><code class=\"lang-markdown\">Replace every call of the form:\n    result, err := doThing(x)\n    if err != nil { return nil, err }\nwith:\n    result, err := doThing(x)\n    if err != nil { return nil, fmt.Errorf(\"doing thing for %s: %w\", x.ID, err) }\n\nExceptions:\n- Do not change anything in internal/legacy/ (frozen).\n- Do not change error handling inside deferred functions.\n- If the error is already wrapped, leave it.</code></pre></div>\n<p><strong>3. Establish the safety net before you start.</strong></p>\n<ul><li>Clean working tree, dedicated branch.</li><li>Full test suite passing, with a recorded baseline.</li><li>A way to check the transformation was applied correctly beyond the tests \u2014 a grep, a linter rule, an AST query.</li></ul>\n<p><strong>4. Batch it.</strong></p>\n<p>Not four hundred files in one change. Twenty to fifty files per commit, grouped by module.</p>\n<p>This matters for two reasons: a reviewable diff size, and the ability to bisect. If something is wrong, you want to know which batch introduced it.</p>\n<p><strong>5. Review the first batch line by line.</strong></p>\n<p>Every line. This is where you catch the systematic error, and a systematic error caught in batch one costs twenty minutes while the same error caught in batch twenty costs a day.</p>\n<p><strong>6. Spot-check subsequent batches, review the anomalies.</strong></p>\n<p>Once the pattern is verified, review by sampling \u2014 but read <em>every</em> diff that looks different from the pattern. Agent output that deviates from the established shape is where the interesting failures are.</p>\n<p><strong>7. Verify mechanically at the end.</strong></p>\n<div class=\"code\"><span class=\"code-lang\">bash</span><pre><code class=\"lang-bash\"># nothing left in the old form\nrg 'return nil, err$' --type go | rg -v 'internal/legacy'</code></pre></div>\n<p>If the transformation is complete, the old pattern should not exist outside the exceptions. This catches the files that were silently skipped, which is a real failure mode.</p>\n<h2 id=\"where-it-goes-wrong\">where it goes wrong<a class=\"anchor\" href=\"#where-it-goes-wrong\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>The agent \"improves\" things you did not ask about.</strong> It renames a variable while fixing the error handling. Individually reasonable, collectively it makes the diff unreviewable because you can no longer scan for the pattern.</p>\n<p>Instruct explicitly: change only what was specified, nothing else.</p>\n<p><strong>Semantic drift across batches.</strong> Batch one wraps errors one way, batch fifteen does it slightly differently, because the context is different and the model made a different reasonable choice.</p>\n<p>Fix: put the exact target form in the instruction file, with examples, and check for consistency at the end with a grep.</p>\n<p><strong>Tests pass and behavior changed.</strong> The most dangerous case. Your tests did not cover the path that broke.</p>\n<p>This is why the mechanical verification in step seven matters. It checks the transformation, not the behavior, and it catches things tests do not.</p>\n<p><strong>Silent skips.</strong> The agent processes 380 of 400 files and reports success. The twenty it skipped are the ones with unusual structure, which are the ones most likely to matter.</p>\n<p>Always count. Always verify the remainder is empty.</p>\n<h2 id=\"the-honest-assessment\">the honest assessment<a class=\"anchor\" href=\"#the-honest-assessment\" aria-label=\"link to this section\">#</a></h2>\n<p>For this class of work the leverage is real and large. A refactor that would have been a week of tedium \u2014 and would therefore never have happened, which is why the codebase has the problem \u2014 becomes an afternoon.</p>\n<p>That last part is the underrated benefit. The refactors that get done are the ones that are cheap enough to do. Lowering the cost means more of them happen, and a codebase where the cross-cutting cleanup actually gets done is a meaningfully better codebase.</p>\n<p>Just count the files at the end.</p>",
   "image": "https://readme.news/cards/refactoring-under-an-agent.png",
   "date_published": "2026-06-12T09:00:00Z",
   "date_modified": "2026-06-12T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "ai",
    "coding-agents",
    "craft"
   ]
  },
  {
   "id": "https://readme.news/the-world-cup-is-the-largest-load-test-ever-run/",
   "url": "https://readme.news/the-world-cup-is-the-largest-load-test-ever-run/",
   "title": "The World Cup is the largest load test ever run",
   "summary": "A month of synchronized global demand across three countries, sixteen cities, and every streaming platform at once.",
   "content_html": "<p>The tournament starts tomorrow: forty-eight teams, three host countries, sixteen venues, and a match schedule designed so that a very large fraction of the planet is watching the same thing at the same moment, repeatedly, for a month.</p>\n<p>From an engineering perspective this is the most demanding recurring event in consumer computing, and almost nothing about how it works gets written up.</p>\n<h2 id=\"the-shape-of-the-demand\">the shape of the demand<a class=\"anchor\" href=\"#the-shape-of-the-demand\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Synchronized, not smooth.</strong> Streaming traffic for a scheduled match is a step function. Millions of concurrent sessions establish within a two-minute window around kickoff, and the ones that fail to establish are a product failure with no recovery \u2014 the user missed the start.</p>\n<p><strong>Multi-peak within a session.</strong> A goal produces a spike in social traffic, in betting platforms, in messaging, in news sites, and in the streams of people switching from another match. These arrive within seconds of each other and are correlated across completely unrelated companies.</p>\n<p><strong>Correlated across the industry.</strong> This is the part that is genuinely unusual. Every CDN, every mobile network, every payment processor, and every messaging platform experiences the peak simultaneously. There is no borrowing capacity from a quiet neighbor, because there is no quiet neighbor.</p>\n<p><strong>Multi-region with different profiles.</strong> Matches in three countries across several time zones means the traffic profile shifts across the tournament in ways that capacity planning must anticipate.</p>\n<h2 id=\"what-breaks-historically\">what breaks, historically<a class=\"anchor\" href=\"#what-breaks-historically\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Payment processing at scale.</strong> Betting and merchandise platforms see enormous transaction spikes at specific moments. Payment processors are a shared dependency and their capacity is a hard constraint nobody downstream controls.</p>\n<p><strong>Mobile networks in venue areas.</strong> Sixty thousand people in one place, all trying to upload video. This is a well-understood problem with an expensive solution (temporary cell capacity) and it still degrades.</p>\n<p><strong>Authentication systems.</strong> Everyone logs in at once. Login is frequently the least scaled part of a streaming stack because it is not on the hot path during normal operation.</p>\n<p><strong>The thundering herd on recovery.</strong> Something fails, comes back, and every client reconnects simultaneously \u2014 causing a second failure. Retry jitter is the single most important line of code in a system like this and it is routinely absent.</p>\n<p><strong>Ad insertion.</strong> Server-side ad insertion personalizes streams per viewer while keeping segments cacheable. It is a hard constraint and it is a disproportionate source of live-stream incidents.</p>\n<h2 id=\"what-the-good-operators-do\">what the good operators do<a class=\"anchor\" href=\"#what-the-good-operators-do\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Pre-scale, do not autoscale.</strong> Autoscaling responds to load after it arrives. Instance startup is measured in tens of seconds; the demand arrives in one. For a known event at a known time, capacity is provisioned in advance and the autoscaler is a safety net, not the mechanism.</p>\n<p><strong>Load shed by feature, not by user.</strong> When capacity is short, disable the recommendations, the comments, the statistics overlay \u2014 keep the video. A degraded experience for everyone beats a perfect experience for 80% and nothing for the rest.</p>\n<p><strong>Multi-CDN with active steering.</strong> Not for capacity alone \u2014 for the fact that any single CDN will have a bad region on a given day. A client-side or DNS-level steering layer that measures real performance and shifts traffic is the difference between a degraded minute and an outage.</p>\n<p><strong>Rehearse.</strong> Group stage matches are the rehearsal for the knockout rounds. The teams that treat early matches as production load tests, with instrumentation and a retrospective after each one, are the ones that survive the final.</p>\n<p><strong>A war room with authority.</strong> Not a monitoring dashboard \u2014 a room with the people who can make decisions, including the decision to turn features off, without an approval chain.</p>\n<h2 id=\"the-generalizable-lesson\">the generalizable lesson<a class=\"anchor\" href=\"#the-generalizable-lesson\" aria-label=\"link to this section\">#</a></h2>\n<p>The interesting property here is <strong>demand you cannot smooth, shed, or refuse</strong>.</p>\n<p>Most systems get to spread load over time, queue it, or degrade gracefully by making users wait. None of those work when the product is a live event \u2014 a queue means the user misses the goal.</p>\n<p>What is left is: provision for peak, make every layer redundant, degrade by feature rather than by user, and rehearse.</p>\n<p>That is expensive, unglamorous, and the only thing that works. It is worth remembering the next time someone proposes autoscaling as the answer to a spike that arrives faster than a machine can boot.</p>\n<p>Some capacity you have to already own.</p>",
   "image": "https://readme.news/cards/the-world-cup-is-the-largest-load-test-ever-run.png",
   "date_published": "2026-06-10T09:00:00Z",
   "date_modified": "2026-06-10T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "infrastructure",
    "reliability"
   ]
  },
  {
   "id": "https://readme.news/wwdc-2026-and-the-platform-that-keeps-its-own-counsel/",
   "url": "https://readme.news/wwdc-2026-and-the-platform-that-keeps-its-own-counsel/",
   "title": "WWDC 2026 and the platform that keeps its own counsel",
   "summary": "New OS versions, more on-device model surface, and a developer relationship that remains complicated.",
   "content_html": "<p>Apple's developer conference ran this week. The pattern of the last few years continues: strong on-device capability, tight platform integration, and a set of platform policy questions that are being settled in courtrooms rather than on stage.</p>\n<h2 id=\"the-on-device-strategy-holding\">the on-device strategy, holding<a class=\"anchor\" href=\"#the-on-device-strategy-holding\" aria-label=\"link to this section\">#</a></h2>\n<p>Apple's position has been consistent and, I think, correct for their situation:</p>\n<ul><li>A small model on device, free, private, offline, exposed to third-party apps through a constrained system API.</li><li>A larger model available for the cases the small one cannot handle.</li><li>Routing handled by the system.</li><li>Privacy as the differentiating property rather than raw capability.</li></ul>\n<p>That is not going to win a benchmark comparison and it was never trying to. It is going to win on the axis where Apple competes, which is a coherent product where the default behavior is the one most users want.</p>\n<p>For developers, the practical implication has not changed: <strong>design features so the small model handles the common case.</strong></p>\n<p>Concretely \u2014 the 90% that the on-device model can do is free, instant, and works on a plane. The 10% that needs escalation costs money and needs a network. Getting that split right is the engineering, and it is a different skill from prompt design.</p>\n<h2 id=\"the-swift-trajectory\">the Swift trajectory<a class=\"anchor\" href=\"#the-swift-trajectory\" aria-label=\"link to this section\">#</a></h2>\n<p>Swift continues its expansion beyond Apple platforms \u2014 server-side, embedded, WebAssembly, cross-platform tooling. The concurrency model's strict checking continues to produce a language that catches data races at compile time, which is a genuinely valuable property that the migration cost has made contentious.</p>\n<p>The honest state: strict concurrency checking is correct and it is a real migration burden for existing codebases, and the error messages when you get it wrong are still harder to act on than they should be.</p>\n<p>Swift is a better language than it gets credit for outside the Apple ecosystem and its adoption outside that ecosystem remains limited, mostly for reasons of tooling gravity rather than language quality.</p>\n<h2 id=\"the-platform-policy-question\">the platform policy question<a class=\"anchor\" href=\"#the-platform-policy-question\" aria-label=\"link to this section\">#</a></h2>\n<p>The regulatory pressure on app distribution and payments continues, differently in different jurisdictions, with the result that the rules now vary by region in ways that are genuinely confusing for developers to comply with.</p>\n<p>I do not have a clean position on the substance. What I will say is the practical part:</p>\n<p><strong>If you ship an app with any monetization, you now have jurisdiction-specific compliance work</strong>, and the rules are still moving. Budget for it, watch the changes, and do not assume a policy you read last year is current.</p>\n<p>The broader observation, which is not about Apple specifically: the era of one global set of platform rules is over. Every major platform now operates under different obligations in the EU, the US, and several other markets, and that fragmentation is going to increase rather than resolve.</p>\n<h2 id=\"the-things-worth-actually-adopting\">the things worth actually adopting<a class=\"anchor\" href=\"#the-things-worth-actually-adopting\" aria-label=\"link to this section\">#</a></h2>\n<p>Filtering out the keynote, the items that will matter to a working developer:</p>\n<p><strong>Anything that reduces app size or launch time.</strong> These are the metrics that correlate with retention and they get the least stage time.</p>\n<p><strong>Testing and debugging improvements in the toolchain.</strong> Consistently the most valuable and least covered part of every WWDC.</p>\n<p><strong>Deprecation notices.</strong> The most important announcements at any platform conference are the ones about what is going away, and they are never in the keynote. Read the release notes.</p>\n<h2 id=\"the-honest-assessment\">the honest assessment<a class=\"anchor\" href=\"#the-honest-assessment\" aria-label=\"link to this section\">#</a></h2>\n<p>Apple ships coherent, well-integrated platforms with genuinely good on-device capability and a privacy posture that is a real product differentiator rather than only marketing.</p>\n<p>It also operates the developer relationship with less flexibility than any comparable platform, and the regulatory environment is the mechanism by which that is being adjusted rather than any change of heart.</p>\n<p>Both of those have been true for a decade and neither is changing this year.</p>",
   "image": "https://readme.news/cards/wwdc-2026-and-the-platform-that-keeps-its-own-counsel.png",
   "date_published": "2026-06-08T09:00:00Z",
   "date_modified": "2026-06-08T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "platforms",
    "languages",
    "ai"
   ]
  },
  {
   "id": "https://readme.news/static-analysis-is-finally-worth-the-false-positives/",
   "url": "https://readme.news/static-analysis-is-finally-worth-the-false-positives/",
   "title": "Static analysis is finally worth the false positives",
   "summary": "The tools got dramatically better while everyone was ignoring them because of a bad experience in 2015.",
   "content_html": "<p>A lot of engineers formed their opinion of static analysis from a tool that produced four thousand warnings on first run, 95% of which were noise, and got disabled within a month.</p>\n<p>That was an accurate assessment of the tools at the time. The tools are substantially different now and the assessment has not updated.</p>\n<h2 id=\"what-changed\">what changed<a class=\"anchor\" href=\"#what-changed\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Flow-sensitive analysis became standard.</strong> Older linters matched patterns in the syntax tree. Modern analyzers track values through the control flow graph, which means they can tell that a variable was checked for null on line 12 and therefore is not null on line 40. That single capability eliminates the largest source of false positives.</p>\n<p><strong>Language servers made it interactive.</strong> A warning in your editor while you type is a different product from a report generated in CI. You fix it in context, in seconds, instead of triaging a list a week later.</p>\n<p><strong>Type systems absorbed much of it.</strong> A lot of what static analysis used to catch is now caught by the compiler in a typed language, for free, with no false positives at all.</p>\n<p><strong>The defaults got sane.</strong> Modern tools ship with a curated recommended set rather than everything enabled. <code>clippy</code>, <code>ruff</code>, <code>biome</code>, <code>staticcheck</code> and their peers are opinionated about what is worth reporting.</p>\n<p><strong>They got fast.</strong> Analyzers written in compiled languages run over a large codebase in seconds. Speed matters more than people credit \u2014 a check that takes two minutes gets run in CI, and a check that takes two seconds gets run on every save, which is where it actually changes behavior.</p>\n<h2 id=\"what-to-actually-run\">what to actually run<a class=\"anchor\" href=\"#what-to-actually-run\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>A fast linter with a good default set</strong>, on save, in the editor. <code>ruff</code> for Python, <code>clippy</code> for Rust, <code>biome</code> or <code>eslint</code> for JavaScript, <code>staticcheck</code> for Go.</p>\n<p><strong>A typechecker in strict mode</strong>, in CI. This is the highest-value item on the list for a gradually-typed language, and the non-strict modes permit exactly the holes that make the guarantees unreliable.</p>\n<p><strong>A security-focused analyzer</strong> if you handle untrusted input. Taint tracking \u2014 does data from a request reach a SQL query, a shell command, or a template without sanitization \u2014 is the specific capability worth having, and it is a genuinely different analysis from ordinary linting.</p>\n<p><strong>A dependency scanner</strong>, ranked by reachability if your tooling supports it. A critical CVE in code you never call is lower priority than a medium in your request path, and a scanner that cannot tell you which is which produces a queue nobody reads.</p>\n<h2 id=\"the-adoption-sequence\">the adoption sequence<a class=\"anchor\" href=\"#the-adoption-sequence\" aria-label=\"link to this section\">#</a></h2>\n<p>Turning on a full rule set against an existing codebase produces thousands of warnings and gets the tool disabled. The sequence that works:</p>\n<p><strong>1. Run it in report-only mode.</strong> Get the number. Do not fix anything yet.</p>\n<p><strong>2. Enable a small subset that has near-zero false positives</strong> and fix those. Usually: unused variables, unreachable code, obviously wrong comparisons, missing awaits. Twenty rules, not four hundred.</p>\n<p><strong>3. Make it blocking for new and changed code only.</strong> Most tools support this, either natively or through a diff-aware wrapper. This is the key move \u2014 the existing violations do not block anyone, and the codebase stops getting worse immediately.</p>\n<p><strong>4. Burn down the backlog opportunistically.</strong> When you touch a file, fix its warnings. No cleanup sprint, no dedicated project.</p>\n<p><strong>5. Add rules gradually</strong>, one at a time, each with a burn-down.</p>\n<p>Steps three and four are where most adoptions succeed or fail. A tool that blocks the whole team on a pre-existing backlog gets turned off; one that only blocks new violations is uncontroversial.</p>\n<h2 id=\"the-rules-worth-arguing-about\">the rules worth arguing about<a class=\"anchor\" href=\"#the-rules-worth-arguing-about\" aria-label=\"link to this section\">#</a></h2>\n<p>Some checks are genuinely contested and you should decide deliberately rather than accepting the default:</p>\n<p><strong>Cyclomatic complexity limits.</strong> Sometimes a function is legitimately complex because the domain is. A hard limit produces artificially split functions that are harder to read, not easier.</p>\n<p><strong>Line length.</strong> Real disagreement, formatter should handle it, not worth a rule.</p>\n<p><strong>Naming conventions.</strong> Worth enforcing, and pick your convention rather than the tool's default if they differ.</p>\n<p><strong>Anything with more than a few percent false positives.</strong> A rule that is wrong one time in ten trains people to ignore it, and that habit generalizes to the rules that are right.</p>\n<h2 id=\"the-honest-limits\">the honest limits<a class=\"anchor\" href=\"#the-honest-limits\" aria-label=\"link to this section\">#</a></h2>\n<p>Static analysis finds a specific class of bug: local, syntactic, pattern-matchable. It does not find logic errors, wrong business rules, race conditions in most cases, or performance problems.</p>\n<p>It is not a substitute for tests, review, or thought. It is a way to spend zero human attention on the errors that do not require human attention, which frees attention for the ones that do.</p>\n<p>That framing \u2014 attention allocation rather than bug finding \u2014 is the one that makes it worth the setup.</p>\n<h2 id=\"the-new-reason-it-matters\">the new reason it matters<a class=\"anchor\" href=\"#the-new-reason-it-matters\" aria-label=\"link to this section\">#</a></h2>\n<p>Machine-generated code has a characteristic error profile: plausible, syntactically valid, and wrong in specific recurring ways. Unchecked errors, missing awaits, resource leaks, off-by-one in boundary conditions.</p>\n<p>Those are exactly the errors static analysis is good at. Running a strict analyzer over generated code is verification you get for free, and it is one of the few places where the verification bottleneck has an automated answer.</p>\n<p>Turn it on.</p>",
   "image": "https://readme.news/cards/static-analysis-is-finally-worth-the-false-positives.png",
   "date_published": "2026-06-05T09:00:00Z",
   "date_modified": "2026-06-05T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "tooling",
    "testing",
    "engineering",
    "craft"
   ]
  },
  {
   "id": "https://readme.news/time-zones-and-why-your-calendar-code-is-wrong/",
   "url": "https://readme.news/time-zones-and-why-your-calendar-code-is-wrong/",
   "title": "Time zones, and why your calendar code is wrong",
   "summary": "An instant is not a time. A time is not a date. The rules change by government decree with a few weeks' notice.",
   "content_html": "<p>Date and time handling is the domain where confident code is most reliably wrong, because the domain is defined by politics rather than physics and it changes without warning.</p>\n<h2 id=\"the-three-distinct-things\">the three distinct things<a class=\"anchor\" href=\"#the-three-distinct-things\" aria-label=\"link to this section\">#</a></h2>\n<p>Most bugs come from conflating these.</p>\n<p><strong>1. An instant.</strong> A specific moment in the history of the universe. Stored as a UTC timestamp or an epoch value. \"The order was placed at 2026-06-03T14:22:11Z.\"</p>\n<p>This is what you want for: logs, audit trails, created-at, anything that records something that happened.</p>\n<p><strong>2. A local date and time with a zone.</strong> \"The meeting is at 09:00 on 2026-09-15 in Europe/Berlin.\"</p>\n<p>Critically, this is <strong>not</strong> convertible to an instant in advance without risk, because the offset for Europe/Berlin on that date depends on the rules in effect at that time, and the rules can change between now and then.</p>\n<p>This is what you want for: future scheduled events, business hours, recurring appointments.</p>\n<p><strong>3. A plain date.</strong> \"2026-06-03.\" No time, no zone. A birthday. A contract date. An invoice period.</p>\n<p>Storing a birthday as an instant is a bug. It will shift across the date boundary for users in some zones and someone will be wished a happy birthday on the wrong day.</p>\n<h2 id=\"the-rules\">the rules<a class=\"anchor\" href=\"#the-rules\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Store instants in UTC.</strong> Always. <code>timestamptz</code> in Postgres, never <code>timestamp</code>. Convert at the display boundary.</p>\n<p><strong>Store future events as local time plus a zone identifier</strong>, not as a UTC instant.</p>\n<p>This is the one people get wrong most often, and it is the most consequential.</p>\n<p>If a user schedules a meeting for 09:00 on 15 September in Berlin, and you convert that to UTC now and store the instant \u2014 and Germany changes its daylight saving rules in July, which is a thing legislatures do \u2014 your meeting is now at the wrong time.</p>\n<p>Store <code>(\"2026-09-15T09:00\", \"Europe/Berlin\")</code>. Convert at display time, using current rules.</p>\n<p><strong>Use IANA zone identifiers</strong>, never offsets. <code>America/New_York</code>, not <code>UTC-5</code> and not <code>EST</code>.</p>\n<p>An offset is a fact about one instant. A zone is a set of rules over time. Abbreviations are worse: <code>CST</code> is Central Standard Time, China Standard Time, and Cuba Standard Time, and there is no way to tell which.</p>\n<p><strong>Never do arithmetic on local times.</strong> Adding 24 hours to a local time is not the same as adding one day. On a DST transition day, one is 23 hours and the other is 25.</p>\n<p>Decide which you mean:</p>\n<ul><li>\"24 hours later\" \u2014 arithmetic on the instant.</li><li>\"the same time tomorrow\" \u2014 arithmetic on the local calendar date, then convert.</li></ul>\n<p>These differ twice a year and the bug reports are seasonal.</p>\n<p><strong>Keep your timezone database updated.</strong> The IANA database is updated several times a year because governments change rules, sometimes with weeks of notice.</p>\n<p>If your application bundles a tzdata snapshot from two years ago, it is wrong for several countries right now. This includes container images, JVM installations, and some language runtimes' bundled copies.</p>\n<h2 id=\"the-specific-traps\">the specific traps<a class=\"anchor\" href=\"#the-specific-traps\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>DST transitions create times that do not exist and times that occur twice.</strong> At a spring-forward transition, 02:30 local does not exist. At fall-back, 01:30 happens twice.</p>\n<p>Your date library has a policy for this. Know what it is. The usual options are throw, shift forward, or pick the earlier occurrence, and the right one is application-specific.</p>\n<p><strong>Leap seconds.</strong> Mostly abstracted away by clock smearing at the infrastructure layer, and worth knowing exists. A minute is not always 60 seconds.</p>\n<p><strong>Countries change zones.</strong> Not just DST rules \u2014 entire offsets. Several countries have shifted their standard time in recent decades.</p>\n<p><strong>\"Midnight\" is ambiguous.</strong> Is 2026-06-03T00:00 the start or the end of 3 June? Prefer half-open intervals: <code>[start, end)</code>. A day is from midnight inclusive to the next midnight exclusive. This eliminates an entire category of off-by-one.</p>\n<p><strong>The user's zone is not their locale is not their country.</strong> A user in Berlin may want dates in US format. Store these as separate preferences.</p>\n<p><strong>Recurring events are their own discipline.</strong> \"Every Tuesday at 09:00\" across a DST boundary means the UTC instant changes. Store the rule, not the instances, and expand at read time.</p>\n<h2 id=\"the-practical-advice\">the practical advice<a class=\"anchor\" href=\"#the-practical-advice\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Use a real date library.</strong> Not string manipulation, not manual arithmetic. The edge cases are numerous, well known, and already handled by people who spent years on them.</p>\n<p>In JavaScript, the Temporal API finally provides the right distinctions natively \u2014 <code>Instant</code>, <code>ZonedDateTime</code>, <code>PlainDate</code> map exactly onto the three things above, which is why it took so long to design.</p>\n<p><strong>Test at boundaries.</strong> Your test suite should include a DST transition, a leap day, a year boundary, and a zone with a non-hour offset \u2014 India is UTC+5:30, Nepal is UTC+5:45, Chatham Islands is UTC+12:45. Code that assumes whole-hour offsets is common and wrong.</p>\n<p><strong>Never store or transmit local time without a zone.</strong> A timestamp without a zone is an ambiguous string and someone downstream will guess, and they will guess UTC, and they will be wrong by up to fourteen hours.</p>",
   "image": "https://readme.news/cards/time-zones-and-why-your-calendar-code-is-wrong.png",
   "date_published": "2026-06-03T09:00:00Z",
   "date_modified": "2026-06-03T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "engineering",
    "craft",
    "type-systems"
   ]
  },
  {
   "id": "https://readme.news/rewrites-when-they-actually-work/",
   "url": "https://readme.news/rewrites-when-they-actually-work/",
   "title": "Rewrites: when they actually work",
   "summary": "The received wisdom is never rewrite. The received wisdom is mostly right and has three real exceptions.",
   "content_html": "<p>The canonical advice is that rewriting from scratch is the single worst strategic mistake a software company can make. That advice is twenty-five years old and it is mostly still correct.</p>\n<p>It has exceptions, and knowing which situation you are in matters more than knowing the rule.</p>\n<h2 id=\"why-rewrites-usually-fail\">why rewrites usually fail<a class=\"anchor\" href=\"#why-rewrites-usually-fail\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>The old system's behavior is undocumented and load-bearing.</strong> Every strange conditional in a decade-old codebase is a bug report somebody filed. You will not find them by reading the code, because the code does not say why. You will find them by shipping the rewrite and having those bugs re-reported.</p>\n<p><strong>The rewrite has no users, so it gets no feedback.</strong> The old system is being exercised by real traffic continuously. The new one is exercised by your test suite, which encodes what you think it should do.</p>\n<p><strong>Feature parity is a moving target.</strong> The old system keeps getting features, because the business does not stop. You are chasing a target that recedes, and the chase consumes the time you budgeted for the rewrite.</p>\n<p><strong>Nobody can justify continuing past month six.</strong> The rewrite has produced nothing users can see. The pressure to redirect the team to visible work is enormous and usually wins, leaving you with two systems.</p>\n<p><strong>The knowledge is in the people, not the code.</strong> And the people who knew have frequently left, which is often why the rewrite was proposed.</p>\n<h2 id=\"the-three-cases-where-it-is-right\">the three cases where it is right<a class=\"anchor\" href=\"#the-three-cases-where-it-is-right\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>1. The platform is dead.</strong></p>\n<p>The framework is unmaintained, the language runtime is past end-of-life and getting CVEs, the vendor discontinued the product, the hardware is unobtainable.</p>\n<p>This is not a preference \u2014 you are being forced. Rewrite, and be grateful you found out before the security incident.</p>\n<p><strong>2. The domain model is fundamentally wrong.</strong></p>\n<p>Not \"the code is messy.\" The core abstractions do not correspond to reality, and every feature requires working around them.</p>\n<p>The test: can you name a specific business capability that is <em>impossible</em>, not just awkward, in the current model? \"We cannot support customers with more than one billing entity, and the fix requires changing what a customer is.\"</p>\n<p>If you cannot name one, you have messy code, not a wrong model, and messy code is fixed by refactoring.</p>\n<p><strong>3. The system is small enough that the rewrite is short.</strong></p>\n<p>If the whole thing is three weeks of work, the calculus changes entirely. The risk of a three-week rewrite is bounded. Just do it.</p>\n<p>The received wisdom is about large systems, and people apply it to small ones where it does not hold.</p>\n<h2 id=\"how-to-do-it-when-you-must\">how to do it when you must<a class=\"anchor\" href=\"#how-to-do-it-when-you-must\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Strangler fig, never big bang.</strong></p>\n<p>Put a routing layer in front of the old system. Move one endpoint at a time to the new implementation. Route traffic gradually. When everything has moved, delete the old system.</p>\n<div class=\"code\"><pre><code>             \u250c\u2500\u2192 new service (endpoints A, B)\nclient \u2192 router\n             \u2514\u2500\u2192 legacy monolith (everything else)</code></pre></div>\n<p>Properties this gives you:</p>\n<ul><li><strong>Value ships continuously.</strong> Every migrated endpoint is a delivered improvement.</li><li><strong>Risk is bounded per endpoint.</strong> If one goes wrong, route it back.</li><li><strong>The project survives leadership changes</strong>, because it is producing visible progress the whole time.</li><li><strong>You learn the old system's real behavior</strong> incrementally, at the point where you have to reimplement it.</li></ul>\n<p><strong>Run both and compare.</strong> For a while, send traffic to both implementations, return the old one's response, and log the differences. This is the single most effective technique for discovering undocumented behavior, and it finds things no amount of code reading would have.</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">old = legacy.compute(req)\ntry:\n    new = rewritten.compute(req)\n    if new != old:\n        log.warn(\"divergence\", request=req, old=old, new=new)\nexcept Exception as e:\n    log.error(\"new path failed\", error=e, request=req)\nreturn old      # old is still authoritative</code></pre></div>\n<p>Run that for weeks. The divergence log is your actual specification.</p>\n<p><strong>Freeze the old system's features.</strong> If the business keeps adding to it, you will never catch up. This requires an explicit organizational decision and it is usually the hardest part.</p>\n<p><strong>Set a deletion date and mean it.</strong> The worst outcome is two systems forever. If the migration stalls at 70%, you have doubled your maintenance surface permanently.</p>\n<h2 id=\"the-question-to-ask-first\">the question to ask first<a class=\"anchor\" href=\"#the-question-to-ask-first\" aria-label=\"link to this section\">#</a></h2>\n<p>Before any rewrite: <strong>what specifically will be better, and how will you know?</strong></p>\n<p>If the answer is \"the code will be cleaner,\" that is not a business outcome and the project will lose its funding to something that is.</p>\n<p>If the answer is \"features in this area will take one week instead of three, and here are the last five that took three,\" you have a case.</p>\n<p>The rewrites that succeed are the ones where somebody could state the payback in a sentence. The ones that fail are the ones motivated by taste, and taste is not wrong \u2014 it is just not fundable.</p>",
   "image": "https://readme.news/cards/rewrites-when-they-actually-work.png",
   "date_published": "2026-06-01T09:00:00Z",
   "date_modified": "2026-06-01T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "architecture",
    "engineering",
    "craft",
    "opinion"
   ]
  },
  {
   "id": "https://readme.news/what-a-great-bug-report-contains/",
   "url": "https://readme.news/what-a-great-bug-report-contains/",
   "title": "What a great bug report contains",
   "summary": "Six fields. Most reports have two. The difference is measured in days of engineering time.",
   "content_html": "<p>A bug report's job is to get someone to the reproduction as fast as possible. Everything else is decoration.</p>\n<p>Most reports fail at this, and the cost is enormous \u2014 a round trip asking for information takes hours or days, during which the bug is not being fixed and the reporter is not being helped.</p>\n<h2 id=\"the-six-fields\">the six fields<a class=\"anchor\" href=\"#the-six-fields\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>1. What you did.</strong> The exact steps. Not \"I tried to log in\" \u2014 the sequence, including things that seem irrelevant. The irrelevant thing is frequently the cause.</p>\n<p><strong>2. What you expected.</strong> This seems redundant and is not. A meaningful fraction of bug reports are misunderstandings, and stating the expectation surfaces that immediately. It also catches the case where the software is working as designed and the design is wrong, which is a different and often more important bug.</p>\n<p><strong>3. What actually happened.</strong> Specifically. The exact error text, copied, not paraphrased and not described. Screenshots for visual issues. The full stack trace, not the last line.</p>\n<p><strong>4. Environment.</strong> Version of the software, operating system, browser, region, account type, relevant feature flags. Half of all \"cannot reproduce\" outcomes are environment differences.</p>\n<p><strong>5. Frequency.</strong> Every time? Once? Intermittently? This changes the debugging approach completely \u2014 a deterministic bug is a logic error, an intermittent one is usually concurrency, caching, or state.</p>\n<p><strong>6. An identifier.</strong> Request ID, trace ID, session ID, timestamp with timezone, account ID. This is what lets an engineer find the actual event in the logs, and it converts \"I have to reproduce this\" into \"I can look at exactly what happened.\"</p>\n<p>That last one is the highest-leverage field and the one most often absent, because most products do not surface an identifier to the user.</p>\n<p><strong>If you build software: put a request ID on your error pages.</strong> It costs nothing and it is the difference between a bug report you can act on and one you cannot.</p>\n<h2 id=\"the-title\">the title<a class=\"anchor\" href=\"#the-title\" aria-label=\"link to this section\">#</a></h2>\n<p>The title is read by dozens of people and determines whether the right person opens it.</p>\n<p>Bad: \"Login broken\" Better: \"Login fails with 500 for SSO users after password reset\"</p>\n<p>The pattern: <strong>what fails, for whom, under what condition.</strong></p>\n<h2 id=\"what-to-leave-out\">what to leave out<a class=\"anchor\" href=\"#what-to-leave-out\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Your theory about the cause</strong>, unless you have evidence. A confident wrong theory sends the investigation in the wrong direction, and reports with a theory get debugged less carefully because the reader anchors on it.</p>\n<p>If you do have a theory, put it at the bottom, labeled as speculation.</p>\n<p><strong>Emotional content.</strong> \"This is completely broken and unacceptable\" adds nothing and makes the reader defensive. The severity is communicated by the impact description, not by adjectives.</p>\n<p><strong>Multiple bugs in one report.</strong> One issue per report. Multi-bug reports get partially fixed and closed, and the remaining bugs are lost.</p>\n<h2 id=\"the-minimal-reproduction\">the minimal reproduction<a class=\"anchor\" href=\"#the-minimal-reproduction\" aria-label=\"link to this section\">#</a></h2>\n<p>If you can produce one, it is worth ten times the rest of the report combined.</p>\n<p>The technique: start with the failing case and remove things until it stops failing. The last thing you removed is involved.</p>\n<p>This is work, and it is work that would otherwise be done by someone with less context than you. A minimal reproduction frequently reveals the cause to the reporter before they finish making it.</p>\n<h2 id=\"the-template\">the template<a class=\"anchor\" href=\"#the-template\" aria-label=\"link to this section\">#</a></h2>\n<p>Put this in your issue template. Teams that do see a measurable improvement in report quality, because most people will fill in the fields you give them and will not invent them:</p>\n<div class=\"code\"><span class=\"code-lang\">markdown</span><pre><code class=\"lang-markdown\">**What I did**\n1.\n2.\n3.\n\n**Expected**\n\n**Actual**\n(exact error text, full stack trace, screenshot)\n\n**Environment**\n- Version:\n- OS / Browser:\n- Account / Region:\n\n**Frequency**\nAlways / Sometimes / Once\n\n**Identifiers**\nRequest ID:\nTimestamp (with timezone):</code></pre></div>\n<h2 id=\"for-the-person-receiving-it\">for the person receiving it<a class=\"anchor\" href=\"#for-the-person-receiving-it\" aria-label=\"link to this section\">#</a></h2>\n<p>The other half of this.</p>\n<p><strong>Reproduce before you theorize.</strong> The single most common debugging failure is building a mental model from the description and investigating that model instead of the actual behavior.</p>\n<p><strong>Ask for the missing field specifically.</strong> \"Can you send the request ID from the error page?\" gets an answer. \"Can you provide more details?\" does not.</p>\n<p><strong>Close the loop.</strong> Tell the reporter what it was. This costs a sentence and it is the reason people file good reports next time \u2014 a report that vanishes into a backlog teaches people not to bother.</p>\n<h2 id=\"the-meta-point\">the meta-point<a class=\"anchor\" href=\"#the-meta-point\" aria-label=\"link to this section\">#</a></h2>\n<p>A bug report is a handoff of context between two people, one of whom has seen the failure and one of whom has to fix it.</p>\n<p>Everything in a good report is in service of moving context across that gap. Everything that does not move context across the gap is noise.</p>",
   "image": "https://readme.news/cards/what-a-great-bug-report-contains.png",
   "date_published": "2026-05-30T09:00:00Z",
   "date_modified": "2026-05-30T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "craft",
    "communication",
    "engineering"
   ]
  },
  {
   "id": "https://readme.news/edge-computing-honestly/",
   "url": "https://readme.news/edge-computing-honestly/",
   "title": "Edge computing, honestly",
   "summary": "Running code close to users is a real win for a narrow set of workloads and a complication for everything else.",
   "content_html": "<p>Edge computing has been sold as a general architectural improvement: run your code in hundreds of locations near users, everything gets faster.</p>\n<p>The physics is real. The applicability is narrower than the marketing, and the reason is data.</p>\n<h2 id=\"the-physics\">the physics<a class=\"anchor\" href=\"#the-physics\" aria-label=\"link to this section\">#</a></h2>\n<p>Light in fiber travels roughly 200,000 km/s. New York to London and back is about 55 ms of pure propagation, before any processing. Add TLS handshakes, TCP setup, and real-world routing that is not a great circle, and a cross-Atlantic round trip is frequently 100 ms or more.</p>\n<p>Running code 20 ms from the user instead of 120 ms is a genuine improvement and it is not achievable any other way.</p>\n<h2 id=\"the-problem\">the problem<a class=\"anchor\" href=\"#the-problem\" aria-label=\"link to this section\">#</a></h2>\n<p>Your data is not at the edge. It is in a database, in one region, and if your edge function needs it, you have moved the compute closer to the user and left the round trip in place \u2014 plus added a hop.</p>\n<div class=\"code\"><pre><code>user \u2192 edge (5ms) \u2192 origin database (120ms) \u2192 edge \u2192 user</code></pre></div>\n<p>That is slower than the user talking to the origin directly, because you added a hop to a path that was always dominated by the database call.</p>\n<p>This is the single most common edge computing mistake and it is easy to make, because the architecture diagram looks right.</p>\n<h2 id=\"what-edge-is-genuinely-good-for\">what edge is genuinely good for<a class=\"anchor\" href=\"#what-edge-is-genuinely-good-for\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Anything that needs no origin data:</strong></p>\n<ul><li><strong>Redirects and rewrites.</strong> URL normalization, locale routing, legacy path mapping.</li><li><strong>Authentication token validation.</strong> A signed JWT can be verified with a public key at the edge, and an invalid request never reaches your origin. This is a real win \u2014 you reject bad traffic at the perimeter.</li><li><strong>A/B test assignment.</strong> Deterministic hash of a cookie into a bucket. No state required.</li><li><strong>Header manipulation.</strong> Security headers, CORS, feature policy.</li><li><strong>Bot filtering and rate limiting.</strong> Reject at the edge, before the request costs you anything.</li><li><strong>Personalization of cached content.</strong> Fetch the cached page, inject the user's name from a cookie, return. The expensive part stays cached.</li></ul>\n<p><strong>Anything where the data is genuinely replicated to the edge:</strong></p>\n<p>Several platforms now offer edge-replicated key-value and SQL storage. If your data is small, read-heavy, and tolerant of replication lag \u2014 configuration, feature flags, product catalogs, translations \u2014 this works well and the latency win is real.</p>\n<p>The constraints are real too: writes go to a primary, replication is eventual, and storage per location is limited.</p>\n<h2 id=\"what-edge-is-bad-for\">what edge is bad for<a class=\"anchor\" href=\"#what-edge-is-bad-for\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Anything write-heavy.</strong> Writes need coordination. Coordination needs a primary. The primary is in one place.</p>\n<p><strong>Anything requiring strong consistency.</strong> By definition, this needs coordination, which needs round trips, which is what you were trying to avoid.</p>\n<p><strong>Anything with a large working set.</strong> You cannot replicate a terabyte to three hundred locations.</p>\n<p><strong>Anything computationally heavy.</strong> Edge runtimes have tight CPU and memory limits. They are designed for milliseconds of work per request.</p>\n<p><strong>Anything that needs a specific runtime.</strong> Most edge platforms run a constrained JavaScript or WebAssembly environment. Your Python dependency with a C extension is not going there.</p>\n<h2 id=\"the-architecture-that-works\">the architecture that works<a class=\"anchor\" href=\"#the-architecture-that-works\" aria-label=\"link to this section\">#</a></h2>\n<p>Layered, with each layer doing what it is good at:</p>\n<div class=\"code\"><pre><code>edge      \u2192 auth check, rate limit, routing, cached content, header work\nregional  \u2192 application logic, caching, session state\norigin    \u2192 the database, the writes, the truth</code></pre></div>\n<p>Most requests are answered at the edge from cache. Some go to a regional application tier. Few reach the origin.</p>\n<p>That is a CDN with programmability, which is what edge computing actually is, and framing it that way produces much better decisions than framing it as \"serverless everywhere.\"</p>\n<h2 id=\"the-thing-to-measure-first\">the thing to measure first<a class=\"anchor\" href=\"#the-thing-to-measure-first\" aria-label=\"link to this section\">#</a></h2>\n<p>Before adopting any of this: <strong>where does your latency actually go?</strong></p>\n<p>Break down a typical request:</p>\n<ul><li>DNS</li><li>TLS handshake</li><li>Network round trip</li><li>Time to first byte at origin</li><li>Origin processing</li><li>Database time within that</li><li>Response transfer</li></ul>\n<p>If origin processing is 400 ms and network is 40 ms, moving compute to the edge addresses 40 ms of a 440 ms problem. Fix the 400 first.</p>\n<p>This is the most common reason edge adoption disappoints: it was applied to a latency problem that was not a network problem.</p>\n<h2 id=\"the-honest-summary\">the honest summary<a class=\"anchor\" href=\"#the-honest-summary\" aria-label=\"link to this section\">#</a></h2>\n<p>Edge is a very good CDN with programmability, and that is genuinely valuable \u2014 it lets you do real work at the perimeter that used to require an origin request.</p>\n<p>It is not a general application platform, and the platforms selling it as one are selling the constraint as a feature.</p>\n<p>Use it for the perimeter. Keep your data where it can be consistent.</p>",
   "image": "https://readme.news/cards/edge-computing-honestly.png",
   "date_published": "2026-05-29T09:00:00Z",
   "date_modified": "2026-05-29T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "infrastructure",
    "architecture",
    "opinion"
   ]
  },
  {
   "id": "https://readme.news/model-evaluation-for-people-who-ship/",
   "url": "https://readme.news/model-evaluation-for-people-who-ship/",
   "title": "Model evaluation for people who ship",
   "summary": "Not research benchmarks. A practical harness you can build in a day that makes every future model decision an hour instead of a week.",
   "content_html": "<p>Every few weeks a new model ships and someone asks whether you should switch. Without an evaluation harness, answering that takes a week of impressions and you will get it wrong. With one, it takes an hour.</p>\n<p>This is the highest-return day of engineering available to anyone building on models, and a surprising number of teams have not done it.</p>\n<h2 id=\"what-it-is-not\">what it is not<a class=\"anchor\" href=\"#what-it-is-not\" aria-label=\"link to this section\">#</a></h2>\n<p>Not MMLU. Not a leaderboard. Not \"vibes after twenty prompts.\"</p>\n<p>Public benchmarks tell you which models are worth testing. They do not predict performance on your task, because your task is not in them and because contamination is universal.</p>\n<h2 id=\"what-it-is\">what it is<a class=\"anchor\" href=\"#what-it-is\" aria-label=\"link to this section\">#</a></h2>\n<p>Fifty to two hundred examples from your actual production traffic, with expected outputs or a grading rubric, run automatically, producing a number.</p>\n<div class=\"code\"><pre><code>evals/\n  cases/\n    001-refund-request.json\n    002-ambiguous-address.json\n    ...\n  run.py\n  results/\n    2026-05-27-model-a.json</code></pre></div>\n<p>Each case:</p>\n<div class=\"code\"><span class=\"code-lang\">json</span><pre><code class=\"lang-json\">{\n  \"id\": \"042\",\n  \"input\": { \"ticket\": \"my order never arrived and I want my money back\" },\n  \"expect\": { \"category\": \"refund\", \"urgency\": \"high\", \"needs_human\": false },\n  \"notes\": \"the word 'never' should not trigger the fraud path\"\n}</code></pre></div>\n<p>That is it. The whole thing is a test suite where the assertions are fuzzier.</p>\n<h2 id=\"where-the-cases-come-from\">where the cases come from<a class=\"anchor\" href=\"#where-the-cases-come-from\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Your production failures.</strong> This is the best source by a wide margin. Every time the system gets something wrong, that becomes a case. Your eval set grows into a precise map of your problem's difficulty.</p>\n<p>Set this up as a workflow: a thumbs-down in the product, or a support escalation, creates a candidate case that someone reviews and adds.</p>\n<p><strong>Your edge cases.</strong> The weird inputs. The empty ones. The ones in another language. The adversarial ones. The ones with an injection attempt.</p>\n<p><strong>A stratified sample of normal traffic.</strong> So you notice when a change breaks the common case while fixing an edge case.</p>\n<p><strong>Cases with no correct answer.</strong> Where the right behavior is to refuse, escalate, or ask a clarifying question. Models are frequently bad at this and it is rarely tested.</p>\n<h2 id=\"grading\">grading<a class=\"anchor\" href=\"#grading\" aria-label=\"link to this section\">#</a></h2>\n<p>Three approaches, and you will use all three.</p>\n<p><strong>Exact or structural match.</strong> For classification, extraction, and structured output. Cheap, deterministic, unambiguous. Use it wherever you can.</p>\n<p><strong>Programmatic checks.</strong> For generated code: does it compile, do the tests pass. For SQL: does it run, does it return the right shape. This is the strongest form of grading and it is available more often than people realize \u2014 if you can verify mechanically, do.</p>\n<p><strong>Model-as-judge.</strong> For open-ended output. A second model grades against a rubric.</p>\n<p>Use it carefully:</p>\n<ul><li><strong>Write a specific rubric</strong>, not \"is this good.\" Score each dimension separately.</li><li><strong>Validate the judge against human ratings</strong> on a sample. If the judge disagrees with you, the judge is wrong and the rubric needs work.</li><li><strong>Use a different model than the one being evaluated</strong>, or at minimum be aware of self-preference bias, which is well documented and large.</li></ul>\n<h2 id=\"the-metrics\">the metrics<a class=\"anchor\" href=\"#the-metrics\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Accuracy on your set</strong>, obviously.</p>\n<p><strong>Cost per case.</strong> Tokens in and out, at current prices. Track this \u2014 a model that is 2% better and 4\u00d7 the cost is usually the wrong choice.</p>\n<p><strong>Latency, at percentiles.</strong> p50 and p95. Reasoning models have high variance and the average hides it.</p>\n<p><strong>Failure mode distribution.</strong> Not just how many wrong \u2014 <em>how</em> wrong. A model that fails by refusing is very different from one that fails by confidently fabricating, and the aggregate score treats them identically.</p>\n<h2 id=\"the-workflow\">the workflow<a class=\"anchor\" href=\"#the-workflow\" aria-label=\"link to this section\">#</a></h2>\n<div class=\"code\"><span class=\"code-lang\">bash</span><pre><code class=\"lang-bash\">python evals/run.py --model model-a --model model-b --parallel 8</code></pre></div>\n<p>Output a table. Commit the results. Diff them across runs.</p>\n<p>Run it:</p>\n<ul><li><strong>When any model ships.</strong> Within an hour of the announcement, you know.</li><li><strong>When you change a prompt.</strong> Prompt changes are code changes with no type system and no compiler; the eval is your only regression check.</li><li><strong>On a schedule.</strong> Providers update models behind stable names. Behavior drifts. You want to know from your dashboard, not from your support queue.</li></ul>\n<p>That last one catches something most teams never notice: the model you deployed against is not the model serving your traffic today.</p>\n<h2 id=\"the-one-that-matters-most\">the one that matters most<a class=\"anchor\" href=\"#the-one-that-matters-most\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Version your prompts and store the eval result with them.</strong></p>\n<p>A prompt is production configuration. It should be in version control, reviewed, and associated with a measured quality number. \"Someone edited the prompt and something got worse three weeks ago\" is a debugging session that should not be possible.</p>\n<h2 id=\"the-payoff\">the payoff<a class=\"anchor\" href=\"#the-payoff\" aria-label=\"link to this section\">#</a></h2>\n<p>Once this exists:</p>\n<ul><li>Model migrations are an afternoon.</li><li>Prompt changes are safe to make.</li><li>You can argue about model choice with data instead of anecdotes.</li><li>You detect provider-side drift.</li><li>Onboarding a new engineer to the AI parts of your system means handing them the eval set, which is the best available documentation of what the system is supposed to do.</li></ul>\n<p>It is a day of work. It is the single highest-leverage day available in this space and it has been for three years.</p>",
   "image": "https://readme.news/cards/model-evaluation-for-people-who-ship.png",
   "date_published": "2026-05-27T09:00:00Z",
   "date_modified": "2026-05-27T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "ai",
    "testing",
    "engineering",
    "craft"
   ]
  },
  {
   "id": "https://readme.news/idempotency-is-the-only-distributed-systems-concept-you-need/",
   "url": "https://readme.news/idempotency-is-the-only-distributed-systems-concept-you-need/",
   "title": "Idempotency is the only distributed systems concept you need",
   "summary": "Not the only one. But if you only internalize one, make it this one, because it makes most of the others survivable.",
   "content_html": "<p>The network will deliver your message twice. Or zero times. Or once, but you will not find out, so you will send it again.</p>\n<p>This is not an edge case. It is the normal operating condition of every distributed system, and idempotency is what makes it survivable.</p>\n<h2 id=\"the-fundamental-problem\">the fundamental problem<a class=\"anchor\" href=\"#the-fundamental-problem\" aria-label=\"link to this section\">#</a></h2>\n<p>You send a request. The connection times out.</p>\n<p>Did it succeed?</p>\n<p>You do not know. There are three possibilities:</p>\n<ol><li>The request never arrived. Retry is correct.</li><li>The request arrived and failed. Retry is correct.</li><li><strong>The request arrived, succeeded, and the response was lost.</strong> Retry charges the card twice.</li></ol>\n<p>You cannot distinguish these from the client. This is not a limitation of your tooling; it is a theorem about asynchronous networks. There is no protocol that resolves it.</p>\n<p>The only resolution is to make retrying safe.</p>\n<h2 id=\"what-idempotency-means\">what idempotency means<a class=\"anchor\" href=\"#what-idempotency-means\" aria-label=\"link to this section\">#</a></h2>\n<p>An operation is idempotent if performing it multiple times has the same effect as performing it once.</p>\n<p>Naturally idempotent:</p>\n<ul><li><code>PUT /user/123 {\"name\": \"Dom\"}</code> \u2014 set to a value.</li><li><code>DELETE /user/123</code> \u2014 the second one finds nothing to do.</li><li><code>SET balance = 100</code> \u2014 absolute assignment.</li></ul>\n<p>Not idempotent:</p>\n<ul><li><code>POST /orders</code> \u2014 creates a new one each time.</li><li><code>UPDATE accounts SET balance = balance + 100</code> \u2014 relative change.</li><li>Sending an email.</li><li>Incrementing a counter.</li></ul>\n<h2 id=\"the-pattern\">the pattern<a class=\"anchor\" href=\"#the-pattern\" aria-label=\"link to this section\">#</a></h2>\n<p>For operations that are not naturally idempotent, the client supplies a key and the server remembers it.</p>\n<div class=\"code\"><span class=\"code-lang\">http</span><pre><code class=\"lang-http\">POST /payments\nIdempotency-Key: 8f14e45f-ea5c-4b0d-9c1a-2f7e3d4a5b6c\n\n{\"amount_cents\": 4200, \"currency\": \"usd\", \"source\": \"card_x\"}</code></pre></div>\n<p>Server side:</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">def create_payment(key, request):\n    with transaction():\n        existing = lookup(key)\n        if existing:\n            if existing.request_hash != hash(request):\n                raise Conflict(\"key reused with different parameters\")\n            return existing.response      # replay, do not re-execute\n\n        result = charge_the_card(request)\n        store(key, hash(request), result)\n        return result</code></pre></div>\n<p>Four details that matter and are usually missed:</p>\n<p><strong>Store the result, not just the key.</strong> A retry should return the original response, not a \"already processed\" error. The client's retry should look like a successful first attempt.</p>\n<p><strong>Hash the request.</strong> If the same key arrives with different parameters, that is a client bug and you should say so rather than silently returning the wrong result.</p>\n<p><strong>Same transaction.</strong> The lookup, the work, and the store must be atomic. Otherwise two concurrent retries both find nothing and both execute.</p>\n<p><strong>Expire the keys.</strong> Days, not forever. Storage is not free and a key from a year ago is not going to be retried.</p>\n<h2 id=\"where-else-this-applies\">where else this applies<a class=\"anchor\" href=\"#where-else-this-applies\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Message consumers.</strong> At-least-once delivery guarantees duplicates. Same pattern: an idempotency key on the message, a record of processed keys, both in one transaction.</p>\n<p><strong>Webhooks you send.</strong> Include an event ID. Your receivers will need it, and if you do not provide one they will invent something worse.</p>\n<p><strong>Webhooks you receive.</strong> Assume duplicates. Every major provider retries and several will send the same event twice under normal operation.</p>\n<p><strong>Deployment and provisioning.</strong> \"Create this resource\" should succeed if it already exists in the right state. This is why declarative infrastructure tools work and imperative scripts do not.</p>\n<p><strong>Migrations and backfills.</strong> They get interrupted. They must be safe to re-run from the beginning.</p>\n<h2 id=\"the-design-advice\">the design advice<a class=\"anchor\" href=\"#the-design-advice\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Prefer absolute over relative.</strong> <code>SET balance = 100</code> is idempotent. <code>ADD 100</code> is not. When you have the choice, take the absolute form.</p>\n<p>Where you cannot \u2014 and you often cannot, because concurrent updates need relative operations \u2014 use a version or a compare-and-set:</p>\n<div class=\"code\"><span class=\"code-lang\">sql</span><pre><code class=\"lang-sql\">UPDATE accounts SET balance = balance + 100, version = version + 1\nWHERE id = $1 AND version = $2;</code></pre></div>\n<p>Zero rows updated means someone else got there first, and you know it.</p>\n<p><strong>Let the client generate the key.</strong> The client knows whether this is a new request or a retry. The server cannot tell. A server-generated key defeats the entire purpose.</p>\n<p><strong>Document it.</strong> If your API supports idempotency keys, say so prominently, say how long they are honored, and say what happens on key reuse with different parameters. Consumers who do not know about the mechanism will not use it, and then they will double-charge someone and it will be your incident too.</p>\n<h2 id=\"why-this-is-the-one-to-internalize\">why this is the one to internalize<a class=\"anchor\" href=\"#why-this-is-the-one-to-internalize\" aria-label=\"link to this section\">#</a></h2>\n<p>Most distributed systems failures are one of: a duplicate, a message lost, or an out-of-order arrival.</p>\n<p>Idempotency makes duplicates safe. Retries make loss survivable \u2014 and retries require idempotency to be safe. Which leaves ordering, which is a narrower problem that fewer systems actually have.</p>\n<p>So: one property, correctly implemented, defuses the majority of what goes wrong. That is a very good ratio, and it is why this is the thing to get right before anything else.</p>",
   "image": "https://readme.news/cards/idempotency-is-the-only-distributed-systems-concept-you-need.png",
   "date_published": "2026-05-26T09:00:00Z",
   "date_modified": "2026-05-26T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "architecture",
    "engineering",
    "craft"
   ]
  },
  {
   "id": "https://readme.news/platform-teams-that-dont-get-resented/",
   "url": "https://readme.news/platform-teams-that-dont-get-resented/",
   "title": "Platform teams that don't get resented",
   "summary": "Internal platforms fail for predictable reasons. The successful ones share four properties.",
   "content_html": "<p>Most internal platform teams end up resented by the engineers they serve. The pattern is consistent enough that the causes are identifiable.</p>\n<h2 id=\"the-failure-pattern\">the failure pattern<a class=\"anchor\" href=\"#the-failure-pattern\" aria-label=\"link to this section\">#</a></h2>\n<ol><li>Platform team forms to reduce duplicated infrastructure work.</li><li>They build an abstraction over the cloud provider.</li><li>The abstraction covers 80% of cases well.</li><li>The remaining 20% is impossible, and the escape hatch is either absent or punished.</li><li>Product teams work around the platform.</li><li>Platform team responds by mandating the platform.</li><li>Everyone is unhappy and the platform is now a tax.</li></ol>\n<p>Every step follows from the previous one. The root is step four.</p>\n<h2 id=\"the-four-properties-of-platforms-that-work\">the four properties of platforms that work<a class=\"anchor\" href=\"#the-four-properties-of-platforms-that-work\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>1. An escape hatch that is not punished.</strong></p>\n<p>The platform covers the common case. It cannot cover every case, and pretending otherwise is what breaks trust.</p>\n<p>There must be a supported path for \"I need something the platform does not do,\" and taking that path must not require an exception process, a meeting, or an apologetic Slack message.</p>\n<p>The best platforms make the escape hatch cheap and then compete on being better than it. The worst make it forbidden, which does not eliminate the need \u2014 it drives it underground.</p>\n<p><strong>2. Adoption is voluntary, at least at first.</strong></p>\n<p>A platform that teams choose is a platform that is good. A platform teams are required to use never gets the feedback that would make it good, because the feedback mechanism \u2014 people leaving \u2014 has been disabled.</p>\n<p>If you cannot get voluntary adoption, that is information. Mandating it does not fix the underlying problem; it hides it and converts a product problem into a political one.</p>\n<p>Mandate later, when it is genuinely better, and the mandate will be uncontroversial because everyone already uses it.</p>\n<p><strong>3. The abstraction leaks deliberately, not accidentally.</strong></p>\n<p>Every abstraction leaks. The question is whether you planned for it.</p>\n<p>A good platform lets you drop a level when you need to: use the paved path for the deployment, and reach the underlying resource directly when you need something specific. A bad one hides the underlying system entirely, so that when it fails you cannot debug it and neither can the platform team, because now there are two systems to understand.</p>\n<p><strong>Concretely:</strong> if your platform generates infrastructure configuration, let people see it. If it wraps a cloud API, let people access the underlying resource. If it runs their container, give them the logs from the actual runtime, not a filtered view.</p>\n<p><strong>4. The platform team is measured on adoption and satisfaction, not on compliance.</strong></p>\n<p>If the platform team's metric is \"percentage of services on the platform,\" they will optimize for mandating it.</p>\n<p>If the metric is \"would you use this if you had a choice,\" they will optimize for making it good.</p>\n<p>Ask that question quarterly, anonymously, and publish the answer.</p>\n<h2 id=\"the-specific-things-that-generate-resentment\">the specific things that generate resentment<a class=\"anchor\" href=\"#the-specific-things-that-generate-resentment\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Slow escape.</strong> A team needs something the platform does not support. The answer is \"file a request, we will look at it next quarter.\" Their deadline is Friday.</p>\n<p><strong>Breaking changes without migration paths.</strong> The platform is infrastructure. Break it and every team stops. Platform teams frequently hold themselves to a lower compatibility standard than they would accept from a vendor.</p>\n<p><strong>Opaque failures.</strong> The deploy failed. The error is a platform-internal message. The product engineer cannot debug it and must escalate, which means waiting.</p>\n<p><strong>Being a gate rather than a service.</strong> A platform that must approve things is a bureaucracy. A platform that makes the right thing easy is infrastructure.</p>\n<p><strong>Solving the platform team's problems.</strong> Standardization is valuable to the platform team and is not automatically valuable to product teams. If the pitch for a migration is \"this makes our lives easier,\" expect a cool reception.</p>\n<h2 id=\"the-framing-that-works\">the framing that works<a class=\"anchor\" href=\"#the-framing-that-works\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>You are building a product. Your users are engineers. They have alternatives.</strong></p>\n<p>That framing produces the right behaviors automatically: user research before building, documentation that assumes nothing, onboarding that works, support that responds, and a roadmap driven by what users need rather than by architectural preference.</p>\n<p>The platform teams I have seen work best behave exactly like a startup selling to a skeptical market, and they say so out loud.</p>\n<p>The ones that fail behave like an internal standards body, and they are usually correct about the standards and wrong about how to get them adopted.</p>\n<h2 id=\"the-measurement-that-matters\">the measurement that matters<a class=\"anchor\" href=\"#the-measurement-that-matters\" aria-label=\"link to this section\">#</a></h2>\n<p>Time from \"a new engineer joins\" to \"their code is running in production.\"</p>\n<p>That single number captures most of what a platform is for, it is measurable, and it is the thing product teams actually care about. If it is going down, the platform is working, regardless of what the adoption dashboard says.</p>",
   "image": "https://readme.news/cards/platform-teams-that-dont-get-resented.png",
   "date_published": "2026-05-22T09:00:00Z",
   "date_modified": "2026-05-22T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "engineering-management",
    "platforms",
    "tooling",
    "opinion"
   ]
  },
  {
   "id": "https://readme.news/open-source-funding-five-models-honestly-compared/",
   "url": "https://readme.news/open-source-funding-five-models-honestly-compared/",
   "title": "Open source funding: five models, honestly compared",
   "summary": "Donations, foundations, dual licensing, open core, and hosted service. Each works for a specific shape of project.",
   "content_html": "<p>The open source sustainability problem is not that the models do not exist. It is that projects pick a model that does not fit their shape, and then conclude that funding open source is impossible.</p>\n<p>Five models. Each works. Each works for a specific kind of project.</p>\n<h2 id=\"1-donations-and-sponsorship\">1. donations and sponsorship<a class=\"anchor\" href=\"#1-donations-and-sponsorship\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Works for:</strong> developer-facing tools with a large individual user base and an identifiable maintainer.</p>\n<p><strong>Does not work for:</strong> libraries deep in a dependency tree, infrastructure nobody knows they use, anything without a personality attached.</p>\n<p>The honest arithmetic: donation income correlates with visibility, not with importance. A well-marketed CLI tool with a charismatic maintainer will out-earn a critical cryptography library by a large multiple.</p>\n<p>The corporate sponsorship version works better than individual donations and requires the maintainer to do sales, which most maintainers are bad at and hate.</p>\n<h2 id=\"2-foundation-governance\">2. foundation governance<a class=\"anchor\" href=\"#2-foundation-governance\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Works for:</strong> infrastructure that multiple large companies depend on and that none of them wants a competitor to control.</p>\n<p><strong>Does not work for:</strong> small projects. The overhead \u2014 governance, legal, trademark, process \u2014 is substantial, and a foundation with one project and no funded staff is just more paperwork.</p>\n<p>The real value of a foundation is not money. It is neutrality: it makes a project safe for competitors to invest in together, which unlocks contribution that would not otherwise happen.</p>\n<h2 id=\"3-dual-licensing\">3. dual licensing<a class=\"anchor\" href=\"#3-dual-licensing\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Works for:</strong> libraries embedded in other products, where the copyleft obligation is genuinely inconvenient for commercial users.</p>\n<p>Ship under a strong copyleft license, sell a commercial license to companies that cannot comply.</p>\n<p><strong>Does not work for:</strong> anything permissively licensed already (no leverage), anything not embedded (the obligation does not bite), or anything with a permissive competitor of similar quality.</p>\n<p>Effective when it fits, and it produces a genuine tension: the license that makes the business work is the one that limits adoption.</p>\n<h2 id=\"4-open-core\">4. open core<a class=\"anchor\" href=\"#4-open-core\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Works for:</strong> products where enterprise features are genuinely separable from the core \u2014 SSO, audit logs, RBAC, compliance reporting, multi-tenancy.</p>\n<p><strong>Does not work for:</strong> libraries. There is no enterprise tier of a date-parsing library.</p>\n<p>The failure mode is well documented: the line between core and commercial moves toward commercial over time, under revenue pressure, and the community that built your adoption watches features they use get moved behind the paywall.</p>\n<p>If you do this, <strong>write down the line publicly, early, and honor it.</strong> \"Anything that a single developer needs is open; anything that exists because you have a compliance department is commercial\" is a defensible line. Moving it later costs more trust than the revenue is worth.</p>\n<h2 id=\"5-hosted-service\">5. hosted service<a class=\"anchor\" href=\"#5-hosted-service\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Works for:</strong> anything that is annoying to operate. Databases, search, queues, observability, CI.</p>\n<p>Give away the software, sell the operation of it. This is the strongest model when it fits, because the value you sell \u2014 not having to run it \u2014 is real and continuous, and it does not require withholding anything.</p>\n<p><strong>The risk:</strong> a hyperscaler offers a managed version of your software, at scale, without contributing back. This has happened repeatedly and it is why the source- available licenses exist.</p>\n<p>Those licenses solve the problem and cost you the open source designation, which costs you contributors, ecosystem inclusion, and some corporate adoption. It is a real trade with real costs on both sides, and the projects that made it mostly survived, which is the empirical answer to whether it works.</p>\n<h2 id=\"what-actually-kills-projects\">what actually kills projects<a class=\"anchor\" href=\"#what-actually-kills-projects\" aria-label=\"link to this section\">#</a></h2>\n<p>Not the absence of a model. Three other things:</p>\n<p><strong>Solo maintainer burnout.</strong> One person, unpaid, receiving an unbounded stream of issues, feature requests, and entitled demands. Funding helps and does not fix it \u2014 the fix is more maintainers, which is a governance problem.</p>\n<p><strong>Success without support.</strong> A project that becomes critical infrastructure while its maintainer count stays at one. This is the most dangerous state and it is extremely common.</p>\n<p><strong>Corporate abandonment.</strong> A company open-sources a project, staffs it with employees, then reorganizes. The external community was never built because it was never needed. Now nobody knows the code.</p>\n<h2 id=\"what-companies-should-do\">what companies should do<a class=\"anchor\" href=\"#what-companies-should-do\" aria-label=\"link to this section\">#</a></h2>\n<p>If your business depends on open source \u2014 and it does \u2014 the highest-leverage actions, in order:</p>\n<ol><li><strong>Pay maintainers of your critical dependencies.</strong> Directly. Small amounts to many projects beat large amounts to a few.</li><li><strong>Assign employee time to upstream contribution.</strong> More valuable than money and much rarer.</li><li><strong>Do not send compliance questionnaires to volunteers.</strong> They owe you nothing and the license says so.</li><li><strong>When you fix a bug in a vendored dependency, upstream it.</strong> The number of companies carrying private patches for bugs everyone has is enormous.</li></ol>\n<p>None of that requires a strategy document. It requires someone with budget deciding it matters, which is the actual bottleneck.</p>",
   "image": "https://readme.news/cards/open-source-funding-five-models-honestly-compared.png",
   "date_published": "2026-05-20T09:00:00Z",
   "date_modified": "2026-05-20T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "open-source",
    "industry",
    "policy",
    "opinion"
   ]
  },
  {
   "id": "https://readme.news/compression-is-underrated/",
   "url": "https://readme.news/compression-is-underrated/",
   "title": "Compression is underrated",
   "summary": "The cheapest performance win available, ignored because it is not glamorous. Where it pays and which algorithm to pick.",
   "content_html": "<p>Compression trades CPU for bytes. On modern hardware, where CPU is abundant and bandwidth is the constraint at nearly every layer, that trade is favorable far more often than people apply it.</p>\n<h2 id=\"the-layers-where-it-pays\">the layers where it pays<a class=\"anchor\" href=\"#the-layers-where-it-pays\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>HTTP responses.</strong> Everyone does this and most do it badly. Check yours:</p>\n<div class=\"code\"><span class=\"code-lang\">bash</span><pre><code class=\"lang-bash\">curl -sI -H 'Accept-Encoding: br, gzip' https://example.com/api/data | grep -i content-encoding</code></pre></div>\n<p>If that returns nothing, you are shipping uncompressed JSON, and JSON compresses extraordinarily well \u2014 commonly 80\u201390% for typical API responses, because it is mostly repeated key names.</p>\n<p>Brotli beats gzip by roughly 15\u201320% on text at comparable CPU cost, and is supported everywhere. Use it for static assets at maximum level (precomputed, so the CPU cost is paid once) and at a moderate level for dynamic responses.</p>\n<p><strong>Database storage.</strong> Most modern databases support per-table or per-column compression. On a table of text or JSON, it frequently halves the storage \u2014 which also halves the I/O, which means more of the working set fits in memory, which is where the real win is.</p>\n<p>The CPU cost of decompression is almost always smaller than the I/O cost you avoided.</p>\n<p><strong>Logs and telemetry.</strong> Log shipping is often a meaningful fraction of internal network traffic and vendor cost. Compressed batching typically reduces it by an order of magnitude, and the batching itself reduces request overhead.</p>\n<p><strong>Backups and object storage.</strong> Storage is cheap and it is not free, and egress definitely is not. Compression at rest is a direct cost reduction with no downside for cold data.</p>\n<p><strong>Container images.</strong> Zstandard-compressed layers pull faster than gzip, which matters for cold starts and for anything that scales by launching new instances.</p>\n<p><strong>Inter-service traffic.</strong> gRPC and Protocol Buffers are already compact. If you are sending JSON between services \u2014 and most people are \u2014 compressing it is a large win that costs one configuration line.</p>\n<h2 id=\"which-algorithm\">which algorithm<a class=\"anchor\" href=\"#which-algorithm\" aria-label=\"link to this section\">#</a></h2>\n<div class=\"table-wrap\"><table><thead><tr><th style=\"text-align:left\">algorithm</th><th style=\"text-align:left\">use for</th></tr></thead><tbody><tr><td style=\"text-align:left\"><strong>zstd</strong></td><td style=\"text-align:left\">almost everything. Wide speed/ratio range, fast decompression.</td></tr><tr><td style=\"text-align:left\"><strong>brotli</strong></td><td style=\"text-align:left\">HTTP text, especially static assets at max level.</td></tr><tr><td style=\"text-align:left\"><strong>gzip</strong></td><td style=\"text-align:left\">compatibility fallback. Never the best choice, always supported.</td></tr><tr><td style=\"text-align:left\"><strong>lz4</strong></td><td style=\"text-align:left\">when speed dominates entirely. In-memory, hot paths, real-time.</td></tr><tr><td style=\"text-align:left\"><strong>xz / lzma</strong></td><td style=\"text-align:left\">archives you compress once and rarely read. Slow, small.</td></tr></tbody></table></div>\n<p>The default answer is <strong>zstd</strong>. It has a level parameter spanning from faster-than-lz4 to nearly-as-small-as-xz, decompression is fast at every level, and it has dictionary support.</p>\n<h2 id=\"the-dictionary-trick\">the dictionary trick<a class=\"anchor\" href=\"#the-dictionary-trick\" aria-label=\"link to this section\">#</a></h2>\n<p>The most underused feature in compression, and the one with the biggest payoff for small messages.</p>\n<p>Compression works by finding repetition. A 200-byte JSON message has almost no internal repetition, so compression barely helps \u2014 sometimes it makes it bigger.</p>\n<p>But across <em>many</em> messages, there is enormous repetition: the same field names, the same enum values, the same URL prefixes.</p>\n<p>A shared dictionary trained on representative samples gives the compressor that repetition up front:</p>\n<div class=\"code\"><span class=\"code-lang\">bash</span><pre><code class=\"lang-bash\">zstd --train samples/*.json -o dict.zst</code></pre></div>\n<p>Then compress each message against the dictionary. Small-message ratios that were 1.1\u00d7 become 3\u00d7 or better. For any system moving many small similar messages \u2014 event streams, queue payloads, cache values \u2014 this is a large and nearly free win.</p>\n<h2 id=\"when-not-to-compress\">when not to compress<a class=\"anchor\" href=\"#when-not-to-compress\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Already-compressed data.</strong> Images, video, audio, archives. You will spend CPU to make it slightly larger.</p>\n<p><strong>Very small payloads without a dictionary.</strong> Under a few hundred bytes, compression overhead can exceed the savings.</p>\n<p><strong>When you are CPU-bound and not bandwidth-bound.</strong> Measure. This is rarer than people assume but it does happen.</p>\n<p><strong>Encrypted data, compressed after encryption.</strong> Pointless \u2014 ciphertext is incompressible. And compressing <em>before</em> encryption can leak information about the plaintext through the ciphertext length, which is the class of attack that includes CRIME and BREACH. If you compress and encrypt, know why it is safe in your context.</p>\n<h2 id=\"the-ten-minute-audit\">the ten-minute audit<a class=\"anchor\" href=\"#the-ten-minute-audit\" aria-label=\"link to this section\">#</a></h2>\n<ol><li>Check that HTTP responses are compressed, including API responses, not just HTML.</li><li>Check that Brotli is enabled, not just gzip.</li><li>Check your log shipping compresses and batches.</li><li>Check whether your largest database tables support compression and whether it is on.</li><li>If you move many small similar messages, train a dictionary.</li></ol>\n<p>That is an afternoon and it routinely produces a larger improvement than a month of application-level optimization, at a fraction of the risk.</p>\n<p>The reason it does not happen is that nobody gets promoted for enabling Brotli.</p>",
   "image": "https://readme.news/cards/compression-is-underrated.png",
   "date_published": "2026-05-18T09:00:00Z",
   "date_modified": "2026-05-18T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "performance",
    "engineering",
    "infrastructure",
    "craft"
   ]
  },
  {
   "id": "https://readme.news/the-queue-is-the-architecture/",
   "url": "https://readme.news/the-queue-is-the-architecture/",
   "title": "The queue is the architecture",
   "summary": "Most scaling problems are solved by making something asynchronous. Most reliability problems are caused by doing it badly.",
   "content_html": "<p>The single most effective architectural move available to most systems is: take the slow thing out of the request path and put it in a queue.</p>\n<p>It is also the move that introduces the most subtle failure modes, and the gap between \"we added a queue\" and \"we added a queue correctly\" is large.</p>\n<h2 id=\"what-it-buys\">what it buys<a class=\"anchor\" href=\"#what-it-buys\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Latency.</strong> The user gets a response when the work is accepted, not when it is done. A checkout that returns in 80 ms and sends the confirmation email asynchronously is a much better product than one that returns in 900 ms.</p>\n<p><strong>Absorbing spikes.</strong> A queue is a buffer. Traffic that would overwhelm a synchronous system accumulates and drains. This is the difference between a slow period and an outage.</p>\n<p><strong>Isolation.</strong> If the email provider is down, checkout still works. The messages accumulate and send later.</p>\n<p><strong>Retry for free.</strong> A failed message goes back on the queue. A failed synchronous call is a user-visible error.</p>\n<h2 id=\"what-it-costs\">what it costs<a class=\"anchor\" href=\"#what-it-costs\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Eventual consistency, everywhere.</strong> The user completed checkout and the confirmation has not arrived. The record exists and the search index does not have it. Every asynchronous boundary introduces a window where the system is inconsistent, and your UI has to be honest about it.</p>\n<p><strong>Debugging across the boundary.</strong> A synchronous stack trace tells you the whole story. An asynchronous failure requires correlating a producer, a broker, and a consumer, possibly hours apart.</p>\n<p><strong>Ordering.</strong> Most queues do not guarantee it, or guarantee it only within a partition. If your consumer must process events in order, that is a design constraint that reaches back into how you partition.</p>\n<p><strong>Duplicate delivery.</strong> Almost all queues are at-least-once. Your consumer <em>will</em> receive the same message twice. If that is not safe, you have a bug that appears under load, weeks after launch.</p>\n<h2 id=\"the-rules\">the rules<a class=\"anchor\" href=\"#the-rules\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>1. Consumers must be idempotent. Non-negotiable.</strong></p>\n<p>At-least-once delivery means duplicates. The consumer must produce the same result whether it processes a message once or five times.</p>\n<p>The usual implementation: a natural idempotency key, and a record of processed keys.</p>\n<div class=\"code\"><span class=\"code-lang\">python</span><pre><code class=\"lang-python\">def handle(msg):\n    key = msg.idempotency_key\n    with tx():\n        if already_processed(key):\n            return\n        do_the_work(msg)\n        mark_processed(key)</code></pre></div>\n<p>The <code>mark_processed</code> must be in the same transaction as the work, or you have moved the race rather than eliminated it.</p>\n<p><strong>2. Every queue needs a dead letter queue, and someone must watch it.</strong></p>\n<p>A message that fails repeatedly must go somewhere. A DLQ nobody monitors is a place where data goes to be silently lost, which is worse than an error, because errors are visible.</p>\n<p>Alert on DLQ depth. Not on it being non-zero \u2014 on it growing.</p>\n<p><strong>3. Retry with backoff and jitter, and cap the attempts.</strong></p>\n<p>Immediate retry on a failing downstream is a denial of service you are performing against yourself. Exponential backoff with jitter, a maximum attempt count, then the DLQ.</p>\n<p><strong>4. Monitor queue depth and age, not just throughput.</strong></p>\n<p>Throughput looks healthy right up until it does not. The metrics that tell you something is wrong:</p>\n<ul><li><strong>Depth</strong> \u2014 how many messages are waiting.</li><li><strong>Oldest message age</strong> \u2014 the most useful single metric. If it is growing, your consumers cannot keep up, and you know how far behind you are in time rather than in count.</li></ul>\n<p><strong>5. Decide what happens when the queue is full.</strong></p>\n<p>It will be. Reject the producer, drop messages, or block? Each is right in different cases and the default is usually wrong for you. An unbounded queue is not a solution; it is a memory leak with extra steps.</p>\n<p><strong>6. Keep the payload small and the reference stable.</strong></p>\n<p>Put an ID in the message, not the whole object. The consumer fetches current state. This avoids stale data in the message and keeps the broker fast.</p>\n<p>The exception: if you need the state <em>as it was</em> when the event occurred, put it in the message deliberately, and say so.</p>\n<p><strong>7. Version your message schema from day one.</strong></p>\n<p>Producers and consumers deploy independently. Old consumers will see new messages. Include a version field. Make additive changes only, or handle both shapes.</p>\n<h2 id=\"the-thing-that-surprises-people\">the thing that surprises people<a class=\"anchor\" href=\"#the-thing-that-surprises-people\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Queues do not reduce load. They defer it.</strong></p>\n<p>If your consumer processes 100 messages per second and you produce 150, you do not have a working system with a buffer. You have a system that is failing slowly, and the queue depth graph is a countdown.</p>\n<p>A queue absorbs <em>bursts</em>. It does not fix a sustained capacity deficit, and the failure mode when you use it that way is a queue that grows for six hours and then an incident where you are simultaneously behind and unable to catch up.</p>\n<p>Alert on the age, watch the trend, and size the consumers for the sustained rate.</p>",
   "image": "https://readme.news/cards/the-queue-is-the-architecture.png",
   "date_published": "2026-05-15T09:00:00Z",
   "date_modified": "2026-05-15T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "architecture",
    "infrastructure",
    "reliability",
    "engineering"
   ]
  },
  {
   "id": "https://readme.news/run-the-incident-before-the-incident/",
   "url": "https://readme.news/run-the-incident-before-the-incident/",
   "title": "Run the incident before the incident",
   "summary": "Game days, failure injection, and the specific reason your untested runbook is wrong.",
   "content_html": "<p>Every runbook that has never been executed is wrong. Not might be \u2014 is. The commands have changed, the dashboard moved, the person who wrote it left, and the system it describes has been modified fourteen times.</p>\n<p>The only way to find out is to run it, and the only good time to run it is when nothing is actually broken.</p>\n<h2 id=\"what-a-game-day-is\">what a game day is<a class=\"anchor\" href=\"#what-a-game-day-is\" aria-label=\"link to this section\">#</a></h2>\n<p>A scheduled exercise where you deliberately break something in a controlled way and have the on-call rotation respond as if it were real.</p>\n<p>Two hours. A stated scenario. The people who would actually respond. Real tools, real runbooks, real dashboards. Someone taking notes on everything that did not work.</p>\n<p>That is the whole practice. It is not chaos engineering in the automated-random sense, though that is a good adjacent practice. It is a rehearsal.</p>\n<h2 id=\"what-you-find-reliably\">what you find, reliably<a class=\"anchor\" href=\"#what-you-find-reliably\" aria-label=\"link to this section\">#</a></h2>\n<p>I have never run one of these that did not find at least four of the following:</p>\n<p><strong>The runbook references something that does not exist.</strong> A dashboard that was renamed, a script in a repository that was archived, an alias nobody has.</p>\n<p><strong>Nobody has the access.</strong> The runbook says to restart the service. The on-call engineer does not have permission and does not know who does. This is the single most common finding.</p>\n<p><strong>The dashboard does not show the thing.</strong> You built the alert. You never built the view that tells you what to do about it.</p>\n<p><strong>The escalation path is a person, not a rotation.</strong> \"Ask Sarah.\" Sarah is on vacation. Sarah left last year.</p>\n<p><strong>Nobody knows the customer impact.</strong> The service is degraded. Which customers? Which features? Nobody can answer, so nobody can decide how urgent it is.</p>\n<p><strong>Recovery has an undocumented step.</strong> The service restarts and does not work because a cache must be cleared first, which is knowledge that lives in one person's head.</p>\n<p><strong>The communication path is unclear.</strong> Who tells the customers? Who updates the status page? Who has the credentials for the status page?</p>\n<p>Every one of those is cheap to fix on a Tuesday afternoon and extremely expensive to discover at 3 a.m.</p>\n<h2 id=\"the-scenarios-worth-running\">the scenarios worth running<a class=\"anchor\" href=\"#the-scenarios-worth-running\" aria-label=\"link to this section\">#</a></h2>\n<p>Start with the boring ones. Exotic failure modes are fun and the common ones are what actually happens.</p>\n<p><strong>The database primary fails over.</strong> Does the application reconnect? How long? Does anything need a manual restart?</p>\n<p><strong>A dependency returns 500s.</strong> Not down \u2014 erroring. Does your circuit breaker work? Does your retry policy make it worse?</p>\n<p><strong>A dependency gets slow.</strong> Harder than down and much more common. Does your timeout fire? Do connections exhaust? Does the slowness propagate to callers?</p>\n<p><strong>Disk fills up.</strong> On the database host, on the log host, on the application host. This is the single most common self-inflicted outage.</p>\n<p><strong>A bad deploy.</strong> Deploy something broken to staging and time the rollback. Not the theoretical rollback \u2014 the actual one, executed by the actual on-call person.</p>\n<p><strong>Certificate expiration.</strong> Set one to expire in staging. Watch what happens. Most teams discover their monitoring does not cover this.</p>\n<p><strong>The person who knows is unavailable.</strong> Run a game day where the subject matter expert is explicitly not allowed to help. This finds the knowledge concentration problems that nothing else does.</p>\n<h2 id=\"how-to-run-one-that-works\">how to run one that works<a class=\"anchor\" href=\"#how-to-run-one-that-works\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Announce it.</strong> Do not surprise people. Surprise exercises generate resentment and teach people to distrust the process. Everyone should know it is a drill.</p>\n<p><strong>Staging first, production eventually.</strong> Staging finds most of the runbook problems. Production finds the ones that only exist because staging is not production, which are real and are the ones that matter most.</p>\n<p><strong>Have a stop condition.</strong> A named person who can call it off, and a defined way to revert whatever you broke.</p>\n<p><strong>Write down every friction point</strong>, including small ones. \"It took four minutes to find the right dashboard\" is a real finding.</p>\n<p><strong>Fix things within a week.</strong> A game day that produces a list nobody acts on is theater, and the second one will have lower attendance.</p>\n<p><strong>Do it quarterly.</strong> Systems change. A runbook validated a year ago is a runbook that has not been validated.</p>\n<h2 id=\"the-cultural-part\">the cultural part<a class=\"anchor\" href=\"#the-cultural-part\" aria-label=\"link to this section\">#</a></h2>\n<p>The purpose is to find gaps in the system, not gaps in people.</p>\n<p>If someone cannot resolve the scenario, that is a finding about documentation, tooling, or access \u2014 not about them. Say this explicitly before you start, and mean it, or people will optimize for looking competent rather than for surfacing problems.</p>\n<p>The best outcome of a game day is a long list of things that went wrong, discovered by people who were not stressed, on a schedule, with time to fix them.</p>",
   "image": "https://readme.news/cards/run-the-incident-before-the-incident.png",
   "date_published": "2026-05-13T09:00:00Z",
   "date_modified": "2026-05-13T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "reliability",
    "engineering",
    "testing",
    "craft"
   ]
  },
  {
   "id": "https://readme.news/python-315-and-free-threadings-second-act/",
   "url": "https://readme.news/python-315-and-free-threadings-second-act/",
   "title": "Python 3.15 and free-threading's second act",
   "summary": "The GIL-free build is officially supported and the ecosystem work is the actual story. A progress report.",
   "content_html": "<p>Python 3.15 is in beta. The headline features are incremental. The story worth tracking is the free-threading migration, which is now in the phase where it succeeds or stalls based on ecosystem work rather than on CPython.</p>\n<h2 id=\"where-free-threading-actually-stands\">where free-threading actually stands<a class=\"anchor\" href=\"#where-free-threading-actually-stands\" aria-label=\"link to this section\">#</a></h2>\n<p>The GIL-free build has been officially supported since 3.14. The question was never whether CPython could do it \u2014 it demonstrably can \u2014 but whether the ecosystem would follow.</p>\n<p><strong>What has gone well:</strong></p>\n<p>The major numerical and data libraries did the work. NumPy, and the array and dataframe ecosystem around it, ship free-threaded wheels. That was the critical path, because those libraries are the reason a large fraction of Python's CPU-bound workload exists.</p>\n<p>Single-threaded performance in the free-threaded build has improved substantially since 3.13. The gap against the default build has narrowed to something most applications would not notice.</p>\n<p>Build infrastructure caught up. Producing free-threaded wheels is now a configuration change rather than a project.</p>\n<p><strong>What has gone slowly:</strong></p>\n<p>The long tail. Thousands of packages with C extensions have not been audited, and most of them will not be until someone hits a problem.</p>\n<p>The subtler issue: <strong>pure-Python code that was accidentally thread-safe.</strong> The GIL made many operations effectively atomic. Code written under that assumption \u2014 a dictionary mutated from multiple threads, a counter incremented without a lock \u2014 worked by accident. Under free-threading it does not.</p>\n<p>Nobody knows how much library code depends on this, because nobody ever had to think about it. It will be discovered one race condition at a time.</p>\n<h2 id=\"the-practical-guidance\">the practical guidance<a class=\"anchor\" href=\"#the-practical-guidance\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Do not switch production workloads casually.</strong> The performance win only exists for CPU-bound multithreaded work. If your workload is I/O-bound, <code>asyncio</code> was already handling it and free-threading gives you nothing.</p>\n<p><strong>Do test your libraries against it.</strong> The compatibility work has to happen and it happens when people run into problems and report them. If you maintain a package with a C extension, this is your work to do.</p>\n<p><strong>Do use it for the workloads it is for.</strong> Data processing, image and signal processing, simulation, anything numeric that currently uses <code>multiprocessing</code> and pays serialization and memory-duplication costs.</p>\n<p>The migration from <code>multiprocessing</code> to threads for those workloads is frequently dramatic \u2014 not just faster, but simpler, because you delete the serialization layer.</p>\n<h2 id=\"the-other-thing-in-315\">the other thing in 3.15<a class=\"anchor\" href=\"#the-other-thing-in-315\" aria-label=\"link to this section\">#</a></h2>\n<p><strong>Subinterpreters continue maturing.</strong> <code>concurrent.interpreters</code> gives you isolated interpreter instances in one process, with much lower overhead than processes and much stronger isolation than threads.</p>\n<p>This is the underrated option. For a workload where you want parallelism and do not want shared mutable state \u2014 which is most workloads \u2014 subinterpreters give you the process model's safety at closer to the thread model's cost.</p>\n<p>It is the right default for a lot of the cases people currently reach for <code>multiprocessing</code> for, and it does not require any of the thread-safety auditing that free-threading does.</p>\n<p><strong>Error messages continue improving</strong>, which has been a multi-release trend and has made Python meaningfully friendlier for people learning it.</p>\n<h2 id=\"the-assessment\">the assessment<a class=\"anchor\" href=\"#the-assessment\" aria-label=\"link to this section\">#</a></h2>\n<p>Python is executing a decade-long project to stop being single-threaded, and it is doing it without breaking the ecosystem, on an opt-in basis, with a fallback.</p>\n<p>That is the right way to do it and it is slower than anyone wants. The alternative \u2014 a hard break, a Python 4 \u2014 would have been faster and would have cost the community what Python 3 cost it, which nobody has the appetite for.</p>\n<p>Ask again in three years. The trajectory is good and the destination is not in doubt; only the timeline is.</p>",
   "image": "https://readme.news/cards/python-315-and-free-threadings-second-act.png",
   "date_published": "2026-05-11T09:00:00Z",
   "date_modified": "2026-05-11T09:00:00Z",
   "authors": [
    {
     "name": "Dom the Developer",
     "url": "https://domthedeveloper.com"
    }
   ],
   "tags": [
    "python",
    "languages",
    "releases"
   ]
  }
 ]
}