<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — opinion</title>
<link>https://readme.news/tags/opinion/</link>
<atom:link href="https://readme.news/tags/opinion/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged opinion.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>The migration that never finished</title><link>https://readme.news/the-migration-that-never-finished/</link><guid isPermaLink="true">https://readme.news/the-migration-that-never-finished/</guid><pubDate>Mon, 21 Sep 2026 09:00:00 +0000</pubDate><description>Half-completed migrations are the most expensive state a system can be in, and the most common one.</description><content:encoded><![CDATA[<p>Every mature codebase has at least one: the migration that got to 70% and stopped. The new system handles most traffic, the old one handles the awkward remainder, and both are maintained forever.</p>
<p>This is worse than either finishing or never starting, and it is the default outcome unless something prevents it.</p>
<h2 id="why-it-costs-more-than-both">why it costs more than both<a class="anchor" href="#why-it-costs-more-than-both" aria-label="link to this section">#</a></h2>
<p><strong>Two systems to maintain.</strong> Every change lands twice, or lands in one and silently diverges in the other.</p>
<p><strong>Two sets of bugs</strong>, plus a third set caused by the interaction.</p>
<p><strong>Nobody knows which is authoritative.</strong> New engineers ask; the answer is "it depends."</p>
<p><strong>The benefit never arrives.</strong> The reason for the migration — delete the old thing, get the performance, simplify the model — is only realised at 100%. At 70% you have paid the full cost and collected none of the return.</p>
<p><strong>It gets harder over time.</strong> The remaining 30% is the hard 30%: the weird integrations, the customer with the bespoke arrangement, the code nobody understands. And it gets harder as the people who understood the original migration leave.</p>
<h2 id="why-it-happens">why it happens<a class="anchor" href="#why-it-happens" aria-label="link to this section">#</a></h2>
<p>Not laziness. The incentives genuinely point this way:</p>
<p><strong>The easy 70% delivers most of the visible benefit.</strong> The graph goes up, the demo works, the announcement is made. The remaining 30% has no visible reward.</p>
<p><strong>The hard cases are hard for a reason.</strong> They were skipped because someone did not know how to handle them, and that has not changed.</p>
<p><strong>Priorities move.</strong> The migration was urgent in Q1. In Q3 there is a launch.</p>
<p><strong>Nobody owns the finish.</strong> The person who drove it moved on, and completion was never anyone's explicit goal.</p>
<h2 id="the-mechanisms-that-actually-work">the mechanisms that actually work<a class="anchor" href="#the-mechanisms-that-actually-work" aria-label="link to this section">#</a></h2>
<p><strong>Name the deletion, not the migration.</strong> The project is not "migrate to the new pricing service." It is <strong>"delete the old pricing service."</strong> The deliverable is the deletion, and the migration is how you get there. This one rewording changes what people track and what counts as done.</p>
<p><strong>Set a date, publicly, at the start.</strong> Not "by end of year" — a date, in the plan, with the deletion as the milestone. Dates without deletions slip silently; a date attached to "and then this code is gone" is checkable.</p>
<p><strong>Make the old path visibly worse.</strong> Log a warning on every use. Add latency — genuinely, deliberately. Put a banner in the internal tool. The old path being comfortable is why nobody leaves it.</p>
<p><strong>Count the stragglers on a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>.</strong> "Requests still on the legacy path" as a number that goes down, reviewed weekly. Anything not measured stalls at whatever level nobody notices.</p>
<p><strong>Do the hard cases first.</strong> Backwards from the usual instinct and correct. The easy 70% will always be doable; the hard 30% is what determines whether the project is possible at all. Finding out in month one that a case cannot be migrated is a much better outcome than finding out in month nine.</p>
<p><strong>Budget the finish before starting.</strong> If you cannot fund the last 30%, do not start. A migration you cannot complete is worse than the system you have.</p>
<h2 id="the-decision-worth-making-explicitly">the decision worth making explicitly<a class="anchor" href="#the-decision-worth-making-explicitly" aria-label="link to this section">#</a></h2>
<p>Sometimes finishing genuinely is not worth it. The remaining cases are rare, the old path works, and the effort is better spent elsewhere.</p>
<p>That is a legitimate call — but it has to be <em>made</em>, written down, and the state made permanent rather than provisional:</p>
<blockquote><p>The legacy importer stays for CSV uploads from the four enterprise customers using it. New integrations use the API. We are not migrating them; the importer is now a supported, frozen component with an owner, not a migration in progress. Reviewed 2026-09-21, revisit 2028.</p></blockquote>
<p>That is a completely different thing from a stalled migration, even though the code looks identical. One is a decision with an owner. The other is a mess nobody has admitted to.</p>
<p>The cost of the second one is not the code. It is that every engineer who touches that area has to reconstruct which way things are supposed to be going, and nobody can tell them.</p>]]></content:encoded></item><item><title>The interface is the product</title><link>https://readme.news/the-interface-is-the-product/</link><guid isPermaLink="true">https://readme.news/the-interface-is-the-product/</guid><pubDate>Mon, 31 Aug 2026 09:00:00 +0000</pubDate><description>Users cannot see your architecture. They can see the six seconds it takes to do the thing they came for.</description><content:encoded><![CDATA[<p>Engineers evaluate software by its internals: the architecture, the correctness, the elegance of the data model. Users evaluate it by the sequence of actions required to get what they came for.</p>
<p>These correlate less than we would like, and the gap is where a lot of otherwise good software fails.</p>
<h2 id="what-users-actually-experience">what users actually experience<a class="anchor" href="#what-users-actually-experience" aria-label="link to this section">#</a></h2>
<p>Not your service boundaries. Not your consistency model. Not the fact that you handle a partition correctly.</p>
<ul><li>How many steps to do the common thing.</li><li>Whether it responds instantly or after a spinner.</li><li>Whether an error tells them what to do.</li><li>Whether it does the same thing twice in a row.</li><li>Whether it remembers what they told it last time.</li></ul>
<p>Every one of those is an interface property, and every one is achievable on top of an ugly implementation — or destroyed on top of a beautiful one.</p>
<h2 id="the-trade-that-is-usually-made-backwards">the trade that is usually made backwards<a class="anchor" href="#the-trade-that-is-usually-made-backwards" aria-label="link to this section">#</a></h2>
<p>There is a real tension between internal cleanliness and external simplicity, and it comes up constantly:</p>
<ul><li>The clean data model exposes three concepts where users think in one.</li><li>The correct API makes the caller specify things they do not care about.</li><li>The properly-separated services mean the UI has to make four calls and handle each failing independently.</li><li>The general solution has eleven configuration options; the specific one would have had zero.</li></ul>
<p>The default resolution is to protect the internals and push the complexity outward, because the internals are what engineers look at and defend in review.</p>
<p>That is backwards. <strong>The interface is used far more times than the implementation is read.</strong> Complexity at the boundary is multiplied by every user and every call; complexity inside is paid by the people who chose it.</p>
<h2 id="what-this-looks-like-in-practice">what this looks like in practice<a class="anchor" href="#what-this-looks-like-in-practice" aria-label="link to this section">#</a></h2>
<p><strong>Collapse concepts at the boundary.</strong> If users think of one thing and your model has three, expose one and do the mapping internally. Yes, that is a lossy abstraction and you will occasionally have to break it. That is the correct place to put the pain.</p>
<p><strong>Default everything.</strong> Every required parameter is a decision you have forced on someone who has less context than you. Make it optional with a sensible default, and let the people who genuinely need control find the option.</p>
<p><strong>Make the common path one step.</strong> Count the actions for the thing 90% of users do 90% of the time. If it is more than two, that is the work.</p>
<p><strong>Absorb the failure.</strong> If your architecture means four things can fail independently, the interface should not surface four independent failures. Retry, degrade, or present one coherent state.</p>
<p><strong>Make it fast where they are waiting.</strong> A local read, an optimistic update, a cached response. Users cannot tell the difference between "fast because it is well-engineered" and "fast because you cheated" — and neither can anyone else.</p>
<h2 id="the-counterweight-honestly">the counterweight, honestly<a class="anchor" href="#the-counterweight-honestly" aria-label="link to this section">#</a></h2>
<p>This is not an argument for shipping a nice facade over a broken system. The internals are what make the interface <em>keep</em> working — under load, at the edges, after a year of changes. Software that is lovely to use and impossible to change dies just as reliably, only slower.</p>
<p>The claim is narrower: <strong>when the two genuinely conflict, and they do, the interface should usually win</strong> — because it is what the software is <em>for</em>, and because interface decisions are much more expensive to reverse than internal ones. You can rewrite the storage layer. You cannot easily take back a concept you taught a million users.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>Watch someone use your software for the first time without helping them.</p>
<p>Count the moments they hesitate. Each one is a place where your model and theirs diverged, and no amount of internal quality closes that gap.</p>
<p>That exercise is uncomfortable, takes twenty minutes, and consistently produces a better backlog than any amount of architectural discussion.</p>]]></content:encoded></item><item><title>Boring technology, revisited</title><link>https://readme.news/boring-technology-revisited/</link><guid isPermaLink="true">https://readme.news/boring-technology-revisited/</guid><pubDate>Sat, 15 Aug 2026 09:00:00 +0000</pubDate><description>The innovation-token argument is a decade old and still right, with one amendment nobody makes.</description><content:encoded><![CDATA[<p>The "choose boring technology" argument goes like this: you have a small number of innovation tokens. Spend them on the thing that is genuinely your problem. Everything else should be the option so well-understood that its failure modes are documented by strangers.</p>
<p>It has held up for a decade. It is worth restating because a decade is long enough that people have started applying it wrongly, in a specific and predictable way.</p>
<h2 id="why-it-works">why it works<a class="anchor" href="#why-it-works" aria-label="link to this section">#</a></h2>
<p><strong>Known failure modes.</strong> The value of a boring technology is not that it is good. It is that when it breaks at 3 a.m., the <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error message</a> has been pasted into a public forum by someone who then explained the fix.</p>
<p><strong>A hiring pool that already knows it.</strong> Every unusual choice is a thing you must teach every new engineer, forever.</p>
<p><strong>Operational surface you can reason about.</strong> Boring things have runbooks, monitoring integrations, backup tooling, and a decade of accumulated operational wisdom.</p>
<p><strong>Someone else finds the bugs.</strong> A widely-deployed system has been run at scales you will never reach, by people who reported what broke.</p>
<h2 id="the-amendment">the amendment<a class="anchor" href="#the-amendment" aria-label="link to this section">#</a></h2>
<p>Here is what the original argument does not say, and what a decade has taught: <strong>boring is a property of a technology at a moment in time, and it moves in both directions.</strong></p>
<p>Things become boring. Things also stop being boring — a project loses its maintainers, its community fragments, its corporate sponsor reorganises, the ecosystem moves and it does not.</p>
<p>So "choose boring" is not a decision you make once. It is a property you have to re-check, and the failure mode nobody plans for is the technology that was the safe choice in 2019 and is now a liability that everyone forgot to reconsider.</p>
<p><strong>The practical version:</strong> once a year, for each load-bearing dependency, ask — is this still boring? Is it still maintained, still hired for, still the thing a new engineer would expect? A "yes" costs a minute. A "no" is the most valuable thing you will learn that quarter.</p>
<h2 id="how-to-tell-whether-something-is-boring-now">how to tell whether something is boring <em>now</em><a class="anchor" href="#how-to-tell-whether-something-is-boring-now" aria-label="link to this section">#</a></h2>
<p>Not by age. Some old things are unmaintained; some five-year-old things are completely settled.</p>
<p>The signals that actually matter:</p>
<ul><li><strong>Multiple independent maintainers</strong>, ideally from more than one employer.</li><li><strong>A release in the last six months</strong>, and a <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a> that shows maintenance rather than churn.</li><li><strong>You can hire for it.</strong> Search jobs, not GitHub stars.</li><li><strong>Operational documentation written by users</strong>, not just by the vendor.</li><li><strong>A migration path off it</strong>, documented by people who took it. A technology nobody has successfully left is a technology you cannot leave.</li><li><strong>Boring failure modes.</strong> Can you name what happens when it runs out of memory, loses its network, or gets a corrupt config? If nobody has written that down, it is not boring yet.</li></ul>
<h2 id="where-to-spend-the-tokens">where to spend the tokens<a class="anchor" href="#where-to-spend-the-tokens" aria-label="link to this section">#</a></h2>
<p>The argument's real content is that innovation tokens are scarce, and the scarcity is not about technology risk. It is about <strong>attention</strong>.</p>
<p>Every unusual choice consumes attention: to learn it, to operate it, to debug it, to explain it. Attention is the actual constraint on an engineering team, and it is far more limited than the budget.</p>
<p>So spend the tokens where the unusual choice is the product. If your differentiator is a query engine, be adventurous about storage and utterly boring about everything else. If your differentiator is a workflow, use the most conventional stack in existence and put all the attention into the workflow.</p>
<p>The teams that get this wrong are almost never wrong about one big choice. They are wrong about six small ones, each individually defensible, that together consumed all the attention that should have gone into the thing customers pay for.</p>
<h2 id="the-counterweight">the counterweight<a class="anchor" href="#the-counterweight" aria-label="link to this section">#</a></h2>
<p>Taken too far this becomes an argument for never learning anything, and that is its own failure. A team that made every choice in 2016 and never revisited one is not disciplined, it is stuck — and the technology it chose has probably stopped being boring in the meantime.</p>
<p>The synthesis: be adventurous in one place at a time, deliberately, with a written reason. Be boring everywhere else. Re-check annually.</p>
<p>That is a harder discipline than either "always use the new thing" or "never use the new thing," and it is the only one that survives a decade.</p>]]></content:encoded></item><item><title>The service you should not have written</title><link>https://readme.news/the-service-you-should-not-have-written/</link><guid isPermaLink="true">https://readme.news/the-service-you-should-not-have-written/</guid><pubDate>Thu, 13 Aug 2026 09:00:00 +0000</pubDate><description>A checklist for the moment before you create a new deployable, when saying no is still free.</description><content:encoded><![CDATA[<p>Creating a new service feels like progress. It has a clean repository, no legacy, a fresh CI pipeline, and none of the compromises of the thing it is splitting away from.</p>
<p>It is also a permanent commitment that somebody will still be paying for in eight years. Here is the conversation worth having first.</p>
<h2 id="what-a-service-costs-in-full">what a service costs, in full<a class="anchor" href="#what-a-service-costs-in-full" aria-label="link to this section">#</a></h2>
<p>Not the code. The code is the cheap part.</p>
<ul><li>A repository, a build, a deploy pipeline, a rollback path.</li><li>A place to run, sized, with capacity headroom.</li><li>Monitoring, alerting, dashboards, an <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> runbook.</li><li>Secrets, credentials, and their rotation.</li><li>Dependency upgrades and security patching, forever.</li><li>A place in the request path that can now time out, and the circuit breaker and retry policy that implies.</li><li>Documentation, and an owner who is still at the company.</li><li>One more thing a new engineer must learn about.</li></ul>
<p>That list is the same whether the service is four hundred lines or forty thousand. The overhead is fixed, which is why small services are the ones whose economics are worst.</p>
<h2 id="the-questions-in-order">the questions, in order<a class="anchor" href="#the-questions-in-order" aria-label="link to this section">#</a></h2>
<p><strong>1. Could this be a module?</strong></p>
<p>The default answer for new functionality is a well-bounded module in something that already exists. It gets all of the above for free.</p>
<p>The follow-up that matters: <em>if this were a module with an enforced boundary, what would we lose?</em> If the honest answer is "nothing, it would just feel less tidy," write the module.</p>
<p><strong>2. What is the deployment argument?</strong></p>
<p>The strongest reason to extract is that a team needs to release on its own cadence and currently cannot. That is real and it scales with organisation size.</p>
<p>But check it: are they actually blocked, or do they merely deploy together? Two teams that could deploy independently and choose not to have a process problem, not an architecture problem, and a new service will not fix it.</p>
<p><strong>3. What is the scaling argument, with numbers?</strong></p>
<p>"It might need to scale differently" is not an argument. "This component is CPU-bound and spiky while the rest is I/O-bound and steady, and we are currently provisioning for the peak of both" is.</p>
<p>If you cannot state the resource profile that differs, there is no scaling argument.</p>
<p><strong>4. Where does the data live?</strong></p>
<p>The question that sinks most extractions. If the new service needs the same tables the old one does, you have not split anything — you have created two writers to one database, which is worse than one writer, and you now have to coordinate migrations across two deploy cycles.</p>
<p>A service that does not own its data is a distributed monolith with extra latency.</p>
<p><strong>5. What happens when it is down?</strong></p>
<p>If the answer is "the main flow breaks," you have added a failure mode and bought nothing in return. Fault isolation is not free; it is the result of deliberate <a class="xref" href="/timeouts-every-one-of-them/" title="Timeouts: every one of them">timeouts</a>, fallbacks and degradation, and if you are going to build those anyway you could have built them around a module.</p>
<p><strong>6. Who owns it in two years?</strong></p>
<p>Name the team. Not the person — the team. Services outlive their authors, and an ownerless service is the one that runs three major versions behind until it is a security incident.</p>
<h2 id="the-cases-where-the-answer-is-yes">the cases where the answer is yes<a class="anchor" href="#the-cases-where-the-answer-is-yes" aria-label="link to this section">#</a></h2>
<p>Being fair, because the reflex against splitting can be as unexamined as the reflex toward it.</p>
<ul><li>A genuinely different resource profile, with numbers.</li><li>A compliance or data-residency boundary that must be enforced structurally.</li><li>A component that has to be written in a different language for a real reason.</li><li>A team that is demonstrably blocked on someone else's release cadence.</li><li>Something with wildly different availability requirements — a batch job that can be down for an hour next to a checkout path that cannot.</li></ul>
<p>In every one of those, the service is buying something specific that a module cannot.</p>
<h2 id="the-reversibility-test">the reversibility test<a class="anchor" href="#the-reversibility-test" aria-label="link to this section">#</a></h2>
<p>Before you create it, ask: <strong>if this turns out to be wrong, how do we merge it back?</strong></p>
<p>If the answer is "we would not, we would just live with it," you are making a one-way decision on a hunch. Those deserve more scrutiny than they usually get, and the moment before the repository exists is the last time the scrutiny is free.</p>]]></content:encoded></item><item><title>Two hundred pieces in</title><link>https://readme.news/two-hundred-pieces-in/</link><guid isPermaLink="true">https://readme.news/two-hundred-pieces-in/</guid><pubDate>Mon, 03 Aug 2026 09:00:00 +0000</pubDate><description>What writing in public taught me about engineering, what I was most wrong about, and why the archive is the point.</description><content:encoded><![CDATA[<p>This is the two hundredth piece published here. That seems like a reasonable moment to write about the writing rather than about the subject.</p>
<h2 id="what-it-did-to-my-engineering">what it did to my engineering<a class="anchor" href="#what-it-did-to-my-engineering" aria-label="link to this section">#</a></h2>
<p><strong>It forced me to actually understand things.</strong> You can hold a fuzzy model of a technology in your head indefinitely and it feels like knowledge. The moment you try to explain it in a paragraph, the fuzziness becomes visible.</p>
<p>I have abandoned drafts because I discovered, four hundred words in, that I did not understand the thing well enough to write about it. Every one of those was more educational than the pieces I finished.</p>
<p><strong>It made me check things.</strong> Writing "X is faster than Y" in public means someone will ask for numbers. Knowing that in advance changes how you form the belief in the first place.</p>
<p><strong>It made me notice my own patterns.</strong> Reading two hundred pieces of my own writing back, the recurring themes are obvious to me now and were invisible while writing: verification over generation, boring over clever, measurement over intuition, and a persistent suspicion of anything that requires you to trust rather than check.</p>
<p>I did not set out with a thesis. It assembled itself.</p>
<p><strong>It taught me to be wrong in public</strong>, which is a skill and is uncomfortable and is the only way to find out you were wrong quickly.</p>
<h2 id="what-i-have-been-most-wrong-about">what I have been most wrong about<a class="anchor" href="#what-i-have-been-most-wrong-about" aria-label="link to this section">#</a></h2>
<p>I graded myself in December and the pattern held for the following eight months.</p>
<p><strong>I am reliably right about technical trajectories and reliably wrong about adoption.</strong></p>
<p>Local models got good; people did not switch, because hosted models got cheap faster than I expected. RAG got less necessary; the infrastructure repositioned instead of dying, which I have now failed to predict three separate times. Registry security controls were obviously needed; they arrived after the incident rather than before.</p>
<p>The lesson I keep relearning: the technology is the easy part to forecast, and the technology was never the hard part. Human and organizational behavior is where the uncertainty lives and where the consequences land.</p>
<p><strong>I under-predict inertia and over-predict rationality.</strong> Almost every wrong call has that shape.</p>
<h2 id="the-thing-about-writing-news">the thing about writing news<a class="anchor" href="#the-thing-about-writing-news" aria-label="link to this section">#</a></h2>
<p>Two hundred pieces, roughly half of them about things that happened in a specific week.</p>
<p>Reading them back, the ones that held up are almost never the ones that reported the event. They are the ones that used the event to explain a mechanism.</p>
<p>Nobody needs my summary of what a company announced. They can read the announcement. What is worth writing is: <em>why does this shape of thing keep happening</em>, and <em>what does it imply for what you should do on Monday</em>.</p>
<p>The news is a prompt. The mechanism is the article.</p>
<p>I did not know that when I started and it took about forty pieces to figure out.</p>
<h2 id="on-the-archive">on the archive<a class="anchor" href="#on-the-archive" aria-label="link to this section">#</a></h2>
<p>Everything published here is still at its original URL. Nothing has been quietly edited, renamed, or removed. Corrections are marked in place.</p>
<p>That is a deliberate choice and it costs something — there are pieces I would write differently now, and a few I think are wrong.</p>
<p>Leaving them up is the point. A publication that silently revises its history is not a record, it is a marketing surface. The value of an archive is that it shows what someone thought at the time, including when that was wrong, and you can only get that by not touching it.</p>
<p>If you want to know whether to trust a technical writer, check whether their old pieces still exist and whether the wrong ones were corrected in the open.</p>
<h2 id="on-the-format">on the format<a class="anchor" href="#on-the-format" aria-label="link to this section">#</a></h2>
<p>Gray background. Monospace headings. One column. No popups, no cookie banner, no newsletter modal, no autoplaying anything, no third-party JavaScript on any page.</p>
<p>This is not minimalism as an aesthetic. It is that every one of those things was added to a website to serve the publisher at the reader's expense, and the cumulative effect has made reading on the web genuinely unpleasant.</p>
<p>A page should render before you notice it loading. Text is the <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>. Nothing moves unless the reader moved it.</p>
<p>Those are not hard constraints to meet. Almost nobody meets them, and the reason is never technical.</p>
<h2 id="the-next-two-hundred">the next two hundred<a class="anchor" href="#the-next-two-hundred" aria-label="link to this section">#</a></h2>
<p>Same beat. Tech, developers, and the code underneath. News where the news teaches something, essays where the news does not.</p>
<p>More on the verification problem, because it is the defining engineering question of this period and it is nowhere near resolved. More on the craft, because the craft is what survives the tooling cycles. Fewer pieces about model launches, because they have stopped being informative.</p>
<p>Thanks for reading. Corrections are always welcome and get priority over everything else.</p>
<p>— Dom</p>]]></content:encoded></item><item><title>Deleting code is the highest-value work nobody schedules</title><link>https://readme.news/deleting-code-is-the-highest-value-work-nobody-schedules/</link><guid isPermaLink="true">https://readme.news/deleting-code-is-the-highest-value-work-nobody-schedules/</guid><pubDate>Wed, 29 Jul 2026 09:00:00 +0000</pubDate><description>Every line you remove is one that cannot break, cannot be misread, and does not need to be maintained.</description><content:encoded><![CDATA[<p>The most valuable pull request I have ever reviewed removed eleven thousand lines and added forty.</p>
<p>Deleting code is the only refactoring that is unambiguously good. It cannot introduce a bug in the deleted code, because there is no deleted code. It reduces build time, test time, cognitive load, security surface, and the probability that someone reads the wrong thing.</p>
<p>And nobody schedules it.</p>
<h2 id="what-accumulates">what accumulates<a class="anchor" href="#what-accumulates" aria-label="link to this section">#</a></h2>
<p><strong>Dead code.</strong> Never called. Nobody noticed, because nothing fails when unused code exists.</p>
<p><strong>Features nobody uses.</strong> Built for a customer who churned, an experiment that ended, a requirement that changed. Still there, still tested, still maintained, still appearing in every search result.</p>
<p><strong>Abandoned abstractions.</strong> Someone built a plugin system for the second plugin, which was never written. Now every call goes through a registry that has one entry.</p>
<p><strong>Configuration for conditions that no longer occur.</strong> A flag for a migration that completed in 2023.</p>
<p><strong>Compatibility shims.</strong> For a version nobody runs, an API that was removed, a browser that no longer exists.</p>
<p><strong>Tests for deleted behavior.</strong> Still running, still slow, testing something that cannot happen.</p>
<p><strong>Vendored copies.</strong> Of a library that is now a real dependency.</p>
<p><strong>Commented-out code.</strong> Always. Delete it. Git remembers.</p>
<h2 id="how-to-find-it">how to find it<a class="anchor" href="#how-to-find-it" aria-label="link to this section">#</a></h2>
<p><strong>Coverage, over a long window.</strong> Run coverage in production if your language supports it, or over your full integration suite. Anything at zero across a month is a candidate.</p>
<p>Be careful: zero coverage does not prove dead. It might be an error path, a rare branch, or something exercised only in a region you did not sample. Verify before deleting.</p>
<p><strong><a class="xref" href="/static-analysis-is-finally-worth-the-false-positives/" title="Static analysis is finally worth the false positives">Static analysis</a> for unreachable code.</strong> Most linters find unreferenced functions within a module. Cross-module dead code is harder and several tools do it.</p>
<p><strong><a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">Feature flags</a> at 100% for over a quarter.</strong> The disabled branch is dead code with a switch on it.</p>
<p><strong>Endpoints with no traffic.</strong> Log every route. Anything with zero requests in ninety days is a candidate. This is trivially easy to check and almost nobody does.</p>
<p><strong>Deprecation warnings nobody triggers.</strong> If you have been logging a deprecation warning for a year and it has never fired, the deprecated thing is unused.</p>
<p><strong><code>git log</code> on the file.</strong> Anything untouched for three years in an actively developed codebase is either perfect or forgotten.</p>
<h2 id="how-to-do-it-safely">how to do it safely<a class="anchor" href="#how-to-do-it-safely" aria-label="link to this section">#</a></h2>
<p><strong>Log before you delete.</strong> For anything you are not certain about, add logging and wait. A month of zero calls is strong evidence.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def old_thing(x):
    log.warning("old_thing called", stack=traceback.format_stack())
    return new_thing(x)</code></pre></div>
<p><strong>Delete in a separate commit</strong> from any other change. A deletion mixed with a refactor is unreviewable and un-revertable.</p>
<p><strong>Delete the tests too.</strong> Tests for deleted code are the most common thing left behind and they will confuse the next person enormously.</p>
<p><strong>Do not comment it out.</strong> Do not move it to an <code>old/</code> directory. Do not keep it "just in case." Git has it. If you genuinely might need it, note the commit hash in the deletion's commit message.</p>
<p><strong>Do it in batches, by area.</strong> One area, fully cleaned, is better than a thousand scattered deletions that are impossible to review.</p>
<h2 id="the-resistance-you-will-meet">the resistance you will meet<a class="anchor" href="#the-resistance-you-will-meet" aria-label="link to this section">#</a></h2>
<p><strong>"What if we need it?"</strong> You will not. And if you do, it is in git. In many years I have never seen a team need to recover deleted code that they could not recover.</p>
<p><strong>"Someone might be using it."</strong> Measure. That is what the logging is for. Guessing in either direction is worse than checking.</p>
<p><strong>"It works, why touch it?"</strong> Because it costs. Every line is read by every person who greps this file, is compiled on every build, is scanned by every security tool, and is a possible place for a future bug.</p>
<p><strong>"That is not a priority."</strong> Correct, and it never will be, which is why it has to be scheduled rather than prioritized. Ten percent of one sprint, quarterly, with a line count as the deliverable.</p>
<h2 id="the-framing-that-gets-it-done">the framing that gets it done<a class="anchor" href="#the-framing-that-gets-it-done" aria-label="link to this section">#</a></h2>
<p>Make it a competition. Track lines removed. Celebrate the largest deletion of the quarter.</p>
<p>This is slightly silly and it works, because it inverts the default incentive. Engineers are implicitly rewarded for adding — features shipped, code written — and never for removing, even though removing is frequently worth more.</p>
<p>Naming it, tracking it, and praising it is enough to change the behavior.</p>
<h2 id="the-number">the number<a class="anchor" href="#the-number" aria-label="link to this section">#</a></h2>
<p>A large fraction of most mature codebases is dead or effectively dead. I have never audited one where it was under 10%, and I have seen 40%.</p>
<p>Every line of that is being read, compiled, tested, scanned, and maintained, at a cost nobody has ever measured.</p>
<p>Go find some.</p>]]></content:encoded></item><item><title>Vendor lock-in: an honest cost model</title><link>https://readme.news/vendor-lock-in-an-honest-cost-model/</link><guid isPermaLink="true">https://readme.news/vendor-lock-in-an-honest-cost-model/</guid><pubDate>Fri, 17 Jul 2026 09:00:00 +0000</pubDate><description>Portability is not free and neither is dependence. A framework for deciding how much abstraction to buy.</description><content:encoded><![CDATA[<p>"Avoid vendor lock-in" is treated as self-evidently good advice. It is not advice, it is a preference, and following it uncritically produces systems that are worse in exchange for optionality nobody will ever exercise.</p>
<p>Here is a way to actually decide.</p>
<h2 id="the-cost-of-lock-in">the cost of lock-in<a class="anchor" href="#the-cost-of-lock-in" aria-label="link to this section">#</a></h2>
<p><strong>Switching cost.</strong> How much engineering time to move off, if you had to.</p>
<p><strong>Pricing power.</strong> A vendor who knows you cannot leave prices accordingly. This is real and it is usually the largest ongoing cost.</p>
<p><strong>Capability ceiling.</strong> You are limited to what they support, on their timeline.</p>
<p><strong>Correlated risk.</strong> They have an outage, you have an outage. They change their terms, you comply. They get acquired and sunset the product, you migrate on their schedule.</p>
<h2 id="the-cost-of-avoiding-lock-in">the cost of avoiding lock-in<a class="anchor" href="#the-cost-of-avoiding-lock-in" aria-label="link to this section">#</a></h2>
<p>This is the half that gets ignored, and it is frequently larger.</p>
<p><strong>The abstraction layer itself.</strong> Code to write, maintain, test, and debug through. It is a permanent tax and it makes stack traces longer.</p>
<p><strong>Lowest common denominator.</strong> Your abstraction can only expose what all candidate providers support. You give up the features that made the good option good.</p>
<p><strong>The abstraction is usually wrong anyway.</strong> It was designed against one provider's model. When you actually try to swap, you discover the abstraction encoded assumptions that do not hold, and you rewrite it.</p>
<p><strong>Delayed value.</strong> Time spent on portability is time not spent on the product.</p>
<h2 id="the-framework">the framework<a class="anchor" href="#the-framework" aria-label="link to this section">#</a></h2>
<p>For each dependency, estimate:</p>
<ol><li><strong>Switching cost</strong> — engineer-weeks to move.</li><li><strong>Probability you switch</strong> in the next three years.</li><li><strong>Cost of the abstraction</strong> — engineer-weeks now, plus ongoing drag.</li></ol>
<p>Then: if <code>switching_cost × probability &lt; abstraction_cost</code>, do not abstract.</p>
<p>The numbers are rough. The exercise still clarifies, because it forces you to state the probability out loud, and stated probabilities are usually much lower than the implied ones people are acting on.</p>
<h2 id="the-categories-worked-through">the categories, worked through<a class="anchor" href="#the-categories-worked-through" aria-label="link to this section">#</a></h2>
<p><strong>Object storage.</strong> Switching cost: low. The S3 API is a de facto standard and every provider implements it. Probability: moderate — people do move for pricing.</p>
<p><strong>Verdict: use the S3 API, do not abstract further.</strong> The API is already the abstraction.</p>
<p><strong>Compute.</strong> Switching cost: moderate if containerized, high if you use provider-specific serverless. Probability: low.</p>
<p><strong>Verdict: containerize</strong> — which is good practice anyway — <strong>and use whatever managed service you want.</strong> Do not build a compute abstraction layer.</p>
<p><strong>Relational database.</strong> Switching cost: high. Probability: low.</p>
<p><strong>Verdict: use the database's features.</strong> Teams that avoid stored procedures, database-specific types, and advanced indexing to stay portable are giving up real capability for an event that will not happen. Postgres-specific SQL is fine. You are not going to migrate to a different engine, and if you do, the SQL dialect will be the smallest part of the pain.</p>
<p><strong>Managed <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a>, streams, and similar.</strong> Switching cost: moderate. The semantics differ enough between providers that a thin abstraction genuinely helps.</p>
<p><strong>Verdict: a thin <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> — publish, subscribe, ack — is worth it.</strong> Not a full abstraction; a boundary.</p>
<p><strong>Authentication.</strong> Switching cost: very high — you have to migrate user credentials, sessions, and integrations. Probability: low, but the consequences of being stuck are severe.</p>
<p><strong>Verdict: use standard protocols.</strong> OIDC and SAML are the abstraction. A provider that supports them is replaceable in principle; one with a proprietary SDK is not.</p>
<p><strong>AI model providers.</strong> Switching cost: low if you kept the interface thin. Probability: <strong>high</strong> — the market is moving fast and the right choice changes quarterly.</p>
<p><strong>Verdict: definitely abstract.</strong> This is the clearest case on the list. A thin interface plus an eval harness in your own repository makes model changes an afternoon. Teams that did this in 2024 have been switching providers casually ever since.</p>
<p><strong>Observability.</strong> Switching cost: moderate to high — instrumentation is everywhere in your code. Probability: moderate, usually driven by cost.</p>
<p><strong>Verdict: OpenTelemetry.</strong> The instrumentation is vendor-neutral, the collector handles routing, and swapping backends is a configuration change.</p>
<h2 id="the-pattern">the pattern<a class="anchor" href="#the-pattern" aria-label="link to this section">#</a></h2>
<p><strong>Abstract where the switching probability is high and the abstraction is cheap.</strong> AI providers, observability backends, object storage.</p>
<p><strong>Do not abstract where switching is unlikely and the abstraction costs you real capability.</strong> Databases, compute platforms, managed services you chose for their specific features.</p>
<p><strong>Use standard protocols wherever they exist.</strong> OIDC, S3, OpenTelemetry, SQL. A standard is an abstraction someone else maintains, and it is always cheaper than yours.</p>
<h2 id="the-thing-that-actually-protects-you">the thing that actually protects you<a class="anchor" href="#the-thing-that-actually-protects-you" aria-label="link to this section">#</a></h2>
<p>Not an abstraction layer. <strong>Your data, in a format you can export, and a documented process for leaving.</strong></p>
<p>Ask, before adopting anything: can I get all my data out, in a usable format, without their cooperation? If yes, you have real optionality regardless of how coupled your code is, because the expensive part of a migration is never the code — it is the data.</p>
<p>If the answer is no, that is a much bigger red flag than any API coupling, and it is the question almost nobody asks during procurement.</p>]]></content:encoded></item><item><title>Cross-platform is a promise you make to your budget</title><link>https://readme.news/cross-platform-is-a-promise-you-make-to-your-budget/</link><guid isPermaLink="true">https://readme.news/cross-platform-is-a-promise-you-make-to-your-budget/</guid><pubDate>Tue, 30 Jun 2026 09:00:00 +0000</pubDate><description>Write once, run anywhere, debug everywhere. The honest accounting of what each approach actually costs.</description><content:encoded><![CDATA[<p>The cross-platform question — one codebase or several — gets argued as a technical matter and is mostly an organizational one. Here is the honest accounting.</p>
<h2 id="the-actual-trade">the actual trade<a class="anchor" href="#the-actual-trade" aria-label="link to this section">#</a></h2>
<p><strong>Native</strong> gives you: full platform capability, best performance, platform-idiomatic <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>, immediate access to new OS features, and the best debugging tools.</p>
<p>It costs you: two or three codebases, two or three teams, features implemented multiple times, and behavior that diverges over years in ways nobody tracks.</p>
<p><strong>Cross-platform</strong> gives you: one codebase, one team, features implemented once, consistent behavior.</p>
<p>It costs you: a framework layer between you and the platform, a lag before new OS features are available, worse debugging when the problem is in the bridge, and an interface that is either non-idiomatic on every platform or requires platform-specific work anyway.</p>
<h2 id="the-thing-the-arguments-miss">the thing the arguments miss<a class="anchor" href="#the-thing-the-arguments-miss" aria-label="link to this section">#</a></h2>
<p><strong>The largest cost of cross-platform is not performance. It is the <a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">escape hatch</a>.</strong></p>
<p>Everything is fine until you need something the framework does not support. Then you are writing platform-specific native code, plus a bridge, plus a fallback, plus tests for all three — and you have the complexity of native development <em>plus</em> the framework.</p>
<p>This happens. It always happens. The question is how often, and that depends entirely on what your app does.</p>
<p><strong>Low escape-hatch pressure:</strong> content, commerce, forms, dashboards, CRUD, most business applications. These use the standard widget set and standard capabilities. Cross-platform works well.</p>
<p><strong>High escape-hatch pressure:</strong> camera and media processing, background execution, Bluetooth and hardware peripherals, complex custom rendering, deep platform integration (widgets, shortcuts, app extensions), anything real-time.</p>
<p>For high-pressure apps, cross-platform frequently costs more than native, because you pay for the framework and then write the native code anyway.</p>
<p><strong>Be honest about which one you are</strong> before choosing. Most teams that regret their choice were high-pressure and assessed themselves as low.</p>
<h2 id="the-options-honestly">the options, honestly<a class="anchor" href="#the-options-honestly" aria-label="link to this section">#</a></h2>
<p><strong>React Native.</strong> Mature, large ecosystem, native widgets. The new architecture removed the old asynchronous bridge, which was the main performance complaint. Best choice if your team is already React.</p>
<p><strong>Flutter.</strong> Renders its own widgets, which means true visual consistency and a non-native feel that some users notice and most do not. Excellent performance for custom interfaces. Dart is a real adoption cost for a team that does not know it.</p>
<p><strong>Kotlin Multiplatform.</strong> Share business logic, write native UI. This is the approach I find most defensible: the logic layer — networking, models, validation, persistence — is where duplication is most wasteful and least visible to users, and the UI layer is where platform idiom matters most.</p>
<p>Requires two UI implementations, which is the point rather than a limitation.</p>
<p><strong>Web technologies in a wrapper.</strong> Fastest to build if you have web engineers, and the platform feel is the weakest. Fine for content-heavy applications, poor for anything interaction-heavy.</p>
<p><strong>Native.</strong> Still correct for a large category, and the category is smaller than native advocates believe.</p>
<h2 id="the-organizational-question-that-actually-decides-it">the organizational question that actually decides it<a class="anchor" href="#the-organizational-question-that-actually-decides-it" aria-label="link to this section">#</a></h2>
<p><strong>Do you have or can you hire two platform teams?</strong></p>
<p>If yes, native is viable and gives you the best result.</p>
<p>If no — and for most companies below a certain size the answer is no — the choice is between cross-platform and shipping on one platform. Framed that way, the decision is usually easy.</p>
<p><strong>What is your feature velocity?</strong></p>
<p>If you ship a large feature monthly, implementing it twice is a permanent 2× cost on your most expensive activity. If you ship quarterly, the duplication matters less.</p>
<p><strong>How much does platform idiom matter to your users?</strong></p>
<p>For a consumer app competing on polish: a lot. For an internal tool: nothing. For a B2B product where the buyer is not the user: less than you think.</p>
<h2 id="the-hybrid-that-most-people-should-consider">the hybrid that most people should consider<a class="anchor" href="#the-hybrid-that-most-people-should-consider" aria-label="link to this section">#</a></h2>
<p>Share the logic, write the UI natively.</p>
<p>The business logic — API clients, data models, validation, offline storage, sync, analytics — is genuinely identical across platforms and duplicating it produces bugs that exist on one platform and not the other, which are the worst bugs to diagnose.</p>
<p>The UI is where platform conventions matter, where users notice, and where the framework abstraction costs the most.</p>
<p>This is more work than full cross-platform and less than full native, and it puts the sharing where the value is.</p>
<h2 id="the-thing-i-would-tell-someone-deciding">the thing I would tell someone deciding<a class="anchor" href="#the-thing-i-would-tell-someone-deciding" aria-label="link to this section">#</a></h2>
<p>Prototype the hardest thing first.</p>
<p>Not the login screen. The thing you are worried about — the camera flow, the background sync, the complex list, the offline behavior. Build that on your candidate stack, in a week.</p>
<p>You will learn more from that week than from any amount of comparison, and you will learn it while changing your mind is still cheap.</p>
<p>The teams that regret their choice almost always chose based on a comparison article, built the easy part first, and discovered the hard part in month five.</p>]]></content:encoded></item><item><title>What a staff engineer does all day</title><link>https://readme.news/what-a-staff-engineer-does-all-day/</link><guid isPermaLink="true">https://readme.news/what-a-staff-engineer-does-all-day/</guid><pubDate>Fri, 19 Jun 2026 09:00:00 +0000</pubDate><description>The role is genuinely ambiguous and that ambiguity is load-bearing. An attempt at a concrete description.</description><content:encoded><![CDATA[<p>"Staff engineer" is the most poorly-defined common title in the industry. It means different things at different companies, and even within one company two staff engineers may do almost nothing in common.</p>
<p>Here is an attempt at what the role actually is, based on watching people do it well.</p>
<h2 id="what-it-is-not">what it is not<a class="anchor" href="#what-it-is-not" aria-label="link to this section">#</a></h2>
<p><strong>Not "a senior engineer who has been there longer."</strong> Time in seat produces a senior engineer with more context, which is valuable and is not this.</p>
<p><strong>Not "the best coder."</strong> The highest-output individual contributor is a valuable role and it is a different one. Staff engineers frequently write less code than seniors.</p>
<p><strong>Not "a manager who did not want to manage."</strong> The scope is comparable to a manager's; the mechanism is entirely different.</p>
<h2 id="what-it-is">what it is<a class="anchor" href="#what-it-is" aria-label="link to this section">#</a></h2>
<p>The concise version: <strong>a staff engineer is responsible for the technical success of work that spans more than one team, without having authority over those teams.</strong></p>
<p>Everything distinctive about the role follows from "without authority." You cannot assign work. You cannot approve headcount. You cannot make a decision stick by deciding it. Every outcome has to be achieved through information, credibility, and persuasion.</p>
<h2 id="the-actual-activities">the actual activities<a class="anchor" href="#the-actual-activities" aria-label="link to this section">#</a></h2>
<p><strong>Finding the problem nobody owns.</strong> The most common form of staff-level impact. Every organization has problems that fall between teams: the shared library nobody maintains, the integration that fails and each side thinks is the other's, the performance issue that is caused by the interaction of three services.</p>
<p>Nobody owns these, so nobody fixes them, so they persist for years. Identifying one, proving it matters, and getting it fixed is a substantial contribution and it usually requires a person who is not on any of the teams involved.</p>
<p><strong>Making a decision that spans teams.</strong> Which of three approaches, when each team has a preference and none of them can see the whole picture. This is where the "no authority" constraint bites hardest: the decision has to be made in a way that the people affected accept, which means the reasoning has to be visible and the objections have to be genuinely addressed.</p>
<p><strong>Writing the document that ends the argument.</strong> A recurring disagreement that resurfaces every quarter because nobody wrote down the resolution. One good document — with the options, the trade-offs, the decision, and the reasoning — can end a multi-year debate.</p>
<p><strong>Being the person who read the whole system.</strong> Most engineers know their team's code. Somebody needs to know how it all fits together, including the parts nobody has touched in three years. That knowledge is what makes the cross-cutting problems visible.</p>
<p><strong>Raising the floor.</strong> Not by writing better code. By making the good pattern easy: a library that removes a class of bug, a template that encodes the right defaults, a lint rule that prevents the mistake, documentation that means nobody has to ask. Impact through leverage rather than output.</p>
<p><strong>Mentoring, specifically on judgment.</strong> Not "how do I use this API." "Should we build this at all," "how do I disagree with my manager about a technical decision," "how do I tell if this design will be a problem in a year."</p>
<p><strong>Saying no with a reason, at a level where it lands.</strong> Frequently the most valuable thing a staff engineer does, and the reason it requires seniority is that the no has to come with an alternative and with credibility behind it.</p>
<h2 id="the-day-concretely">the day, concretely<a class="anchor" href="#the-day-concretely" aria-label="link to this section">#</a></h2>
<p>A real week looks roughly like:</p>
<ul><li>20% writing code, usually the hard or risky part of something, or a prototype that settles an argument.</li><li>25% writing documents — designs, decisions, analyses, postmortems.</li><li>25% in conversations — reviews, one-on-ones, arguing about designs, being asked "does this seem right to you."</li><li>15% reading — code, incident reports, other people's designs, the thing everyone is complaining about.</li><li>15% on whatever is currently on fire.</li></ul>
<p>The proportion of coding is the part that surprises people moving into the role, and the discomfort of not shipping visible code is the most common reason people bounce out of it.</p>
<h2 id="how-to-tell-if-someone-is-good-at-it">how to tell if someone is good at it<a class="anchor" href="#how-to-tell-if-someone-is-good-at-it" aria-label="link to this section">#</a></h2>
<p><strong>Do things get decided?</strong> Not "do they have opinions." Do arguments they are involved in reach a resolution that holds.</p>
<p><strong>Do other engineers get better?</strong> Look at the people around them over a year.</p>
<p><strong>Do they work on things nobody asked them to?</strong> The highest-value staff work is usually self-directed, because if it were obvious and assigned it would already have an owner.</p>
<p><strong>Are they trusted by people who disagree with them?</strong> This is the real test. Someone who is only trusted by people who already agree has influence, not credibility.</p>
<h2 id="the-failure-modes">the failure modes<a class="anchor" href="#the-failure-modes" aria-label="link to this section">#</a></h2>
<p><strong>Becoming an architecture astronaut.</strong> Designs, diagrams, opinions, no contact with running code. The credibility runs out within about a year, and it is very hard to get back.</p>
<p><strong>Becoming a very expensive senior engineer.</strong> Doing excellent work with single-team scope. Comfortable, valuable, and not the job.</p>
<p><strong>Spreading too thin.</strong> Involved in twelve things, effective in none. The scope is tempting and the constraint is real: two or three significant efforts at a time is the realistic maximum.</p>
<p><strong>Losing the ability to build.</strong> The role requires enough hands-on work to stay credible and calibrated. An engineer who has not shipped in a year is guessing.</p>]]></content:encoded></item><item><title>Rewrites: when they actually work</title><link>https://readme.news/rewrites-when-they-actually-work/</link><guid isPermaLink="true">https://readme.news/rewrites-when-they-actually-work/</guid><pubDate>Mon, 01 Jun 2026 09:00:00 +0000</pubDate><description>The received wisdom is never rewrite. The received wisdom is mostly right and has three real exceptions.</description><content:encoded><![CDATA[<p>The canonical advice is that rewriting from scratch is the single worst strategic mistake a software company can make. That advice is twenty-five years old and it is mostly still correct.</p>
<p>It has exceptions, and knowing which situation you are in matters more than knowing the rule.</p>
<h2 id="why-rewrites-usually-fail">why rewrites usually fail<a class="anchor" href="#why-rewrites-usually-fail" aria-label="link to this section">#</a></h2>
<p><strong>The old system's behavior is undocumented and load-bearing.</strong> Every strange conditional in a decade-old codebase is a bug report somebody filed. You will not find them by reading the code, because the code does not say why. You will find them by shipping the rewrite and having those bugs re-reported.</p>
<p><strong>The rewrite has no users, so it gets no feedback.</strong> The old system is being exercised by real traffic continuously. The new one is exercised by your test suite, which encodes what you think it should do.</p>
<p><strong>Feature parity is a moving target.</strong> The old system keeps getting features, because the business does not stop. You are chasing a target that recedes, and the chase consumes the time you budgeted for the rewrite.</p>
<p><strong>Nobody can justify continuing past month six.</strong> The rewrite has produced nothing users can see. The pressure to redirect the team to visible work is enormous and usually wins, leaving you with two systems.</p>
<p><strong>The knowledge is in the people, not the code.</strong> And the people who knew have frequently left, which is often why the rewrite was proposed.</p>
<h2 id="the-three-cases-where-it-is-right">the three cases where it is right<a class="anchor" href="#the-three-cases-where-it-is-right" aria-label="link to this section">#</a></h2>
<p><strong>1. The platform is dead.</strong></p>
<p>The framework is unmaintained, the language runtime is past end-of-life and getting CVEs, the vendor discontinued the product, the hardware is unobtainable.</p>
<p>This is not a preference — you are being forced. Rewrite, and be grateful you found out before the security incident.</p>
<p><strong>2. The domain model is fundamentally wrong.</strong></p>
<p>Not "the code is messy." The core abstractions do not correspond to reality, and every feature requires working around them.</p>
<p>The test: can you name a specific business capability that is <em>impossible</em>, not just awkward, in the current model? "We cannot support customers with more than one billing entity, and the fix requires changing what a customer is."</p>
<p>If you cannot name one, you have messy code, not a wrong model, and messy code is fixed by refactoring.</p>
<p><strong>3. The system is small enough that the rewrite is short.</strong></p>
<p>If the whole thing is three weeks of work, the calculus changes entirely. The risk of a three-week rewrite is bounded. Just do it.</p>
<p>The received wisdom is about large systems, and people apply it to small ones where it does not hold.</p>
<h2 id="how-to-do-it-when-you-must">how to do it when you must<a class="anchor" href="#how-to-do-it-when-you-must" aria-label="link to this section">#</a></h2>
<p><strong>Strangler fig, never big bang.</strong></p>
<p>Put a routing layer in front of the old system. Move one endpoint at a time to the new implementation. Route traffic gradually. When everything has moved, delete the old system.</p>
<div class="code"><pre><code>             ┌─→ new service (endpoints A, B)
client → router
             └─→ legacy monolith (everything else)</code></pre></div>
<p>Properties this gives you:</p>
<ul><li><strong>Value ships continuously.</strong> Every migrated endpoint is a delivered improvement.</li><li><strong>Risk is bounded per endpoint.</strong> If one goes wrong, route it back.</li><li><strong>The project survives leadership changes</strong>, because it is producing visible progress the whole time.</li><li><strong>You learn the old system's real behavior</strong> incrementally, at the point where you have to reimplement it.</li></ul>
<p><strong>Run both and compare.</strong> For a while, send traffic to both implementations, return the old one's response, and log the differences. This is the single most effective technique for discovering undocumented behavior, and it finds things no amount of code reading would have.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">old = legacy.compute(req)
try:
    new = rewritten.compute(req)
    if new != old:
        log.warn("divergence", request=req, old=old, new=new)
except Exception as e:
    log.error("new path failed", error=e, request=req)
return old      # old is still authoritative</code></pre></div>
<p>Run that for weeks. The divergence log is your actual specification.</p>
<p><strong>Freeze the old system's features.</strong> If the business keeps adding to it, you will never catch up. This requires an explicit organizational decision and it is usually the hardest part.</p>
<p><strong>Set a deletion date and mean it.</strong> The worst outcome is two systems forever. If the migration stalls at 70%, you have doubled your maintenance surface permanently.</p>
<h2 id="the-question-to-ask-first">the question to ask first<a class="anchor" href="#the-question-to-ask-first" aria-label="link to this section">#</a></h2>
<p>Before any rewrite: <strong>what specifically will be better, and how will you know?</strong></p>
<p>If the answer is "the code will be cleaner," that is not a business outcome and the project will lose its funding to something that is.</p>
<p>If the answer is "features in this area will take one week instead of three, and here are the last five that took three," you have a case.</p>
<p>The rewrites that succeed are the ones where somebody could state the payback in a sentence. The ones that fail are the ones motivated by taste, and taste is not wrong — it is just not fundable.</p>]]></content:encoded></item><item><title>Edge computing, honestly</title><link>https://readme.news/edge-computing-honestly/</link><guid isPermaLink="true">https://readme.news/edge-computing-honestly/</guid><pubDate>Fri, 29 May 2026 09:00:00 +0000</pubDate><description>Running code close to users is a real win for a narrow set of workloads and a complication for everything else.</description><content:encoded><![CDATA[<p>Edge computing has been sold as a general architectural improvement: run your code in hundreds of locations near users, everything gets faster.</p>
<p>The physics is real. The applicability is narrower than the marketing, and the reason is data.</p>
<h2 id="the-physics">the physics<a class="anchor" href="#the-physics" aria-label="link to this section">#</a></h2>
<p>Light in fiber travels roughly 200,000 km/s. New York to London and back is about 55 ms of pure propagation, before any processing. Add TLS handshakes, TCP setup, and real-world routing that is not a great circle, and a cross-Atlantic round trip is frequently 100 ms or more.</p>
<p>Running code 20 ms from the user instead of 120 ms is a genuine improvement and it is not achievable any other way.</p>
<h2 id="the-problem">the problem<a class="anchor" href="#the-problem" aria-label="link to this section">#</a></h2>
<p>Your data is not at the edge. It is in a database, in one region, and if your edge function needs it, you have moved the compute closer to the user and left the round trip in place — plus added a hop.</p>
<div class="code"><pre><code>user → edge (5ms) → origin database (120ms) → edge → user</code></pre></div>
<p>That is slower than the user talking to the origin directly, because you added a hop to a path that was always dominated by the database call.</p>
<p>This is the single most common edge computing mistake and it is easy to make, because the architecture diagram looks right.</p>
<h2 id="what-edge-is-genuinely-good-for">what edge is genuinely good for<a class="anchor" href="#what-edge-is-genuinely-good-for" aria-label="link to this section">#</a></h2>
<p><strong>Anything that needs no origin data:</strong></p>
<ul><li><strong>Redirects and rewrites.</strong> URL normalization, locale routing, legacy path mapping.</li><li><strong>Authentication token validation.</strong> A signed JWT can be verified with a public key at the edge, and an invalid request never reaches your origin. This is a real win — you reject bad traffic at the perimeter.</li><li><strong>A/B test assignment.</strong> Deterministic hash of a cookie into a bucket. No state required.</li><li><strong>Header manipulation.</strong> Security headers, CORS, feature policy.</li><li><strong>Bot filtering and <a class="xref" href="/rate-limiting-the-four-algorithms-and-when-each-is-wrong/" title="Rate limiting: the four algorithms and when each is wrong">rate limiting</a>.</strong> Reject at the edge, before the request costs you anything.</li><li><strong>Personalization of cached content.</strong> Fetch the cached page, inject the user's name from a cookie, return. The expensive part stays cached.</li></ul>
<p><strong>Anything where the data is genuinely replicated to the edge:</strong></p>
<p>Several platforms now offer edge-replicated key-value and SQL storage. If your data is small, read-heavy, and tolerant of replication lag — configuration, <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">feature flags</a>, product catalogs, translations — this works well and the latency win is real.</p>
<p>The constraints are real too: writes go to a primary, replication is eventual, and storage per location is limited.</p>
<h2 id="what-edge-is-bad-for">what edge is bad for<a class="anchor" href="#what-edge-is-bad-for" aria-label="link to this section">#</a></h2>
<p><strong>Anything write-heavy.</strong> Writes need coordination. Coordination needs a primary. The primary is in one place.</p>
<p><strong>Anything requiring strong consistency.</strong> By definition, this needs coordination, which needs round trips, which is what you were trying to avoid.</p>
<p><strong>Anything with a large working set.</strong> You cannot replicate a terabyte to three hundred locations.</p>
<p><strong>Anything computationally heavy.</strong> Edge runtimes have tight CPU and memory limits. They are designed for milliseconds of work per request.</p>
<p><strong>Anything that needs a specific runtime.</strong> Most edge platforms run a constrained JavaScript or <a class="xref" href="/the-component-model-and-the-plugin-problem/" title="The component model and the plugin problem">WebAssembly</a> environment. Your Python dependency with a C extension is not going there.</p>
<h2 id="the-architecture-that-works">the architecture that works<a class="anchor" href="#the-architecture-that-works" aria-label="link to this section">#</a></h2>
<p>Layered, with each layer doing what it is good at:</p>
<div class="code"><pre><code>edge      → auth check, rate limit, routing, cached content, header work
regional  → application logic, caching, session state
origin    → the database, the writes, the truth</code></pre></div>
<p>Most requests are answered at the edge from cache. Some go to a regional application tier. Few reach the origin.</p>
<p>That is a CDN with programmability, which is what edge computing actually is, and framing it that way produces much better decisions than framing it as "serverless everywhere."</p>
<h2 id="the-thing-to-measure-first">the thing to measure first<a class="anchor" href="#the-thing-to-measure-first" aria-label="link to this section">#</a></h2>
<p>Before adopting any of this: <strong>where does your latency actually go?</strong></p>
<p>Break down a typical request:</p>
<ul><li>DNS</li><li>TLS handshake</li><li>Network round trip</li><li>Time to first byte at origin</li><li>Origin processing</li><li>Database time within that</li><li>Response transfer</li></ul>
<p>If origin processing is 400 ms and network is 40 ms, moving compute to the edge addresses 40 ms of a 440 ms problem. Fix the 400 first.</p>
<p>This is the most common reason edge adoption disappoints: it was applied to a latency problem that was not a network problem.</p>
<h2 id="the-honest-summary">the honest summary<a class="anchor" href="#the-honest-summary" aria-label="link to this section">#</a></h2>
<p>Edge is a very good CDN with programmability, and that is genuinely valuable — it lets you do real work at the perimeter that used to require an origin request.</p>
<p>It is not a general application platform, and the platforms selling it as one are selling the constraint as a feature.</p>
<p>Use it for the perimeter. Keep your data where it can be consistent.</p>]]></content:encoded></item><item><title>Platform teams that don't get resented</title><link>https://readme.news/platform-teams-that-dont-get-resented/</link><guid isPermaLink="true">https://readme.news/platform-teams-that-dont-get-resented/</guid><pubDate>Fri, 22 May 2026 09:00:00 +0000</pubDate><description>Internal platforms fail for predictable reasons. The successful ones share four properties.</description><content:encoded><![CDATA[<p>Most internal platform teams end up resented by the engineers they serve. The pattern is consistent enough that the causes are identifiable.</p>
<h2 id="the-failure-pattern">the failure pattern<a class="anchor" href="#the-failure-pattern" aria-label="link to this section">#</a></h2>
<ol><li>Platform team forms to reduce duplicated infrastructure work.</li><li>They build an abstraction over the cloud provider.</li><li>The abstraction covers 80% of cases well.</li><li>The remaining 20% is impossible, and the escape hatch is either absent or punished.</li><li>Product teams work around the platform.</li><li>Platform team responds by mandating the platform.</li><li>Everyone is unhappy and the platform is now a tax.</li></ol>
<p>Every step follows from the previous one. The root is step four.</p>
<h2 id="the-four-properties-of-platforms-that-work">the four properties of platforms that work<a class="anchor" href="#the-four-properties-of-platforms-that-work" aria-label="link to this section">#</a></h2>
<p><strong>1. An escape hatch that is not punished.</strong></p>
<p>The platform covers the common case. It cannot cover every case, and pretending otherwise is what breaks trust.</p>
<p>There must be a supported path for "I need something the platform does not do," and taking that path must not require an exception process, a meeting, or an apologetic Slack message.</p>
<p>The best platforms make the escape hatch cheap and then compete on being better than it. The worst make it forbidden, which does not eliminate the need — it drives it underground.</p>
<p><strong>2. Adoption is voluntary, at least at first.</strong></p>
<p>A platform that teams choose is a platform that is good. A platform teams are required to use never gets the feedback that would make it good, because the feedback mechanism — people leaving — has been disabled.</p>
<p>If you cannot get voluntary adoption, that is information. Mandating it does not fix the underlying problem; it hides it and converts a product problem into a political one.</p>
<p>Mandate later, when it is genuinely better, and the mandate will be uncontroversial because everyone already uses it.</p>
<p><strong>3. The abstraction leaks deliberately, not accidentally.</strong></p>
<p>Every abstraction leaks. The question is whether you planned for it.</p>
<p>A good platform lets you drop a level when you need to: use the paved path for the deployment, and reach the underlying resource directly when you need something specific. A bad one hides the underlying system entirely, so that when it fails you cannot debug it and neither can the platform team, because now there are two systems to understand.</p>
<p><strong>Concretely:</strong> if your platform generates infrastructure configuration, let people see it. If it wraps a cloud API, let people access the underlying resource. If it runs their container, give them the logs from the actual runtime, not a filtered view.</p>
<p><strong>4. The platform team is measured on adoption and satisfaction, not on compliance.</strong></p>
<p>If the platform team's metric is "percentage of services on the platform," they will optimize for mandating it.</p>
<p>If the metric is "would you use this if you had a choice," they will optimize for making it good.</p>
<p>Ask that question quarterly, anonymously, and publish the answer.</p>
<h2 id="the-specific-things-that-generate-resentment">the specific things that generate resentment<a class="anchor" href="#the-specific-things-that-generate-resentment" aria-label="link to this section">#</a></h2>
<p><strong>Slow escape.</strong> A team needs something the platform does not support. The answer is "file a request, we will look at it next quarter." Their deadline is Friday.</p>
<p><strong>Breaking changes without migration paths.</strong> The platform is infrastructure. Break it and every team stops. Platform teams frequently hold themselves to a lower compatibility standard than they would accept from a vendor.</p>
<p><strong>Opaque failures.</strong> The deploy failed. The error is a platform-internal message. The product engineer cannot debug it and must escalate, which means waiting.</p>
<p><strong>Being a gate rather than a service.</strong> A platform that must approve things is a bureaucracy. A platform that makes the right thing easy is infrastructure.</p>
<p><strong>Solving the platform team's problems.</strong> Standardization is valuable to the platform team and is not automatically valuable to product teams. If the pitch for a migration is "this makes our lives easier," expect a cool reception.</p>
<h2 id="the-framing-that-works">the framing that works<a class="anchor" href="#the-framing-that-works" aria-label="link to this section">#</a></h2>
<p><strong>You are building a product. Your users are engineers. They have alternatives.</strong></p>
<p>That framing produces the right behaviors automatically: user research before building, documentation that assumes nothing, onboarding that works, support that responds, and a roadmap driven by what users need rather than by architectural preference.</p>
<p>The platform teams I have seen work best behave exactly like a startup selling to a skeptical market, and they say so out loud.</p>
<p>The ones that fail behave like an internal standards body, and they are usually correct about the standards and wrong about how to get them adopted.</p>
<h2 id="the-measurement-that-matters">the measurement that matters<a class="anchor" href="#the-measurement-that-matters" aria-label="link to this section">#</a></h2>
<p>Time from "a new engineer joins" to "their code is running in production."</p>
<p>That single number captures most of what a platform is for, it is measurable, and it is the thing product teams actually care about. If it is going down, the platform is working, regardless of what the adoption <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> says.</p>]]></content:encoded></item><item><title>Open source funding: five models, honestly compared</title><link>https://readme.news/open-source-funding-five-models-honestly-compared/</link><guid isPermaLink="true">https://readme.news/open-source-funding-five-models-honestly-compared/</guid><pubDate>Wed, 20 May 2026 09:00:00 +0000</pubDate><description>Donations, foundations, dual licensing, open core, and hosted service. Each works for a specific shape of project.</description><content:encoded><![CDATA[<p>The open source sustainability problem is not that the models do not exist. It is that projects pick a model that does not fit their shape, and then conclude that funding open source is impossible.</p>
<p>Five models. Each works. Each works for a specific kind of project.</p>
<h2 id="1-donations-and-sponsorship">1. donations and sponsorship<a class="anchor" href="#1-donations-and-sponsorship" aria-label="link to this section">#</a></h2>
<p><strong>Works for:</strong> developer-facing tools with a large individual user base and an identifiable maintainer.</p>
<p><strong>Does not work for:</strong> libraries deep in a dependency tree, infrastructure nobody knows they use, anything without a personality attached.</p>
<p>The honest arithmetic: donation income correlates with visibility, not with importance. A well-marketed CLI tool with a charismatic maintainer will out-earn a critical cryptography library by a large multiple.</p>
<p>The corporate sponsorship version works better than individual donations and requires the maintainer to do sales, which most maintainers are bad at and hate.</p>
<h2 id="2-foundation-governance">2. foundation governance<a class="anchor" href="#2-foundation-governance" aria-label="link to this section">#</a></h2>
<p><strong>Works for:</strong> infrastructure that multiple large companies depend on and that none of them wants a competitor to control.</p>
<p><strong>Does not work for:</strong> small projects. The overhead — governance, legal, trademark, process — is substantial, and a foundation with one project and no funded staff is just more paperwork.</p>
<p>The real value of a foundation is not money. It is neutrality: it makes a project safe for competitors to invest in together, which unlocks contribution that would not otherwise happen.</p>
<h2 id="3-dual-licensing">3. dual licensing<a class="anchor" href="#3-dual-licensing" aria-label="link to this section">#</a></h2>
<p><strong>Works for:</strong> libraries embedded in other products, where the copyleft obligation is genuinely inconvenient for commercial users.</p>
<p>Ship under a strong copyleft license, sell a commercial license to companies that cannot comply.</p>
<p><strong>Does not work for:</strong> anything permissively licensed already (no leverage), anything not embedded (the obligation does not bite), or anything with a permissive competitor of similar quality.</p>
<p>Effective when it fits, and it produces a genuine tension: the license that makes the business work is the one that limits adoption.</p>
<h2 id="4-open-core">4. open core<a class="anchor" href="#4-open-core" aria-label="link to this section">#</a></h2>
<p><strong>Works for:</strong> products where enterprise features are genuinely separable from the core — SSO, audit logs, RBAC, compliance reporting, multi-tenancy.</p>
<p><strong>Does not work for:</strong> libraries. There is no enterprise tier of a date-parsing library.</p>
<p>The failure mode is well documented: the line between core and commercial moves toward commercial over time, under revenue pressure, and the community that built your adoption watches features they use get moved behind the paywall.</p>
<p>If you do this, <strong>write down the line publicly, early, and honor it.</strong> "Anything that a single developer needs is open; anything that exists because you have a compliance department is commercial" is a defensible line. Moving it later costs more trust than the revenue is worth.</p>
<h2 id="5-hosted-service">5. hosted service<a class="anchor" href="#5-hosted-service" aria-label="link to this section">#</a></h2>
<p><strong>Works for:</strong> anything that is annoying to operate. Databases, search, <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a>, observability, CI.</p>
<p>Give away the software, sell the operation of it. This is the strongest model when it fits, because the value you sell — not having to run it — is real and continuous, and it does not require withholding anything.</p>
<p><strong>The risk:</strong> a hyperscaler offers a managed version of your software, at scale, without contributing back. This has happened repeatedly and it is why the source- available licenses exist.</p>
<p>Those licenses solve the problem and cost you the open source designation, which costs you contributors, ecosystem inclusion, and some corporate adoption. It is a real trade with real costs on both sides, and the projects that made it mostly survived, which is the empirical answer to whether it works.</p>
<h2 id="what-actually-kills-projects">what actually kills projects<a class="anchor" href="#what-actually-kills-projects" aria-label="link to this section">#</a></h2>
<p>Not the absence of a model. Three other things:</p>
<p><strong>Solo maintainer burnout.</strong> One person, unpaid, receiving an unbounded stream of issues, feature requests, and entitled demands. Funding helps and does not fix it — the fix is more maintainers, which is a governance problem.</p>
<p><strong>Success without support.</strong> A project that becomes critical infrastructure while its maintainer count stays at one. This is the most dangerous state and it is extremely common.</p>
<p><strong>Corporate abandonment.</strong> A company open-sources a project, staffs it with employees, then reorganizes. The external community was never built because it was never needed. Now nobody knows the code.</p>
<h2 id="what-companies-should-do">what companies should do<a class="anchor" href="#what-companies-should-do" aria-label="link to this section">#</a></h2>
<p>If your business depends on open source — and it does — the highest-leverage actions, in order:</p>
<ol><li><strong>Pay maintainers of your critical dependencies.</strong> Directly. Small amounts to many projects beat large amounts to a few.</li><li><strong>Assign employee time to upstream contribution.</strong> More valuable than money and much rarer.</li><li><strong>Do not send compliance questionnaires to volunteers.</strong> They owe you nothing and the license says so.</li><li><strong>When you fix a bug in a vendored dependency, upstream it.</strong> The number of companies carrying private patches for bugs everyone has is enormous.</li></ol>
<p>None of that requires a strategy document. It requires someone with budget deciding it matters, which is the actual bottleneck.</p>]]></content:encoded></item><item><title>Hiring juniors in 2026</title><link>https://readme.news/hiring-juniors-in-2026/</link><guid isPermaLink="true">https://readme.news/hiring-juniors-in-2026/</guid><pubDate>Thu, 30 Apr 2026 09:00:00 +0000</pubDate><description>The entry-level pipeline is breaking in a way that will be expensive in five years. Some of the fixes are cheap.</description><content:encoded><![CDATA[<p>Entry-level software hiring has contracted sharply. The reasons are partly cyclical and partly a genuine belief among hiring managers that AI tooling has reduced the need for junior engineers.</p>
<p>The first part will recover. The second part is a mistake, and it is the kind of mistake that is invisible for four years and then extremely expensive.</p>
<h2 id="the-argument-for-not-hiring-juniors">the argument for not hiring juniors<a class="anchor" href="#the-argument-for-not-hiring-juniors" aria-label="link to this section">#</a></h2>
<p>Stated honestly, because it is not stupid:</p>
<p>A junior engineer's first-year output is mostly small, well-specified tasks — the CRUD endpoint, the test coverage, the bug in a file someone pointed them at. That work is now substantially automatable. Meanwhile the junior requires mentoring time from senior engineers, which is the scarcest resource.</p>
<p>So the ROI on a junior looks worse than it did. That reasoning is coherent.</p>
<h2 id="why-it-is-wrong">why it is wrong<a class="anchor" href="#why-it-is-wrong" aria-label="link to this section">#</a></h2>
<p><strong>Seniors come from juniors.</strong> There is no other supply. An organization that hires only seniors is free-riding on other organizations' training, and if everyone does it, the pipeline empties. This is a classic collective action failure and the industry is walking into it with open eyes.</p>
<p><strong>The judgment that makes seniors valuable comes from doing the work.</strong> The ability to look at plausible code and know it is wrong comes from having written the wrong version and debugged it at 3 a.m. You cannot read your way to it and you cannot prompt your way to it.</p>
<p>If the apprenticeship stops, the next generation of senior engineers does not exist, and the current one retires.</p>
<p><strong>Juniors are better at the new tools.</strong> Consistently, in my experience. They have no prior workflow to defend and they explore. A team of only senior engineers adopts new tooling slowly and grudgingly.</p>
<p><strong>Mentoring makes seniors better.</strong> The engineer who has to explain why a design is wrong understands it better afterward. Teams with no juniors lose that forcing function and get sloppier about articulating their own reasoning.</p>
<h2 id="what-actually-has-to-change">what actually has to change<a class="anchor" href="#what-actually-has-to-change" aria-label="link to this section">#</a></h2>
<p>The old model — hire a junior, give them small tickets for a year, gradually increase scope — does not work as well when the small tickets are automated. The model has to change, not the hiring.</p>
<p><strong>Start them on reading, not writing.</strong> Give a new engineer a real system and a week to understand and explain it. Have them write the architecture document that does not exist. This builds the skill that actually matters now and it produces something useful.</p>
<p><strong>Give them debugging, not features.</strong> Debugging is the skill that generalizes, that AI is least reliable at, and that cannot be learned from a course. Pair them on incidents. Give them the flaky test nobody wants.</p>
<p><strong>Make them review agent output.</strong> Reviewing machine-generated code with a senior engineer walking through what is wrong with it is an extraordinarily efficient teaching mechanism. You get a stream of plausible-but-flawed code, which is exactly the training material you want and which used to be expensive to produce.</p>
<p><strong>Require them to write the tests first.</strong> Specifying behavior before implementing teaches design, and it is the part of the workflow that has become more important rather than less.</p>
<p><strong>Do not let them delegate the hard part.</strong> For the first year, some things get done by hand, deliberately, because the point is the learning rather than the output. Say this out loud so it does not feel like an arbitrary restriction.</p>
<h2 id="the-hiring-signal-that-works-now">the hiring signal that works now<a class="anchor" href="#the-hiring-signal-that-works-now" aria-label="link to this section">#</a></h2>
<p>Traditional junior screens — implement this algorithm, complete this take-home — are substantially defeated and were never good predictors anyway.</p>
<p>What works better:</p>
<p><strong>A code review exercise.</strong> Give them a pull request with three problems. Watch what they find and how they talk about it. This is the job.</p>
<p><strong>A debugging exercise on a real repository.</strong> Failing test, thirty minutes, any tools they want including AI. Watch the process, not the outcome. Do they read the error? Form a hypothesis? Check it? Notice when the model's suggestion is wrong?</p>
<p><strong>A conversation about something they built.</strong> Follow-up questions until you hit the edge of their understanding. Where that edge sits, and how they handle reaching it, tells you almost everything.</p>
<h2 id="the-case-to-make-internally">the case to make internally<a class="anchor" href="#the-case-to-make-internally" aria-label="link to this section">#</a></h2>
<p>If you are arguing for junior headcount:</p>
<p>The cost of a junior is roughly a senior's partial attention for a year plus a below-market salary. The cost of a senior hire in three years, in a market where nobody trained anyone, is going to be considerably higher than it is now.</p>
<p>Every organization that stopped training in 2009 spent 2013 through 2016 paying enormous premiums for the engineers who had been trained elsewhere. It is the same trade and it is being made again.</p>
<p>The organizations that keep training through this period will have a meaningful advantage in five years, and it will be very hard to catch up to them quickly.</p>]]></content:encoded></item><item><title>The three kinds of technical debt</title><link>https://readme.news/the-three-kinds-of-technical-debt/</link><guid isPermaLink="true">https://readme.news/the-three-kinds-of-technical-debt/</guid><pubDate>Fri, 24 Apr 2026 09:00:00 +0000</pubDate><description>Deliberate, accidental, and structural. They need completely different responses and everyone calls them the same thing.</description><content:encoded><![CDATA[<p>"Technical debt" has become a phrase that means "code I do not like," which has made it useless in the conversations where it matters — the ones where you are asking for time to fix something.</p>
<p>Three distinct things wear the label. They have different causes, different costs, and different correct responses.</p>
<h2 id="1-deliberate-debt">1. deliberate debt<a class="anchor" href="#1-deliberate-debt" aria-label="link to this section">#</a></h2>
<p>You knew the better solution. You chose the faster one on purpose, for a reason.</p>
<p>"We are hard-coding this to hit the deadline. We will parameterize it in Q3."</p>
<p>This is the original meaning of the metaphor and it is a legitimate engineering tool. Borrowing against future effort to ship now is frequently correct, especially when you are not yet sure the thing will survive.</p>
<p><strong>The failure is not taking on the debt. It is not recording it.</strong></p>
<p>What makes deliberate debt manageable:</p>
<ul><li>A comment at the site, explaining the trade-off and the condition that would trigger the fix.</li><li>A ticket, linked from the comment.</li><li>A named condition: "when we have more than five customers on this path" is much better than "later."</li></ul>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python"># DEBT: hardcoded to the US tax table to ship for the March launch.
# Parameterize when we take our first non-US customer. See ENG-4417.
TAX_RATE = 0.0725</code></pre></div>
<p>That comment costs thirty seconds and it is the difference between debt and mystery.</p>
<h2 id="2-accidental-debt">2. accidental debt<a class="anchor" href="#2-accidental-debt" aria-label="link to this section">#</a></h2>
<p>You did not know better at the time. The requirements changed. The library you chose turned out to be wrong. The abstraction fit the problem you had and not the one you have now.</p>
<p>This is the largest category and it is not anyone's fault. It is the natural consequence of building things under uncertainty.</p>
<p><strong>The response is refactoring as part of ordinary work</strong>, not as a project.</p>
<p>The practice that works: when you touch a file, leave it slightly better. Not a rewrite — rename the confusing variable, extract the function that is doing two things, delete the dead branch. Small, continuous, in the same commit as the feature.</p>
<p>Refactoring projects — a quarter dedicated to cleanup — mostly fail. They are unfunded after the first month, they conflict with in-flight work, and they produce large risky changes with no user-visible benefit to justify them.</p>
<p>Continuous small improvement compounds and is nearly invisible in the process. That is the whole trick.</p>
<h2 id="3-structural-debt">3. structural debt<a class="anchor" href="#3-structural-debt" aria-label="link to this section">#</a></h2>
<p>The architecture is wrong for what the system now does.</p>
<p>The monolith needs to be split, or the microservices need to be merged. The data model does not represent the domain. The synchronous design cannot support the scale. The framework choice from 2018 is blocking everything.</p>
<p>This is the expensive category and it is qualitatively different from the other two, because <strong>you cannot fix it incrementally without a plan.</strong> Small improvements to a wrong structure make the wrong structure more entrenched.</p>
<p><strong>The response is a real project, with a real justification, sized honestly.</strong></p>
<p>That requires:</p>
<ul><li><strong>A specific cost statement.</strong> Not "the architecture is bad." "Every new feature in this area takes three times as long as an equivalent feature elsewhere, and here are the last four examples with dates."</li><li><strong>A specific benefit.</strong> What becomes possible or fast afterward.</li><li><strong>A path with intermediate value.</strong> A twelve-month rewrite with no deliverable until month twelve will be cancelled in month seven. Structure it so each phase ships something.</li><li><strong>A strangler pattern, not a rewrite.</strong> Build the new alongside the old, migrate incrementally, delete the old. Big-bang rewrites have a well-documented failure rate and the reason is always the same: the old system's behavior includes a decade of undocumented edge cases nobody enumerated.</li></ul>
<h2 id="how-to-talk-about-it">how to talk about it<a class="anchor" href="#how-to-talk-about-it" aria-label="link to this section">#</a></h2>
<p>The conversation that fails: "we need to spend time on technical debt."</p>
<p>The conversation that works: "the last four features in the billing area each took about three weeks. Equivalent features elsewhere take one. The difference is the data model, specifically that subscriptions and invoices share a table. Fixing it is about six weeks and would bring billing features back to normal velocity. We have eleven billing features on the roadmap."</p>
<p>The second version has: a measurement, a cause, a cost, and a payback period. It is a business case, and business cases get funded.</p>
<p>The first version is a complaint, and complaints do not.</p>
<h2 id="the-thing-nobody-says">the thing nobody says<a class="anchor" href="#the-thing-nobody-says" aria-label="link to this section">#</a></h2>
<p>Some debt should never be paid.</p>
<p>Code in a system that will be retired, or that nobody has changed in three years, or that works and has no pending requirements — leave it. Ugly code that is stable and unmodified costs you nothing. It is not debt if you never pay interest on it.</p>
<p>The question is never "is this code good." It is "is this code costing us anything." A lot of what gets called technical debt is code that offends someone's taste in a module nobody touches, and refactoring it is a hobby, not engineering.</p>]]></content:encoded></item><item><title>You still do not need Kubernetes</title><link>https://readme.news/you-still-do-not-need-kubernetes/</link><guid isPermaLink="true">https://readme.news/you-still-do-not-need-kubernetes/</guid><pubDate>Mon, 13 Apr 2026 09:00:00 +0000</pubDate><description>It&#x27;s excellent software solving a real problem that most teams do not have. The honest threshold, and what to do below it.</description><content:encoded><![CDATA[<p>Kubernetes is genuinely good software. It solves a real problem well. It has an enormous ecosystem and a large pool of people who know it.</p>
<p>It is also, for a majority of the teams running it, a substantial amount of complexity in exchange for benefits they do not receive, and saying so is still mildly heretical.</p>
<h2 id="the-problem-it-actually-solves">the problem it actually solves<a class="anchor" href="#the-problem-it-actually-solves" aria-label="link to this section">#</a></h2>
<p>Kubernetes was built for: many services, many teams, heterogeneous workloads, on a fleet of machines, where you want bin-packing efficiency and declarative self-healing, and where the platform is operated by people whose job that is.</p>
<p>If you have all of those, it is the right answer and there is no close second.</p>
<h2 id="what-it-costs">what it costs<a class="anchor" href="#what-it-costs" aria-label="link to this section">#</a></h2>
<p><strong>A permanent learning tax.</strong> Pods, deployments, services, ingresses, configmaps, secrets, persistent volume claims, storage classes, service accounts, roles, network policies, resource quotas, and a YAML dialect for each. Every engineer who deploys anything must learn a meaningful fraction of it.</p>
<p><strong>Operational surface.</strong> Control plane upgrades, node upgrades, CNI plugin, CSI driver, ingress controller, cert manager, metrics server, log shipper. Each is a component that can break and that must be upgraded on someone else's schedule.</p>
<p><strong>Debugging distance.</strong> "Why is my service not reachable" has a dozen possible answers across five layers, and diagnosing it requires understanding all of them.</p>
<p><strong>Cost, frequently.</strong> A managed control plane plus nodes sized for the platform's own overhead plus the observability stack it needs is often more than the equivalent capacity on simpler infrastructure.</p>
<p><strong>Resume-driven adoption.</strong> This is real and worth naming. Kubernetes on your CV is worth money. That is a genuine incentive pointed away from the simplest solution that works.</p>
<h2 id="the-honest-threshold">the honest threshold<a class="anchor" href="#the-honest-threshold" aria-label="link to this section">#</a></h2>
<p>You probably want Kubernetes if:</p>
<ul><li>More than roughly fifteen to twenty distinct services, deployed independently.</li><li>More than a handful of teams that need to deploy without coordinating.</li><li>You have someone whose job includes operating the platform, not as a side task.</li><li>Genuinely heterogeneous workloads with different scaling characteristics.</li><li>Multi-tenancy requirements with real isolation needs.</li></ul>
<p>You probably do not if:</p>
<ul><li>Under ten services.</li><li>One or two teams.</li><li>Nobody owns the platform.</li><li>Traffic is predictable.</li><li>You are running one application with a database.</li></ul>
<h2 id="what-to-do-below-the-threshold">what to do below the threshold<a class="anchor" href="#what-to-do-below-the-threshold" aria-label="link to this section">#</a></h2>
<p>The options are better than they were, and all of them are boring:</p>
<p><strong>A platform-as-a-service.</strong> Push code, it runs. This is the correct answer for a very large number of applications and the reason people avoid it is usually aesthetic.</p>
<p><strong>Containers on a managed container service</strong> without the orchestrator — the various "run this container, scale it, load balance it" products every cloud offers. You get containers, autoscaling, and rolling deploys without the platform.</p>
<p><strong>A couple of servers and a process manager.</strong> Systemd units, a reverse proxy with automatic certificates, and a deploy script. This runs an enormous amount of traffic, is trivially debuggable, and every engineer already understands it.</p>
<p><strong>Docker Compose on one machine.</strong> For staging, for internal tools, for anything where a single host is enough. Unfashionable, works.</p>
<h2 id="the-migration-path-argument">the migration-path argument<a class="anchor" href="#the-migration-path-argument" aria-label="link to this section">#</a></h2>
<p>"We will need Kubernetes eventually, so we should start now."</p>
<p>This is the most common justification and it is usually wrong, for two reasons.</p>
<p>First, the complexity cost is paid every day from now until then, and "eventually" frequently never arrives.</p>
<p>Second, containerizing your application is the actual hard part of a future migration, and you can do that without an orchestrator. A containerized application running under a simple process manager can move to Kubernetes later in a couple of weeks.</p>
<p>Build the container. Skip the platform until you have the problem it solves.</p>
<h2 id="the-position-i-will-defend">the position I will defend<a class="anchor" href="#the-position-i-will-defend" aria-label="link to this section">#</a></h2>
<p>The teams I have seen most successfully run Kubernetes are large organizations with dedicated <a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">platform teams</a>, where it is genuinely the right tool.</p>
<p>The teams I have seen most damaged by it are small ones where a single engineer set it up, that engineer left, and the remaining team is operating a system nobody understands and is afraid to touch.</p>
<p>That second failure is common, expensive, and entirely predictable from the staffing at adoption time. If nobody's job is going to be operating the platform, you should not have a platform.</p>]]></content:encoded></item><item><title>The database you should have chosen</title><link>https://readme.news/the-database-you-should-have-chosen/</link><guid isPermaLink="true">https://readme.news/the-database-you-should-have-chosen/</guid><pubDate>Mon, 30 Mar 2026 09:00:00 +0000</pubDate><description>Postgres. That&#x27;s the article. Here&#x27;s the longer version, including the cases where it&#x27;s wrong.</description><content:encoded><![CDATA[<p>For a new application, the default database choice is Postgres, and the burden of proof is on anything else.</p>
<p>This is not a controversial position anymore, which is itself notable, because a decade ago it was.</p>
<h2 id="what-it-absorbed">what it absorbed<a class="anchor" href="#what-it-absorbed" aria-label="link to this section">#</a></h2>
<p>The reason the argument ended is that Postgres kept adding the things people left it for:</p>
<p><strong>JSON.</strong> <code>jsonb</code> with indexing, operators, and path queries. The document database use case — schemaless-ish data with nested structure — is handled well enough that the specialized option is rarely worth a second system.</p>
<p><strong>Full-text search.</strong> Built in, with ranking, stemming, and index support. Not as good as a dedicated search engine at large scale or with complex relevance requirements. Good enough for the search box on most applications, and one fewer system to operate.</p>
<p><strong>Vector search.</strong> <code>pgvector</code> with HNSW indexing. Again: not the best available at extreme scale, and entirely adequate for most retrieval workloads, with the enormous advantage that your vectors sit next to your metadata and you can filter and join in one query.</p>
<p>That last point is underrated. A dedicated vector database that cannot join to your relational data means you do the join in application code, badly.</p>
<p><strong>Time series.</strong> Partitioning, BRIN indexes, and extensions cover a lot of the use case.</p>
<p><strong>Geospatial.</strong> PostGIS remains the best geospatial implementation in any database, period.</p>
<p><strong><a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">Queues</a>.</strong> <code>SELECT ... FOR UPDATE SKIP LOCKED</code> gives you a correct, transactional job queue in about twenty lines. For anything under high-thousands of jobs per second — which is most systems — you do not need a dedicated queue, and the transactional property (enqueue in the same transaction as the state change) removes an entire class of consistency bug.</p>
<h2 id="the-operational-reality">the operational reality<a class="anchor" href="#the-operational-reality" aria-label="link to this section">#</a></h2>
<p><strong>One system to operate.</strong> Every additional datastore is a backup strategy, a monitoring setup, an upgrade path, a failure mode, an <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> runbook, and a set of consistency questions between it and everything else.</p>
<p>The cost of a second datastore is not the license. It is the permanent operational surface, and teams consistently underestimate it by a large factor.</p>
<p><strong>Transactions across your data.</strong> If your users and your documents and your embeddings are in one database, a change to all three is one transaction. Split across three systems, it is a distributed consistency problem you now own.</p>
<h2 id="when-it-is-genuinely-wrong">when it is genuinely wrong<a class="anchor" href="#when-it-is-genuinely-wrong" aria-label="link to this section">#</a></h2>
<p>Being honest about this matters, because "just use Postgres" as a reflex is the same failure mode as any other reflex.</p>
<p><strong>Extreme write throughput on a single logical dataset.</strong> Postgres scales writes vertically, well, up to a point. Past that point you need sharding, and Postgres's sharding story is real but less mature than systems designed for it from the start. If you genuinely need millions of writes per second, look elsewhere.</p>
<p><strong>Analytical queries over very large datasets.</strong> Postgres is a row store. Scanning a billion rows to compute an aggregate is what column stores are for, and the difference is orders of magnitude. Use a column store — several integrate cleanly with Postgres.</p>
<p><strong>Global multi-region with low write latency everywhere.</strong> Postgres replication is primary-based. If you need writes accepted in multiple regions with low latency, that is a distributed database problem and Postgres is not one.</p>
<p><strong>Extremely high-volume <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a>.</strong> Redis exists and is better at being a cache. This is a legitimate second system.</p>
<p><strong>Embedded / <a class="xref" href="/pixel-10-and-the-on-device-model-as-a-platform-feature/" title="Pixel 10 and the on-device model as a platform feature">on-device</a>.</strong> SQLite. Different problem, different answer.</p>
<h2 id="the-thing-to-actually-do">the thing to actually do<a class="anchor" href="#the-thing-to-actually-do" aria-label="link to this section">#</a></h2>
<p>Start with Postgres. Put everything in it. Measure.</p>
<p>When something is genuinely the bottleneck — and you will know, because you will have the metrics — extract that one thing to a specialized system with a clear reason.</p>
<p>That order matters. Teams that start with five specialized systems on the theory that they will need them end up operating five systems for a workload one would have handled, and the complexity is permanent.</p>
<h2 id="the-version-note">the version note<a class="anchor" href="#the-version-note" aria-label="link to this section">#</a></h2>
<p>Whatever you are running, be less than two major versions behind. Postgres's release quality is high, the upgrades are usually boring, and the performance improvements between versions are substantial and free.</p>
<p>The people who have bad Postgres upgrade experiences are the ones who skipped four versions and tried to do it in one jump during an outage.</p>]]></content:encoded></item><item><title>The state of self-hosting</title><link>https://readme.news/the-state-of-self-hosting/</link><guid isPermaLink="true">https://readme.news/the-state-of-self-hosting/</guid><pubDate>Fri, 20 Mar 2026 09:00:00 +0000</pubDate><description>Running your own infrastructure got dramatically easier while the industry was arguing about the cloud. A practical assessment.</description><content:encoded><![CDATA[<p>The default answer to "where should this run" has been "the cloud" for fifteen years, and for most of that time it was correct.</p>
<p>Several things changed and the answer is now more nuanced than the reflex suggests.</p>
<h2 id="what-changed-in-favor-of-self-hosting">what changed in favor of self-hosting<a class="anchor" href="#what-changed-in-favor-of-self-hosting" aria-label="link to this section">#</a></h2>
<p><strong>Machines got enormous.</strong> A single server you can rent for a few hundred dollars a month has more cores, more memory, and dramatically more I/O than a rack of hardware from 2012. A very large number of applications fit on one machine with room to spare.</p>
<p><strong>The tooling got good.</strong> Configuration management, container runtimes, reverse proxies with automatic certificates, backup tooling. What required a team a decade ago requires a competent person and a weekend.</p>
<p><strong>Cloud egress pricing did not fall.</strong> Compute prices came down. Bandwidth pricing at the major clouds is still a large multiple of what it costs, and for bandwidth-heavy applications it dominates the bill.</p>
<p><strong>Managed service prices are high relative to the alternative.</strong> A managed database costs several times what the equivalent instance costs, for operational convenience that is real and is not always worth the multiple.</p>
<h2 id="what-changed-against-it">what changed against it<a class="anchor" href="#what-changed-against-it" aria-label="link to this section">#</a></h2>
<p><strong>Security expectations rose.</strong> Patching, hardening, monitoring, incident response. Self-hosting means you own all of it, and the threat environment is worse than it was.</p>
<p><strong>Compliance frameworks assume cloud controls.</strong> SOC 2, ISO 27001, and their relatives are achievable self-hosted and the evidence collection is more work.</p>
<p><strong>The talent assumption inverted.</strong> A decade ago every team had someone who knew Linux systems administration. Now a lot of teams do not, and hiring for it is harder than hiring for cloud skills.</p>
<h2 id="the-honest-decision-framework">the honest decision framework<a class="anchor" href="#the-honest-decision-framework" aria-label="link to this section">#</a></h2>
<p><strong>Self-host when:</strong></p>
<ul><li>Your workload is steady rather than spiky. Cloud's core value proposition is elasticity, and you are paying for elasticity you do not use.</li><li>Bandwidth is a large share of your bill.</li><li>You have or can hire operational competence.</li><li>Data locality or sovereignty is a requirement.</li><li>You are at a scale where the cloud premium is a meaningful number — which starts lower than most people assume.</li></ul>
<p><strong>Use the cloud when:</strong></p>
<ul><li>Traffic is spiky or unpredictable.</li><li>You are early and optimizing for speed of iteration over unit economics.</li><li>You need global presence and do not want to operate it.</li><li>Your team's time is better spent on the product, which for most early-stage companies it is.</li><li>Compliance requirements are easier to satisfy with a provider's attestations.</li></ul>
<p><strong>The hybrid that most people should consider:</strong> run the steady baseline on owned or rented hardware, burst to cloud for peaks, keep object storage and CDN with a provider. This captures most of the cost advantage without giving up elasticity where it matters.</p>
<h2 id="the-middle-option-nobody-talks-about">the middle option nobody talks about<a class="anchor" href="#the-middle-option-nobody-talks-about" aria-label="link to this section">#</a></h2>
<p>Between "hyperscaler" and "rack in a colo" there is a large market of dedicated server providers and mid-size clouds: a real machine, in a real datacenter, with network and power handled, for a monthly fee.</p>
<p>You get root, predictable performance without noisy neighbors, and bandwidth allowances that are not priced as a profit center. You do not get managed databases, autoscaling, or a hundred adjacent services.</p>
<p>For a very large number of applications this is the correct answer and it is under-considered because the discourse is binary.</p>
<h2 id="the-operational-minimum">the operational minimum<a class="anchor" href="#the-operational-minimum" aria-label="link to this section">#</a></h2>
<p>If you self-host, these are non-negotiable:</p>
<ul><li><strong>Automated, tested restores.</strong> Not backups — <em>restores</em>. A backup you have never restored is a hypothesis. Test it quarterly, on a schedule, with a timer running.</li><li><strong>Unattended security updates</strong>, at least for the OS.</li><li><strong>Monitoring with alerting that reaches a human.</strong> Disk full is the most common self-hosted outage and it is entirely preventable.</li><li><strong><a class="xref" href="/infrastructure-as-code-ten-years-of-lessons/" title="Infrastructure as code, ten years of lessons">Infrastructure as code</a>.</strong> The machine must be reproducible. If rebuilding it requires someone's memory, you have a single point of failure that is a person.</li><li><strong>A documented runbook</strong> for the failures you expect: disk, certificate expiration, service crash, host failure.</li></ul>
<p>That is a weekend of setup and a few hours a month. If nobody on the team will own those hours, use the cloud — that is a legitimate reason and it is the actual deciding factor more often than cost is.</p>
<h2 id="the-thing-that-changed-my-mind">the thing that changed my mind<a class="anchor" href="#the-thing-that-changed-my-mind" aria-label="link to this section">#</a></h2>
<p>I used to treat "we run our own servers" as a red flag. I now treat "we are on the cloud and have never modeled the alternative" as an equal one.</p>
<p>Both are defaults applied without analysis. The analysis takes an afternoon and the answer is frequently not what the reflex says.</p>]]></content:encoded></item><item><title>The cost of a meeting, in engineering terms</title><link>https://readme.news/the-cost-of-a-meeting-in-engineering-terms/</link><guid isPermaLink="true">https://readme.news/the-cost-of-a-meeting-in-engineering-terms/</guid><pubDate>Wed, 11 Mar 2026 09:00:00 +0000</pubDate><description>Not the hourly rate. The fragmentation. Here&#x27;s why a 30-minute meeting costs four hours.</description><content:encoded><![CDATA[<p>The standard argument against meetings is arithmetic: eight people, thirty minutes, four person-hours, multiply by salary.</p>
<p>That understates it by a large factor, and the reason is the same reason context switching is expensive for a CPU.</p>
<h2 id="the-fragmentation-cost">the fragmentation cost<a class="anchor" href="#the-fragmentation-cost" aria-label="link to this section">#</a></h2>
<p>Deep engineering work requires holding a system in your head: the call graph, the invariants, the specific thing you were about to check, the four hypotheses you had narrowed to two.</p>
<p>Building that state takes time — call it twenty to forty minutes for nontrivial work. Losing it takes one interruption.</p>
<p>So a meeting at 2 p.m. does not cost thirty minutes. It costs:</p>
<ul><li>The twenty minutes before it, where you cannot start anything substantial because you will be interrupted.</li><li>The thirty minutes of the meeting.</li><li>The twenty to forty minutes after it, rebuilding the state you lost.</li></ul>
<p>That is roughly ninety minutes for a thirty-minute meeting, per person, and it is the optimistic case.</p>
<p>Worse: a meeting placed in the middle of a morning does not remove ninety minutes from a four-hour block. It removes the <em>block</em>. Two ninety-minute fragments are not equivalent to one three-hour stretch, because the hardest work requires depth that ninety minutes cannot reach.</p>
<h2 id="the-schedule-shapes-that-work">the schedule shapes that work<a class="anchor" href="#the-schedule-shapes-that-work" aria-label="link to this section">#</a></h2>
<p><strong>Meeting-free days.</strong> Two full days a week with no recurring meetings. Not "try to keep them clear" — a calendar policy. This is the single highest-impact scheduling change available and it is nearly free.</p>
<p><strong>Meetings at the edges.</strong> Cluster them at the start or end of the day. A day with three meetings from 9 to 11 and nothing after is dramatically more productive than one with three meetings at 10, 1, and 3.</p>
<p><strong>Default to 25 and 50 minutes.</strong> Not 30 and 60. The buffer prevents the cascade where every meeting starts late and the person coming from the previous one misses the first five minutes.</p>
<p><strong>Async by default for status.</strong> Standup is a meeting to exchange information that could be a written update. The argument for synchronous standup is that it surfaces blockers, which it does, and a written update with a "blocked on" field surfaces them too, in a searchable form, without costing everyone their morning.</p>
<h2 id="when-meetings-are-actually-right">when meetings are actually right<a class="anchor" href="#when-meetings-are-actually-right" aria-label="link to this section">#</a></h2>
<p>They are, frequently, and the anti-meeting position overcorrects.</p>
<p><strong>High-bandwidth disagreement.</strong> Two people who disagree about a design will resolve it in twenty minutes of conversation and in nine days of comment threads. Real-time is enormously better for anything with back-and-forth.</p>
<p><strong>Ambiguity resolution.</strong> When nobody is sure what the problem even is, a conversation converges much faster than writing, because you can ask the clarifying question immediately.</p>
<p><strong>Anything with emotional content.</strong> Feedback, conflict, bad news. Writing is the wrong medium and using it is usually avoidance.</p>
<p><strong>Building relationships.</strong> Real, unmeasurable, and the reason fully-async organizations struggle in ways they cannot diagnose.</p>
<p>The rule: <strong>meet for things that require interaction. Write for things that require information.</strong></p>
<h2 id="the-tests">the tests<a class="anchor" href="#the-tests" aria-label="link to this section">#</a></h2>
<p>Before scheduling, three questions:</p>
<p><strong>What decision will be made?</strong> If none, this is a status update. Write it.</p>
<p><strong>Who must be there for that decision?</strong> The answer is usually two to four people. Everyone else can read the outcome.</p>
<p><strong>What would happen if we did not have it?</strong> If the answer is "nothing," you have found a recurring meeting that outlived its purpose. Every organization has several.</p>
<h2 id="the-audit-worth-doing">the audit worth doing<a class="anchor" href="#the-audit-worth-doing" aria-label="link to this section">#</a></h2>
<p>Take a recurring meeting and cancel it for a month. Not "make it optional" — cancel it.</p>
<p>If someone notices and asks for it back, with a reason, reinstate it. Most of the time nobody notices, which tells you what you needed to know.</p>
<p>This works because recurring meetings are created for a reason that expires and never re-evaluated, and the social cost of proposing cancellation is high enough that nobody does it. A trial cancellation removes the social cost.</p>
<h2 id="the-part-managers-should-hear">the part managers should hear<a class="anchor" href="#the-part-managers-should-hear" aria-label="link to this section">#</a></h2>
<p>Your calendar is not your team's calendar. A manager's day is legitimately made of meetings — that is the job, and thirty-minute chunks are the natural unit.</p>
<p>An engineer's day is not, and scheduling as if it were is the single most common way well-meaning managers destroy their team's output while looking at metrics that do not show it.</p>
<p>Protect the blocks. It is most of what schedule-level management can do.</p>]]></content:encoded></item><item><title>The engineer's guide to saying no</title><link>https://readme.news/the-engineers-guide-to-saying-no/</link><guid isPermaLink="true">https://readme.news/the-engineers-guide-to-saying-no/</guid><pubDate>Sat, 28 Feb 2026 09:00:00 +0000</pubDate><description>Refusing work badly is a career problem. Refusing it well is one of the most valuable things a senior engineer does.</description><content:encoded><![CDATA[<p>Most engineers are bad at saying no. They either cannot do it — and end up with a commitment they cannot meet — or they do it in a way that reads as obstruction, and get routed around.</p>
<p>Both failures come from the same mistake: treating "no" as a verdict rather than as the opening of a conversation about trade-offs.</p>
<h2 id="what-the-request-actually-is">what the request actually is<a class="anchor" href="#what-the-request-actually-is" aria-label="link to this section">#</a></h2>
<p>When someone asks you to build something, they are not asking for the thing. They are asking for an outcome, and the thing is their guess at how to get it.</p>
<p>That means the highest-value response is frequently not yes or no. It is a better guess.</p>
<blockquote><p>"You are asking for a real-time <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>. What decision are you going to make from it? If it is 'should we page someone,' an alert is better than a dashboard and it is a day of work instead of a month."</p></blockquote>
<p>That is a no to the request and a yes to the outcome, and nobody experiences it as obstruction.</p>
<p>Ask what the outcome is before you evaluate the request. Half the time the request evaporates.</p>
<h2 id="the-four-honest-noes">the four honest noes<a class="anchor" href="#the-four-honest-noes" aria-label="link to this section">#</a></h2>
<p><strong>"Not this, that."</strong> The alternative approach that achieves the same outcome for less. This is the best one and it requires actually understanding the problem.</p>
<p><strong>"Yes, and here is what it displaces."</strong> Not a no. A statement about capacity, which is factual and is the other person's decision to make.</p>
<blockquote><p>"I can do that in this cycle. It means the API migration slips to next quarter. Which do you want?"</p></blockquote>
<p>This is enormously more effective than "we do not have time," because it hands the prioritization decision to the person whose job it is, with the information they need.</p>
<p><strong>"Yes, after X."</strong> A sequencing objection. "We can build this on top of the new data model. Building it on the old one means we build it twice."</p>
<p><strong>"No, and here is the risk I am flagging."</strong> The real no, reserved for correctness, security, legal, or ethical problems. Use it rarely so that it lands when you do.</p>
<blockquote><p>"I am not going to implement this the way it is specified because it stores plaintext credentials. I will implement it with a token exchange, which takes three extra days. If that is unacceptable, I want the decision documented and made by someone who can accept the risk."</p></blockquote>
<p>That is a hard sentence to say and it is the sentence the job sometimes requires.</p>
<h2 id="the-ones-that-do-not-work">the ones that do not work<a class="anchor" href="#the-ones-that-do-not-work" aria-label="link to this section">#</a></h2>
<p><strong>"That is not possible."</strong> Almost always false, and the person will find someone who says it is. Say "that would take six months" instead, which is the real constraint and is checkable.</p>
<p><strong>"That is a bad idea."</strong> Without an alternative, this is just friction.</p>
<p><strong>Silence.</strong> Not responding is a no that damages trust, because it looks like you did not care rather than that you disagreed.</p>
<p><strong>"Sure"</strong> followed by not doing it. The worst one. It destroys your reliability, which is the only currency you actually have.</p>
<h2 id="the-timing">the timing<a class="anchor" href="#the-timing" aria-label="link to this section">#</a></h2>
<p>Say no early. The cost of a no rises with every day of planning that assumed a yes.</p>
<p>An objection raised in the design review is a discussion. The same objection raised two weeks before launch is a crisis, and people will remember that you could have said it earlier — correctly.</p>
<p>If you have doubts, voice them while they are cheap.</p>
<h2 id="the-part-about-capital">the part about capital<a class="anchor" href="#the-part-about-capital" aria-label="link to this section">#</a></h2>
<p>Every no spends something. Every yes earns something. If you never say yes, your noes stop landing, because you have become the person who says no.</p>
<p>Deliver reliably on what you agree to. That is what makes the refusal credible when it matters. Engineers who are trusted to ship get enormous latitude to push back, and engineers who are not, do not — regardless of whether they are right.</p>
<h2 id="the-hardest-case">the hardest case<a class="anchor" href="#the-hardest-case" aria-label="link to this section">#</a></h2>
<p>Sometimes you are overruled on something you believe is wrong, and it is not a correctness or ethics issue — it is a judgment call and someone with the authority made a different one.</p>
<p>Disagree and commit is the right practice here and it is genuinely hard. Say your piece once, clearly, in writing. Then build the thing well.</p>
<p>Do not build it badly to prove a point. Do not relitigate it in every standup. Do not say "I told you so" if it goes wrong — the written record already said it, and gloating costs you the ability to be listened to next time.</p>
<p>You will be wrong about some of these. That is the actual reason to commit gracefully: you are not always right, and a culture where disagreement is followed by good-faith execution is one where being wrong is survivable for everyone, including you.</p>]]></content:encoded></item><item><title>Documentation is a product and you should staff it like one</title><link>https://readme.news/documentation-is-a-product-and-you-should-staff-it-like-one/</link><guid isPermaLink="true">https://readme.news/documentation-is-a-product-and-you-should-staff-it-like-one/</guid><pubDate>Fri, 20 Feb 2026 09:00:00 +0000</pubDate><description>Every team says docs matter. Almost none of them assign an owner, a budget, or a metric.</description><content:encoded><![CDATA[<p>Ask any engineering team whether documentation matters and they will say yes. Ask who owns it and you will get a pause.</p>
<p>That pause is the entire problem.</p>
<h2 id="the-four-kinds-and-why-mixing-them-fails">the four kinds, and why mixing them fails<a class="anchor" href="#the-four-kinds-and-why-mixing-them-fails" aria-label="link to this section">#</a></h2>
<p>The taxonomy that fixed documentation for me — and it is not mine, it is the Diátaxis framework — is that there are four distinct kinds and they have incompatible goals.</p>
<p><strong>Tutorials</strong> teach a beginner by having them do something that works. The goal is a successful experience, not completeness. A tutorial that mentions every option has failed. It should be prescriptive, opinionated, and it should work exactly as written, every time.</p>
<p><strong>How-to guides</strong> help someone accomplish a specific task they already understand. "How to configure TLS." Goal-oriented, assumes competence, skips explanation.</p>
<p><strong>Reference</strong> describes the machinery precisely and completely. Every parameter, every return value, every error. Boring by design. Generated where possible.</p>
<p><strong>Explanation</strong> provides understanding. Why is it designed this way? What are the trade-offs? What is the mental model? This is the kind that is almost always missing and the kind that most reduces support burden.</p>
<p>Most documentation fails because it tries to be all four at once. A tutorial that stops to explain architecture loses the beginner. A reference page with a narrative is hard to scan. Separate them, label them, and each one gets better.</p>
<h2 id="what-staff-it-like-a-product-means">what "staff it like a product" means<a class="anchor" href="#what-staff-it-like-a-product-means" aria-label="link to this section">#</a></h2>
<p><strong>One named owner.</strong> Not "the team." A person whose review includes it.</p>
<p><strong>A budget in the sprint.</strong> Documentation work sized and scheduled alongside features, not appended to the end of a ticket where it gets cut.</p>
<p><strong>Metrics.</strong> Support tickets that a doc would have prevented. Search queries with no results. Time-to-first-successful-request for a new user. Page-level feedback. Every one of these is measurable and almost nobody measures them.</p>
<p><strong>A definition of done that includes it.</strong> A feature is not shipped until it is documented. This is either enforced or it is a slogan; there is no middle.</p>
<h2 id="the-practices-that-actually-move-the-needle">the practices that actually move the needle<a class="anchor" href="#the-practices-that-actually-move-the-needle" aria-label="link to this section">#</a></h2>
<p><strong>Docs live with the code.</strong> Same repository, same pull request, same review. Documentation in a separate wiki drifts within one quarter, guaranteed, without exception.</p>
<p><strong>Test the examples.</strong> Every code sample in your docs should be extracted and run in CI. Broken examples are worse than no examples — they destroy trust in the whole document, and every set of docs has them, because they were correct when written.</p>
<p><strong>Write the failure cases.</strong> The single highest-value section in any documentation is "common errors and what they mean." This is what <a class="xref" href="/stack-overflows-traffic-fell-off-a-cliff-and-it-is-not-coming-back/" title="Stack Overflow&#x27;s traffic fell off a cliff and it is not coming back">Stack Overflow</a> existed to provide and it is the thing your docs almost certainly lack.</p>
<p>Go read your support queue. Every recurring question is a documentation gap with a measured frequency attached.</p>
<p><strong>Date and version everything.</strong> "This page describes v4.2, last updated 2026-01-15." Undated documentation is untrustworthy documentation, because the reader cannot tell whether it is current.</p>
<p><strong>Make the first example work.</strong> The single most common documentation failure: the quickstart does not run. Someone changed a default, renamed a parameter, required a new config field. Test the quickstart in CI, on a clean environment, on every release.</p>
<h2 id="the-argument-that-gets-budget">the argument that gets budget<a class="anchor" href="#the-argument-that-gets-budget" aria-label="link to this section">#</a></h2>
<p>Documentation is deflection. Every question answered by a doc is a question not asked of an engineer.</p>
<p>Count your support load. Estimate the fraction that is documentation-shaped — "how do I," "what does this error mean," "does it support." In most organizations it is more than half.</p>
<p>Now price that in engineer-hours. That is your documentation ROI, and it is usually large enough to fund a technical writer, which is the actual right answer and which almost nobody does.</p>
<h2 id="the-new-reason-it-matters">the new reason it matters<a class="anchor" href="#the-new-reason-it-matters" aria-label="link to this section">#</a></h2>
<p>Your documentation is now also a model's training data and a model's retrieval corpus.</p>
<p>When a developer asks an assistant about your library, the answer is synthesized from your docs. If your docs are wrong, incomplete, or ambiguous, the assistant confidently produces wrong code, and the user blames your library.</p>
<p>You have less control over how your project is explained than you did three years ago, and the only lever you have is the quality of the source material.</p>
<p>That is a strange new incentive and it is the strongest argument for good documentation that has ever existed.</p>]]></content:encoded></item><item><title>Every company is briefly a model company</title><link>https://readme.news/every-company-is-briefly-a-model-company/</link><guid isPermaLink="true">https://readme.news/every-company-is-briefly-a-model-company/</guid><pubDate>Wed, 18 Feb 2026 09:00:00 +0000</pubDate><description>The fine-tuning wave, the RAG wave, and the agent wave all followed the same arc. Here&#x27;s where the value actually settled.</description><content:encoded><![CDATA[<p>Three times in three years, a wave of companies concluded that the way to build an AI product was to own a layer that turned out not to be theirs.</p>
<p>The pattern is consistent enough to be predictive, which makes it worth naming.</p>
<h2 id="wave-one-fine-tuning">wave one: fine-tuning<a class="anchor" href="#wave-one-fine-tuning" aria-label="link to this section">#</a></h2>
<p><strong>The thesis (2023):</strong> general models are generic. Fine-tune on your domain data and you get a model that is specifically good at your problem and that competitors cannot replicate.</p>
<p><strong>What happened:</strong> base models improved faster than fine-tunes could keep up. A fine-tuned model from six months ago was worse than the new base model with a good prompt. Every fine-tune had to be redone on every model release, which is a treadmill.</p>
<p><strong>Where it settled:</strong> fine-tuning is genuinely valuable for narrow, stable, high-volume tasks — classification into your specific taxonomy, output in your specific format, a <a class="xref" href="/haiku-45-and-the-collapsing-cost-of-good-enough/" title="Haiku 4.5 and the collapsing cost of good-enough">small model</a> matching a large model's behavior on one task. It is not a moat and it is not a product strategy.</p>
<h2 id="wave-two-rag">wave two: RAG<a class="anchor" href="#wave-two-rag" aria-label="link to this section">#</a></h2>
<p><strong>The thesis (2023-24):</strong> the model does not know your data. Build a retrieval pipeline — chunk, embed, index, retrieve, rerank — and you have a defensible system built on proprietary knowledge.</p>
<p><strong>What happened:</strong> context windows grew by two orders of magnitude, long-context quality improved, and prompt <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> made large contexts economical. A large fraction of naive RAG got replaced by putting the documents in the prompt.</p>
<p>Simultaneously, the pipeline components commoditized. Embedding models became interchangeable. Vector search became a feature of every database rather than a product.</p>
<p><strong>Where it settled:</strong> retrieval did not go away. It moved. Retrieval is now a <em>tool the model calls</em> rather than a preprocessing step, and it is genuinely necessary at large corpus sizes, where citation is required, and where cost or latency rules out large contexts.</p>
<p>The infrastructure repositioned rather than dying, which is what usually happens.</p>
<h2 id="wave-three-agent-frameworks">wave three: agent frameworks<a class="anchor" href="#wave-three-agent-frameworks" aria-label="link to this section">#</a></h2>
<p><strong>The thesis (2024-25):</strong> models cannot plan reliably. Build the orchestration — task decomposition, tool routing, retry logic, state management — and own the layer that makes agents work.</p>
<p><strong>What happened:</strong> the models absorbed it. In-context tool use during reasoning removed the need for an external loop. Native parallel tool calling removed the need for a dispatcher. Memory tools and <a class="xref" href="/claude-sonnet-45-and-the-agent-that-runs-for-thirty-hours/" title="Claude Sonnet 4.5 and the agent that runs for thirty hours">context editing</a> removed the need for external state management.</p>
<p>Each capability the framework provided became a model feature within about a year.</p>
<p><strong>Where it settled:</strong> in progress, but the shape is clear. Orchestration frameworks are converging on thin conveniences. The durable parts are evaluation, observability, and the domain-specific policy that no model will ever have.</p>
<h2 id="the-pattern">the pattern<a class="anchor" href="#the-pattern" aria-label="link to this section">#</a></h2>
<p>Every wave follows the same arc:</p>
<ol><li>The model has a limitation.</li><li>Companies build infrastructure to work around the limitation.</li><li>The limitation gets fixed in the model.</li><li>The infrastructure either finds a new position or disappears.</li></ol>
<p>The consistent error is <strong>building on a gap rather than on an asset</strong>. A gap is temporary by construction — the labs are actively working to close it, with more resources than you have.</p>
<h2 id="what-has-actually-been-durable">what has actually been durable<a class="anchor" href="#what-has-actually-been-durable" aria-label="link to this section">#</a></h2>
<p>Across all three waves, the same things kept their value:</p>
<p><strong>Proprietary data.</strong> Not "we have documents" — everyone has documents. Data that is genuinely yours: production logs, customer interactions, labeled outcomes, domain expertise encoded as examples. Nobody can buy it and no model was trained on it.</p>
<p><strong>Evaluation specific to your task.</strong> The company that knows, precisely, how well a system performs on their actual problem can adopt a new model in a day. The one that does not spends a month on vibes. That gap compounds every release cycle.</p>
<p><strong>Distribution and workflow integration.</strong> Being where the user already works. The most boring answer and the most durable.</p>
<p><strong>Domain constraints.</strong> The rules, regulations, edge cases, and institutional knowledge that make a generic capability into a usable product. This is unglamorous and it is the actual work.</p>
<p><strong>Trust.</strong> Security posture, compliance, reliability, support. Enterprises buy this and it takes years to build.</p>
<h2 id="the-test-to-apply">the test to apply<a class="anchor" href="#the-test-to-apply" aria-label="link to this section">#</a></h2>
<p>Before building on top of a model limitation, ask: <strong>if this limitation disappeared next quarter, what would I have left?</strong></p>
<p>If the answer is "nothing," you are building a bridge over a river that is being drained.</p>
<p>If the answer is "the data, the evaluations, the integrations, and the customer relationships," build it, and expect to throw the bridge away.</p>]]></content:encoded></item><item><title>What happened to microservices</title><link>https://readme.news/what-happened-to-microservices/</link><guid isPermaLink="true">https://readme.news/what-happened-to-microservices/</guid><pubDate>Wed, 11 Feb 2026 09:00:00 +0000</pubDate><description>The pendulum swung back and the useful part is the specific reasons, not the vibe.</description><content:encoded><![CDATA[<p>Around 2015, splitting your application into many small services was the default recommendation. Around 2022, "modular monolith" started appearing in the same conference slots. Now the consensus is roughly "start with a monolith, extract when you have a reason."</p>
<p>The pendulum framing is unhelpful. What is useful is naming the specific reasons the original argument was wrong, because those reasons generalize.</p>
<h2 id="what-the-case-for-microservices-actually-was">what the case for microservices actually was<a class="anchor" href="#what-the-case-for-microservices-actually-was" aria-label="link to this section">#</a></h2>
<p>Four claims, all plausible:</p>
<ol><li><strong>Independent deployment.</strong> Teams ship without coordinating.</li><li><strong>Independent scaling.</strong> Scale the busy part, not everything.</li><li><strong>Technology freedom.</strong> Each service picks its own stack.</li><li><strong>Fault isolation.</strong> One service's failure does not take down the rest.</li></ol>
<h2 id="how-each-one-went">how each one went<a class="anchor" href="#how-each-one-went" aria-label="link to this section">#</a></h2>
<p><strong>Independent deployment: real, and the main benefit.</strong> This one held up. If two teams can deploy without coordinating, that is genuinely valuable and it scales with organization size.</p>
<p>The catch: it requires the services to be genuinely independent. If service A cannot deploy without service B deploying a compatible change first, you have a distributed monolith — all the operational cost, none of the benefit. That is the outcome most organizations actually got, because splitting a system along the wrong seams produces services that must change together.</p>
<p><strong>Independent scaling: real and usually irrelevant.</strong> Most applications are not scale-constrained in a way where this matters. A monolith on a bigger machine handles a very large amount of traffic, and modern machines are enormous.</p>
<p>Where it matters — one component with wildly different resource requirements, like a video transcoder or an ML inference path — extract that one thing. That is not microservices, that is a service.</p>
<p><strong>Technology freedom: real and a mistake.</strong> Six languages means six toolchains, six sets of libraries, six deployment stories, six <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> rotations that cannot help each other, and an engineer who cannot move between teams without a month of ramp-up.</p>
<p>Most organizations that got this eventually mandated standardization anyway, which means they paid the cost and gave up the benefit.</p>
<p><strong>Fault isolation: inverted.</strong> This is the one that went most wrong.</p>
<p>In a monolith, a failing component throws an exception. In a distributed system, a failing component times out — after 30 seconds, while holding a connection, in every caller, simultaneously. Cascading failure, retry amplification, and metastable states are <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> that a monolith simply does not have.</p>
<p>You do not get fault isolation for free by adding a network. You get it by implementing circuit breakers, <a class="xref" href="/timeouts-every-one-of-them/" title="Timeouts: every one of them">timeouts</a>, bulkheads, and <a class="xref" href="/backpressure-is-the-concept-your-system-is-missing/" title="Backpressure is the concept your system is missing">backpressure</a> — deliberate engineering that most teams did not do because they thought the architecture gave it to them.</p>
<h2 id="the-costs-nobody-budgeted">the costs nobody budgeted<a class="anchor" href="#the-costs-nobody-budgeted" aria-label="link to this section">#</a></h2>
<p><strong>Debugging.</strong> A single user action touches nine services. Understanding what happened requires <a class="xref" href="/distributed-tracing-that-people-actually-use/" title="Distributed tracing that people actually use">distributed tracing</a>, correlated logs, and a mental model of the whole call graph. This is dramatically harder than a stack trace and it is a permanent tax on every incident.</p>
<p><strong>Data consistency.</strong> A transaction across services is not a transaction. You get sagas, compensating actions, eventual consistency, and a class of bug where the system is in a state that should be impossible.</p>
<p><strong>Local development.</strong> Running the system on a laptop is a project. Teams end up developing against shared environments, which serializes work and reintroduces the coordination cost that microservices were supposed to remove.</p>
<p><strong>Latency.</strong> In-process call: nanoseconds. Network call: milliseconds. A chain of nine is a user-visible delay that did not exist before.</p>
<p><strong>Operational surface.</strong> Nine services means nine deployment pipelines, nine sets of dashboards, nine on-call runbooks, nine dependency trees to patch.</p>
<h2 id="the-synthesis">the synthesis<a class="anchor" href="#the-synthesis" aria-label="link to this section">#</a></h2>
<p><strong>Start with a monolith.</strong> One deployable, one database, clear internal module boundaries with enforced dependency rules.</p>
<p><strong>Enforce the boundaries in code.</strong> This is the part people skip and it is the part that matters. Modules with explicit public interfaces, dependency rules that fail the build, and no reaching into another module's tables. If you do this, you have most of the benefit and none of the network.</p>
<p><strong>Extract a service when you have a specific reason.</strong> Good reasons: a genuinely different scaling profile, a team that needs independent release cadence and can have it, a compliance boundary, a component that must be written in a different language.</p>
<p>Bad reasons: it feels cleaner, we read a blog post, our architecture diagram would look better.</p>
<p><strong>The extraction should be easy</strong> if the boundaries were real. If pulling a module out is a six-month project, that tells you the boundary was not real, and it also tells you extracting it would not have given you independence.</p>
<h2 id="the-general-lesson">the general lesson<a class="anchor" href="#the-general-lesson" aria-label="link to this section">#</a></h2>
<p>Every architectural decision trades one kind of complexity for another. Microservices trade <em>code complexity</em> for <em>operational complexity</em>.</p>
<p>For a large organization with many teams, that is frequently a good trade — operational complexity can be handled by a <a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">platform team</a>, and coordination complexity cannot be handled by anyone.</p>
<p>For a team of twelve, it is a terrible trade, and the industry spent five years recommending it to teams of twelve.</p>]]></content:encoded></item><item><title>Go and the case for boring</title><link>https://readme.news/go-and-the-case-for-boring/</link><guid isPermaLink="true">https://readme.news/go-and-the-case-for-boring/</guid><pubDate>Wed, 04 Feb 2026 09:00:00 +0000</pubDate><description>A language that ships small releases, breaks nothing, and refuses features. Its detractors and its users are describing the same thing.</description><content:encoded><![CDATA[<p>Go's release notes are the least exciting in the industry, and its adoption in infrastructure software is close to total. Those two facts are the same fact.</p>
<h2 id="the-design-decision-that-defines-it">the design decision that defines it<a class="anchor" href="#the-design-decision-that-defines-it" aria-label="link to this section">#</a></h2>
<p>Go's compatibility promise is stronger than almost any language's: code written against Go 1 continues to compile and run. That has held since 2012.</p>
<p>Combined with a deliberate resistance to adding features, the effect is that Go code from 2015 reads like Go code from today. There is one way to do most things. There is no dialect. There is no "modern Go" versus "legacy Go."</p>
<p>For a language used to write infrastructure that must be maintained for a decade by rotating teams, that property is worth more than any individual feature.</p>
<h2 id="what-the-criticism-gets-right">what the criticism gets right<a class="anchor" href="#what-the-criticism-gets-right" aria-label="link to this section">#</a></h2>
<p><strong>Error handling is verbose.</strong> <code>if err != nil</code> is a large share of the lines in a lot of Go code. The proposals to fix it have all been rejected, repeatedly, and the reasoning — that explicit error handling at every call site is a feature, not a bug — is coherent and is genuinely annoying to live with.</p>
<p><strong>Generics arrived late and are limited.</strong> No sum types. No method type parameters. The type inference gives up in cases where you expect it to work. Generics in Go are useful and they are clearly a retrofit.</p>
<p><strong>The nil <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> trap.</strong> A nil pointer stored in an interface produces a non-nil interface. This bites everyone once and it is a genuine wart.</p>
<p><strong>No sum types.</strong> The single most requested feature. Modeling "this is one of three things" in Go requires an interface with unexported methods and a type switch, which is neither exhaustive-checked nor pleasant.</p>
<h2 id="what-the-criticism-gets-wrong">what the criticism gets wrong<a class="anchor" href="#what-the-criticism-gets-wrong" aria-label="link to this section">#</a></h2>
<p>The complaints are almost all about <em>expressiveness</em>, and expressiveness is not the thing Go optimizes for. It optimizes for a large team maintaining a large codebase over a long time, with new people joining constantly.</p>
<p>Under that objective function:</p>
<ul><li><strong>Verbosity is fine</strong> if it makes the code obvious without context.</li><li><strong>Few features is good</strong>, because every feature is a dialect someone will write in and someone else will have to read.</li><li><strong>Fast compilation</strong> matters enormously, because it determines the <a class="xref" href="/why-your-tests-are-slow/" title="Why your tests are slow">feedback loop</a> and therefore how people work.</li><li><strong>A good standard library</strong> means fewer dependencies, which means fewer supply chain problems and fewer upgrade treadmills.</li><li><strong><code>gofmt</code></strong> ended formatting arguments permanently, which saved the industry an incalculable number of hours.</li></ul>
<p>Go is not trying to be a good language for writing clever code. It is trying to be a good language for reading code someone else wrote three years ago while you are on call at 3 a.m.</p>
<p>It is very good at that.</p>
<h2 id="the-concurrency-story-honestly">the concurrency story, honestly<a class="anchor" href="#the-concurrency-story-honestly" aria-label="link to this section">#</a></h2>
<p>Goroutines and channels were revolutionary in 2012 and are now table stakes. Several languages have equivalent or better models.</p>
<p>What has held up: the scheduler is excellent, the tooling around concurrency — the race detector especially — is best in class, and <code>context</code> for cancellation propagation is a good design that most ecosystems lack an equivalent of.</p>
<p>What has not: channels are overused by people who learned Go from the tour. Most Go code should use a mutex and a plain function call. "Do not communicate by sharing memory; share memory by communicating" is good advice that has produced a lot of unnecessarily channel-based code.</p>
<h2 id="where-it-wins">where it wins<a class="anchor" href="#where-it-wins" aria-label="link to this section">#</a></h2>
<p>Network services, CLI tools, infrastructure daemons, anything deployed as a single static binary, anything with a large team.</p>
<p>The ecosystem effect is real: a very large fraction of the cloud-native stack is Go. If you work in infrastructure, you will read Go whether or not you write it.</p>
<h2 id="where-it-does-not">where it does not<a class="anchor" href="#where-it-does-not" aria-label="link to this section">#</a></h2>
<p>Anything needing precise memory control. Anything numeric or scientific, where the ecosystem is thin and the language does not help. Anything where expressiveness genuinely pays — a compiler, a complex domain model, anything with rich algebraic structure.</p>
<h2 id="the-thing-worth-stealing">the thing worth stealing<a class="anchor" href="#the-thing-worth-stealing" aria-label="link to this section">#</a></h2>
<p>Whatever language you use: Go's compatibility promise, its formatting standardization, and its resistance to feature accumulation are choices, not accidents.</p>
<p>Most projects would benefit from adopting the spirit of all three. Pick one way to do things. Automate the formatting argument out of existence. Say no to the feature that adds a second way to express something that already has one.</p>
<p>That is available to you regardless of your language, and it is most of what makes Go codebases pleasant.</p>]]></content:encoded></item><item><title>Rust's eleventh year and the shape of a finished language</title><link>https://readme.news/rusts-eleventh-year-and-the-shape-of-a-finished-language/</link><guid isPermaLink="true">https://readme.news/rusts-eleventh-year-and-the-shape-of-a-finished-language/</guid><pubDate>Tue, 20 Jan 2026 09:00:00 +0000</pubDate><description>Six-week releases, no breakage, and a backlog that is finally shorter than it was. What&#x27;s left is the hard part.</description><content:encoded><![CDATA[<p>Another <a class="xref" href="/rusts-six-week-metronome-seven-years-on/" title="Rust&#x27;s six-week metronome, seven years on">Rust release</a>, another set of const stabilizations and library additions. The release notes have been pleasantly dull for a year, which is the best thing you can say about a language with load-bearing production deployments.</p>
<p>Rather than enumerate, it is worth asking what "finished" would look like and how close this is.</p>
<h2 id="what-got-finished">what got finished<a class="anchor" href="#what-got-finished" aria-label="link to this section">#</a></h2>
<p>The list of "this obviously should work and does not" items has genuinely shrunk. Trait upcasting, let chains, RPIT capture, anonymous pipes, const in more places, SIMD intrinsics, async closures. Each of those was a multi-year open question and each is now just how the language works.</p>
<p>That is the signature of a language moving from the growth phase into the maintenance phase, and it is a good place to be. The compiler is faster than it was. The <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a> remain the best in the industry. The tooling story — cargo, clippy, rustfmt, rust-analyzer — is coherent in a way that most languages never achieve.</p>
<h2 id="what-is-left-honestly">what is left, honestly<a class="anchor" href="#what-is-left-honestly" aria-label="link to this section">#</a></h2>
<p><strong>Async.</strong> Still the sharpest edge. Async traits work. Async closures work. What does not work smoothly: cancellation semantics that do not surprise you, async drop, <code>Send</code> bound propagation through generic async code, and the absence of a standard executor <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> which fragments the ecosystem.</p>
<p>The <code>Pin</code> API remains the piece of Rust that experienced engineers most often describe as genuinely confusing, and the ergonomic improvements have been incremental.</p>
<p>This is the area where Rust is meaningfully harder than it needs to be, and where a new user is most likely to conclude the language is too complicated. It matters because async is how you write network services, and network services are most of software.</p>
<p><strong><a class="xref" href="/rust-189-and-the-long-tail-of-const-generics/" title="Rust 1.89 and the long tail of const generics">Const generic</a> expressions.</strong> <code>[T; N * 2]</code> still needs nightly. The blocker is real — deciding when two type-level expressions are equal is undecidable in general and the team will not ship a rule they cannot explain.</p>
<p><strong>GUI.</strong> Ten-plus years, several serious projects, no consensus. Partly a Rust problem: the borrow checker's relationship with the retained-mode widget tree that every traditional GUI framework uses is genuinely awkward. Partly not a Rust problem: <a class="xref" href="/cross-platform-is-a-promise-you-make-to-your-budget/" title="Cross-platform is a promise you make to your budget">cross-platform</a> GUI is unsolved everywhere.</p>
<p><strong>Compile times.</strong> Better every year. Still slower than Go. Structurally will remain so, because monomorphization and the optimization Rust performs are not free and are the reason the output is fast.</p>
<h2 id="the-adoption-picture-unsentimentally">the adoption picture, unsentimentally<a class="anchor" href="#the-adoption-picture-unsentimentally" aria-label="link to this section">#</a></h2>
<p><strong>Won decisively:</strong> systems programming, CLI tooling, JavaScript build tooling, cryptography implementations, embedded, <a class="xref" href="/the-component-model-and-the-plugin-problem/" title="The component model and the plugin problem">WebAssembly</a> targets, kernel drivers in both Linux and Windows.</p>
<p><strong>Winning slowly:</strong> backend services, where Go's simplicity and compile speed are a real competitor and where the async ergonomics cost is felt daily.</p>
<p><strong>Not winning and probably will not:</strong> application development, data science, scripting. Different constraints, different right answers.</p>
<p>That is a good outcome. A language does not need to win everything, and the places Rust won are the places where memory safety bugs were causing the most damage.</p>
<h2 id="the-thing-to-appreciate">the thing to appreciate<a class="anchor" href="#the-thing-to-appreciate" aria-label="link to this section">#</a></h2>
<p>Eleven years of six-week releases and no broken stable program. Three editions with automated migration and no ecosystem split. A governance structure that has survived real disagreements in public.</p>
<p>For a language that has changed as much as Rust has, that stability record is the actual achievement, and it is the reason people are willing to put infrastructure on it. Features are cheap. Trust is not.</p>
<h2 id="for-someone-deciding-whether-to-learn-it-in-2026">for someone deciding whether to learn it in 2026<a class="anchor" href="#for-someone-deciding-whether-to-learn-it-in-2026" aria-label="link to this section">#</a></h2>
<p>If you write systems software, networking infrastructure, embedded code, or anything where a memory safety bug is a security incident: yes, and you probably already know that.</p>
<p>If you write web applications: the async ergonomics are the thing you will fight, and Go or a good typed scripting language will make you more productive faster. Learn Rust anyway, eventually, because understanding ownership makes you better in every language.</p>
<p>If you tried it in 2021 and bounced off the compile times or the error messages: try again. Both are substantially different now.</p>]]></content:encoded></item>
</channel>
</rss>
