<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — ai</title>
<link>https://readme.news/tags/ai/</link>
<atom:link href="https://readme.news/tags/ai/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged ai.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>The AI Act's high-risk obligations take effect today</title><link>https://readme.news/the-ai-acts-high-risk-obligations-take-effect-today/</link><guid isPermaLink="true">https://readme.news/the-ai-acts-high-risk-obligations-take-effect-today/</guid><pubDate>Sun, 02 Aug 2026 09:00:00 +0000</pubDate><description>Two years after entry into force, the substantive requirements arrive. What changes, what was delayed, and what to do now.</description><content:encoded><![CDATA[<p>The EU AI Act's obligations for high-risk AI systems apply from today. This has been scheduled since the Act entered into force in August 2024, and it is the point at which the most substantive requirements become enforceable.</p>
<h2 id="what-applies-now">what applies now<a class="anchor" href="#what-applies-now" aria-label="link to this section">#</a></h2>
<p>For providers of high-risk systems — those used in employment, education access, credit, essential services, law enforcement, migration, justice, and as safety components in regulated products:</p>
<ul><li><strong>Risk management system</strong>, documented and maintained across the lifecycle.</li><li><strong>Data governance</strong>, with documented training, validation, and test data and attention to bias.</li><li><strong>Technical documentation</strong> sufficient for conformity assessment.</li><li><strong>Automatic logging</strong> of the system's operation, retained.</li><li><strong>Transparency</strong> to deployers about capabilities, limitations, and intended use.</li><li><strong>Human oversight</strong> designed into the system.</li><li><strong>Accuracy, robustness, and cybersecurity</strong> proportionate to the purpose.</li><li><strong>Conformity assessment</strong> and CE marking before placing on the market.</li><li><strong>Registration</strong> in the EU database.</li><li><strong>Post-market monitoring</strong> and serious incident reporting.</li></ul>
<p>Deployers — organizations using a high-risk system rather than supplying it — have a lighter but real set: use it according to instructions, ensure input data is relevant, monitor operation, retain logs, and assign competent human oversight.</p>
<h2 id="the-parts-that-moved">the parts that moved<a class="anchor" href="#the-parts-that-moved" aria-label="link to this section">#</a></h2>
<p>The implementation timeline has been adjusted more than once during the run-up, which is normal for regulation of this scope and which made planning genuinely difficult for anyone trying to comply.</p>
<p>The important practical consequences:</p>
<p><strong>Some obligations phase in later</strong> for systems already on the market, and for high-risk systems that are safety components of products covered by other EU legislation.</p>
<p><strong>Harmonised standards are still being finalised.</strong> Conformity assessment is easier when you can demonstrate compliance against a standard; where standards are not yet available, providers must demonstrate compliance against the requirements directly, which is more work and more uncertain.</p>
<p><strong>Enforcement capacity varies by member state.</strong> Market surveillance authorities are at different stages of readiness. That is not a reason to assume non-enforcement — it is a reason to expect inconsistency in the first period.</p>
<p>Check the current state before making decisions. The details have moved and may move again.</p>
<h2 id="what-to-do-if-you-are-in-scope">what to do if you are in scope<a class="anchor" href="#what-to-do-if-you-are-in-scope" aria-label="link to this section">#</a></h2>
<p><strong>If you have not classified your systems, do that today.</strong> It is the prerequisite for everything else and it is a legal question with technical inputs. A number of organizations discover they are in scope for one feature they had not considered — a CV screening tool, a proctoring feature, a creditworthiness proxy.</p>
<p><strong>Documentation is the bulk of the work.</strong> Training data provenance, validation methodology, known <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a>, intended use and misuse. Most engineering teams do not have this written down and reconstructing it is slower than producing it as you go.</p>
<p><strong>Logging is a technical requirement.</strong> Automatic recording of operation, retained, with enough detail to trace a decision. If your system does not do this, it is engineering work with a deadline that has passed.</p>
<p><strong>Human oversight is a design constraint, not a checkbox.</strong> "A person reviews the output" is insufficient if the person cannot meaningfully understand, question, or override it. Surfacing confidence, explaining the inputs that drove a decision, and making override easy and recorded are <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> requirements.</p>
<h2 id="if-you-are-not-in-scope">if you are not in scope<a class="anchor" href="#if-you-are-not-in-scope" aria-label="link to this section">#</a></h2>
<p>Most software is not. The transparency obligations — disclose that users are interacting with an AI system, label synthetic media — apply much more broadly and are much lighter.</p>
<p>The thing worth doing regardless of scope: <strong>know which category you are in, and write down why.</strong> "We assessed this and concluded it is not high-risk because X" is a one-page document that saves an enormous amount of time when someone asks, and someone will ask.</p>
<h2 id="the-honest-assessment">the honest assessment<a class="anchor" href="#the-honest-assessment" aria-label="link to this section">#</a></h2>
<p>The Act is the most comprehensive AI regulation any jurisdiction has attempted. It has real criticisms — compliance cost falls hardest on small companies, the risk categories map imperfectly onto how systems are actually built, and the standards process has lagged the deadlines.</p>
<p>It is also law, it has extraterritorial reach, and penalties scale with global turnover.</p>
<p>The organizations that started classification work a year ago are in reasonable shape today. The ones that waited for clarity are discovering that regulatory clarity tends to arrive after the deadline rather than before it, which is a lesson that generalizes well beyond this particular Act.</p>]]></content:encoded></item><item><title>Two years of agentic coding: what stuck</title><link>https://readme.news/two-years-of-agentic-coding-what-stuck/</link><guid isPermaLink="true">https://readme.news/two-years-of-agentic-coding-what-stuck/</guid><pubDate>Fri, 31 Jul 2026 09:00:00 +0000</pubDate><description>The workflows that survived contact with real work, the ones that did not, and what the whole thing actually changed.</description><content:encoded><![CDATA[<p>Terminal coding agents went from <a class="xref" href="/operator-and-the-long-road-to-an-agent-that-can-click/" title="Operator, and the long road to an agent that can click">research preview</a> to standard tooling in about two years. Enough time has passed to separate what stuck from what was a phase.</p>
<h2 id="what-stuck">what stuck<a class="anchor" href="#what-stuck" aria-label="link to this section">#</a></h2>
<p><strong>Mechanical refactors at scale.</strong> The clearest win, by a wide margin. Renaming a concept across four hundred files, migrating a deprecated API, converting a pattern used everywhere. Verifiable, tedious, and exactly what the tools are good at.</p>
<p>The important second-order effect: <strong>refactors that were too expensive to do now happen.</strong> A codebase where cross-cutting cleanup is affordable is a meaningfully better codebase, and that is a permanent improvement rather than a productivity number.</p>
<p><strong>Working in unfamiliar territory.</strong> A language you do not know, a framework you have not used, an API you have never touched. The median output in an unfamiliar domain is better than your first attempt, and reading it teaches you the idioms.</p>
<p>This is the use case I would defend most strongly and it is discussed least.</p>
<p><strong>Test generation from a specification.</strong> Not "write tests for this function" — that produces tests that assert the implementation. But "here is the behavior, write tests that verify it" works well and it inverts the effort in the right direction.</p>
<p><strong>Investigation.</strong> Reading logs, bisecting history, tracing a call path, summarizing a large diff. Parallelizable, cheap, and it saves the expensive resource, which is your attention.</p>
<p><strong>Repository-level instruction files.</strong> <code>AGENTS.md</code> and its equivalents became standard practice, and the discipline of writing down how your project actually works improved documentation for humans as a side effect.</p>
<h2 id="what-did-not-stick">what did not stick<a class="anchor" href="#what-did-not-stick" aria-label="link to this section">#</a></h2>
<p><strong>Fully autonomous feature development.</strong> The demo works. The real version produces a plausible implementation of a subtly different feature, because the requirements that live in someone's head were never written down.</p>
<p><strong>Agent fleets at high concurrency.</strong> The generation scales; the review does not. Two to three concurrent agents with one reviewer turned out to be the practical limit, and the constraint is entirely on the human side.</p>
<p><strong>Orchestration frameworks.</strong> Absorbed into the models, as function-calling libraries and JSON-repair libraries were before them. The durable layer was never orchestration.</p>
<p><strong>"Just describe it and it builds."</strong> For anything with design decisions, the description that is precise enough to produce the right result is approximately as long as the code, and writing it is the same work.</p>
<h2 id="what-actually-changed-about-the-job">what actually changed about the job<a class="anchor" href="#what-actually-changed-about-the-job" aria-label="link to this section">#</a></h2>
<p><strong>Review is the bottleneck, permanently.</strong> Generation got roughly two orders of magnitude cheaper. Verification got no cheaper at all. Everything downstream follows from that asymmetry and nothing in two years has changed it.</p>
<p><strong>Tests became the primary artifact.</strong> If the implementation is cheap and verification is expensive, effort moves to specification. The teams getting the most out of these tools are the ones with strong test suites, and the correlation is not subtle.</p>
<p><strong><a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">Type systems</a> got a promotion.</strong> Every constraint the compiler checks is verification you do not perform by reading. Teams that were ambivalent about strict typing became evangelists, and the reason is always the same: it catches the class of error generated code produces most.</p>
<p><strong>Small diffs became non-negotiable.</strong> A machine can produce two thousand lines effortlessly. Accepting it is not a favor to anyone.</p>
<p><strong>The skill that separates people is judgment, not speed.</strong> It always was. It is now the only thing, and the gap between engineers who can tell when output is wrong and engineers who cannot is much more visible than it was.</p>
<h2 id="the-thing-still-unresolved">the thing still unresolved<a class="anchor" href="#the-thing-still-unresolved" aria-label="link to this section">#</a></h2>
<p>The apprenticeship problem.</p>
<p>The judgment that makes a senior engineer valuable was acquired by writing a lot of code badly and then debugging it. That work is being automated. Nobody has a replacement for how the next generation acquires it, and junior hiring contracted sharply during exactly the period when the training mechanism was being removed.</p>
<p>I have written about this several times and I still do not have an answer beyond: deliberately do the hard part yourself sometimes, review generated code carefully as a learning exercise, and hire juniors anyway.</p>
<p>That is a partial answer to a structural problem and I am not satisfied with it.</p>
<h2 id="the-honest-summary-two-years-in">the honest summary, two years in<a class="anchor" href="#the-honest-summary-two-years-in" aria-label="link to this section">#</a></h2>
<p>These tools are genuinely useful and the useful envelope is narrower and more specific than either the enthusiasts or the skeptics claimed.</p>
<p>They are excellent at bounded, verifiable, tedious work. They are unreliable at anything requiring judgment about what should be built. They multiply output and do not multiply throughput, because throughput is limited by review.</p>
<p>The engineers getting the most from them are the ones who were already good at specifying problems precisely and at telling when something is wrong. That is not a new skill and it was never evenly distributed.</p>
<p>Which is roughly what every previous tooling revolution did: raised the floor, moved the bottleneck, and made expertise more valuable rather than less.</p>]]></content:encoded></item><item><title>The crawler tolls, one year on</title><link>https://readme.news/the-crawler-tolls-one-year-on/</link><guid isPermaLink="true">https://readme.news/the-crawler-tolls-one-year-on/</guid><pubDate>Wed, 01 Jul 2026 09:00:00 +0000</pubDate><description>Default blocking and pay-per-crawl changed who can read the web. An assessment of what actually happened.</description><content:encoded><![CDATA[<p>A year ago today a CDN sitting in front of a large fraction of the web flipped its default: AI crawlers blocked unless explicitly allowed, with a marketplace for charging per crawl.</p>
<p>Enough time has passed to say what actually happened rather than what everyone predicted.</p>
<h2 id="what-changed">what changed<a class="anchor" href="#what-changed" aria-label="link to this section">#</a></h2>
<p><strong>The norm inverted.</strong> Before, crawling was permitted by default and <code>robots.txt</code> was a request. Now, for a large share of the web, crawling is denied by default and access is a negotiation.</p>
<p>That is a genuine structural change to how the web works and it happened through one company's configuration default rather than through any standards process, legislation, or public deliberation.</p>
<p><strong>Licensing deals concentrated.</strong> Large AI companies negotiated bulk access with large publishers. That was always the likely outcome: the parties with lawyers and leverage made arrangements, and the arrangements are private.</p>
<p><strong>Small publishers got very little.</strong> The pay-per-crawl mechanism works, technically. The revenue for a site with modest traffic is negligible — the arithmetic never supported anything else. The publishers who most needed a new economic model got the one that pays least.</p>
<p><strong>Non-commercial crawling got harder.</strong> Academic researchers, archivists, and independent search projects have no licensing budget and no negotiating position. Carve-outs exist and are discretionary, which means the ability to study the web is now something you apply for.</p>
<p>That is the outcome I was most worried about and it is the one that materialized most clearly.</p>
<h2 id="what-did-not-change">what did not change<a class="anchor" href="#what-did-not-change" aria-label="link to this section">#</a></h2>
<p><strong>Training data supply.</strong> The frontier labs have enormous existing corpora, licensed sources, and synthetic generation. The marginal value of newly crawled web text was already declining. Restricting it did not create the leverage publishers hoped for.</p>
<p><strong>Traffic.</strong> Referral traffic from search to publishers continued its decline, driven by AI answers in search results, which is a completely separate mechanism from training crawlers and which blocking crawlers does nothing about.</p>
<p>This is the part that was most misunderstood at the time. The traffic problem and the training problem have different causes and blocking crawlers only addresses one of them — the one with less economic impact.</p>
<h2 id="the-thing-to-actually-take-from-it">the thing to actually take from it<a class="anchor" href="#the-thing-to-actually-take-from-it" aria-label="link to this section">#</a></h2>
<p><strong>Infrastructure defaults are policy.</strong> A configuration change at a chokepoint reshaped access to a large fraction of the web, with no process and no appeal.</p>
<p>That is not a criticism of the specific decision, which was popular and defensible. It is an observation about where power actually sits, and it generalizes: the entities that can change the web's behavior are the ones with concentration at a layer everyone depends on, and there are about five of them.</p>
<p><strong>For your own site</strong>, the decision remains yours and it is worth making deliberately rather than accepting a default:</p>
<ul><li><strong>Documentation sites</strong> frequently want to be in the training data. Being the thing the model knows about is worth more than the pageview you did not get.</li><li><strong>Original reporting and analysis</strong> has a stronger case for restriction.</li><li><strong>Anything you want found</strong> should still permit search crawlers, which are a different category and are frequently blocked by accident when people configure this.</li></ul>
<p>Check what you are actually blocking. A meaningful number of sites blocked their own search indexing in the first months of this and did not notice for weeks.</p>
<h2 id="the-unresolved-thing">the unresolved thing<a class="anchor" href="#the-unresolved-thing" aria-label="link to this section">#</a></h2>
<p>The web's economic model — publish freely, get traffic, monetize traffic — is breaking, and nothing has replaced it.</p>
<p>Crawler tolls are not the replacement; the arithmetic does not work at the scale of the actual web, where most content is made by people with no ability to negotiate anything.</p>
<p>Licensing deals are not the replacement either; they work for a few hundred large publishers and for nobody else.</p>
<p>I do not know what the replacement is. I am increasingly convinced that nobody does, and that the interval between the old model failing and a new one existing is going to be long and is going to be bad for the open web.</p>
<p>That is a genuinely pessimistic conclusion and I have not found a way around it in a year of thinking about it.</p>]]></content:encoded></item><item><title>Refactoring under an agent</title><link>https://readme.news/refactoring-under-an-agent/</link><guid isPermaLink="true">https://readme.news/refactoring-under-an-agent/</guid><pubDate>Fri, 12 Jun 2026 09:00:00 +0000</pubDate><description>Large mechanical refactors are the clearest win available from coding agents. Here is the process that keeps them safe.</description><content:encoded><![CDATA[<p>Large mechanical refactors — rename this concept across four hundred files, migrate every call site to a new API, convert a pattern used everywhere — are the single clearest win available from coding agents.</p>
<p>They are also where an unattended agent can do the most damage quietly. Here is the process that has worked.</p>
<h2 id="why-this-is-the-sweet-spot">why this is the sweet spot<a class="anchor" href="#why-this-is-the-sweet-spot" aria-label="link to this section">#</a></h2>
<p>Mechanical refactors have exactly the properties agents are good at:</p>
<ul><li><strong>Verifiable.</strong> The tests either pass or they do not.</li><li><strong>Repetitive.</strong> The same transformation, many times, which is where humans make mistakes from fatigue.</li><li><strong>Well-specified.</strong> You can state the transformation precisely.</li><li><strong>Boring.</strong> Nobody enjoys this work and nobody does it carefully after hour two.</li></ul>
<p>And they have the property that makes human refactoring risky: <strong>the scale defeats attention.</strong> A person converting four hundred call sites will be careful for the first fifty.</p>
<h2 id="the-process">the process<a class="anchor" href="#the-process" aria-label="link to this section">#</a></h2>
<p><strong>1. Do ten by hand first.</strong></p>
<p>Before writing any instruction, do a representative sample yourself. You will discover:</p>
<ul><li>The cases where the mechanical transformation is wrong.</li><li>The variations you did not know existed.</li><li>What the actual rule is, as opposed to what you thought it was.</li></ul>
<p>This is the step people skip and it is the one that determines whether the whole thing works. You cannot specify a transformation you have not performed.</p>
<p><strong>2. Write the transformation down precisely.</strong></p>
<p>Not "modernize the error handling." The exact before and after, with the exceptions named:</p>
<div class="code"><span class="code-lang">markdown</span><pre><code class="lang-markdown">Replace every call of the form:
    result, err := doThing(x)
    if err != nil { return nil, err }
with:
    result, err := doThing(x)
    if err != nil { return nil, fmt.Errorf("doing thing for %s: %w", x.ID, err) }

Exceptions:
- Do not change anything in internal/legacy/ (frozen).
- Do not change error handling inside deferred functions.
- If the error is already wrapped, leave it.</code></pre></div>
<p><strong>3. Establish the safety net before you start.</strong></p>
<ul><li>Clean working tree, dedicated branch.</li><li>Full test suite passing, with a recorded baseline.</li><li>A way to check the transformation was applied correctly beyond the tests — a grep, a linter rule, an AST query.</li></ul>
<p><strong>4. Batch it.</strong></p>
<p>Not four hundred files in one change. Twenty to fifty files per commit, grouped by module.</p>
<p>This matters for two reasons: a reviewable diff size, and the ability to bisect. If something is wrong, you want to know which batch introduced it.</p>
<p><strong>5. Review the first batch line by line.</strong></p>
<p>Every line. This is where you catch the systematic error, and a systematic error caught in batch one costs twenty minutes while the same error caught in batch twenty costs a day.</p>
<p><strong>6. Spot-check subsequent batches, review the anomalies.</strong></p>
<p>Once the pattern is verified, review by sampling — but read <em>every</em> diff that looks different from the pattern. Agent output that deviates from the established shape is where the interesting failures are.</p>
<p><strong>7. Verify mechanically at the end.</strong></p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash"># nothing left in the old form
rg 'return nil, err$' --type go | rg -v 'internal/legacy'</code></pre></div>
<p>If the transformation is complete, the old pattern should not exist outside the exceptions. This catches the files that were silently skipped, which is a real failure mode.</p>
<h2 id="where-it-goes-wrong">where it goes wrong<a class="anchor" href="#where-it-goes-wrong" aria-label="link to this section">#</a></h2>
<p><strong>The agent "improves" things you did not ask about.</strong> It renames a variable while fixing the error handling. Individually reasonable, collectively it makes the diff unreviewable because you can no longer scan for the pattern.</p>
<p>Instruct explicitly: change only what was specified, nothing else.</p>
<p><strong>Semantic drift across batches.</strong> Batch one wraps errors one way, batch fifteen does it slightly differently, because the context is different and the model made a different reasonable choice.</p>
<p>Fix: put the exact target form in the instruction file, with examples, and check for consistency at the end with a grep.</p>
<p><strong>Tests pass and behavior changed.</strong> The most dangerous case. Your tests did not cover the path that broke.</p>
<p>This is why the mechanical verification in step seven matters. It checks the transformation, not the behavior, and it catches things tests do not.</p>
<p><strong>Silent skips.</strong> The agent processes 380 of 400 files and reports success. The twenty it skipped are the ones with unusual structure, which are the ones most likely to matter.</p>
<p>Always count. Always verify the remainder is empty.</p>
<h2 id="the-honest-assessment">the honest assessment<a class="anchor" href="#the-honest-assessment" aria-label="link to this section">#</a></h2>
<p>For this class of work the leverage is real and large. A refactor that would have been a week of tedium — and would therefore never have happened, which is why the codebase has the problem — becomes an afternoon.</p>
<p>That last part is the underrated benefit. The refactors that get done are the ones that are cheap enough to do. Lowering the cost means more of them happen, and a codebase where the cross-cutting cleanup actually gets done is a meaningfully better codebase.</p>
<p>Just count the files at the end.</p>]]></content:encoded></item><item><title>WWDC 2026 and the platform that keeps its own counsel</title><link>https://readme.news/wwdc-2026-and-the-platform-that-keeps-its-own-counsel/</link><guid isPermaLink="true">https://readme.news/wwdc-2026-and-the-platform-that-keeps-its-own-counsel/</guid><pubDate>Mon, 08 Jun 2026 09:00:00 +0000</pubDate><description>New OS versions, more on-device model surface, and a developer relationship that remains complicated.</description><content:encoded><![CDATA[<p>Apple's developer conference ran this week. The pattern of the last few years continues: strong <a class="xref" href="/pixel-10-and-the-on-device-model-as-a-platform-feature/" title="Pixel 10 and the on-device model as a platform feature">on-device</a> capability, tight platform integration, and a set of platform policy questions that are being settled in courtrooms rather than on stage.</p>
<h2 id="the-on-device-strategy-holding">the on-device strategy, holding<a class="anchor" href="#the-on-device-strategy-holding" aria-label="link to this section">#</a></h2>
<p>Apple's position has been consistent and, I think, correct for their situation:</p>
<ul><li>A <a class="xref" href="/haiku-45-and-the-collapsing-cost-of-good-enough/" title="Haiku 4.5 and the collapsing cost of good-enough">small model</a> on device, free, private, offline, exposed to third-party apps through a constrained system API.</li><li>A larger model available for the cases the small one cannot handle.</li><li>Routing handled by the system.</li><li>Privacy as the differentiating property rather than raw capability.</li></ul>
<p>That is not going to win a benchmark comparison and it was never trying to. It is going to win on the axis where Apple competes, which is a coherent product where the default behavior is the one most users want.</p>
<p>For developers, the practical implication has not changed: <strong>design features so the small model handles the common case.</strong></p>
<p>Concretely — the 90% that the on-device model can do is free, instant, and works on a plane. The 10% that needs escalation costs money and needs a network. Getting that split right is the engineering, and it is a different skill from prompt design.</p>
<h2 id="the-swift-trajectory">the Swift trajectory<a class="anchor" href="#the-swift-trajectory" aria-label="link to this section">#</a></h2>
<p>Swift continues its expansion beyond Apple platforms — server-side, embedded, <a class="xref" href="/the-component-model-and-the-plugin-problem/" title="The component model and the plugin problem">WebAssembly</a>, <a class="xref" href="/cross-platform-is-a-promise-you-make-to-your-budget/" title="Cross-platform is a promise you make to your budget">cross-platform</a> tooling. The concurrency model's strict checking continues to produce a language that catches data races at compile time, which is a genuinely valuable property that the migration cost has made contentious.</p>
<p>The honest state: strict concurrency checking is correct and it is a real migration burden for existing codebases, and the <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a> when you get it wrong are still harder to act on than they should be.</p>
<p>Swift is a better language than it gets credit for outside the Apple ecosystem and its adoption outside that ecosystem remains limited, mostly for reasons of tooling gravity rather than language quality.</p>
<h2 id="the-platform-policy-question">the platform policy question<a class="anchor" href="#the-platform-policy-question" aria-label="link to this section">#</a></h2>
<p>The regulatory pressure on app distribution and payments continues, differently in different jurisdictions, with the result that the rules now vary by region in ways that are genuinely confusing for developers to comply with.</p>
<p>I do not have a clean position on the substance. What I will say is the practical part:</p>
<p><strong>If you ship an app with any monetization, you now have jurisdiction-specific compliance work</strong>, and the rules are still moving. Budget for it, watch the changes, and do not assume a policy you read last year is current.</p>
<p>The broader observation, which is not about Apple specifically: the era of one global set of platform rules is over. Every major platform now operates under different obligations in the EU, the US, and several other markets, and that fragmentation is going to increase rather than resolve.</p>
<h2 id="the-things-worth-actually-adopting">the things worth actually adopting<a class="anchor" href="#the-things-worth-actually-adopting" aria-label="link to this section">#</a></h2>
<p>Filtering out the keynote, the items that will matter to a working developer:</p>
<p><strong>Anything that reduces app size or launch time.</strong> These are the metrics that correlate with retention and they get the least stage time.</p>
<p><strong>Testing and debugging improvements in the toolchain.</strong> Consistently the most valuable and least covered part of every WWDC.</p>
<p><strong>Deprecation notices.</strong> The most important announcements at any platform conference are the ones about what is going away, and they are never in the keynote. Read the release notes.</p>
<h2 id="the-honest-assessment">the honest assessment<a class="anchor" href="#the-honest-assessment" aria-label="link to this section">#</a></h2>
<p>Apple ships coherent, well-integrated platforms with genuinely good on-device capability and a privacy posture that is a real product differentiator rather than only marketing.</p>
<p>It also operates the developer relationship with less flexibility than any comparable platform, and the regulatory environment is the mechanism by which that is being adjusted rather than any change of heart.</p>
<p>Both of those have been true for a decade and neither is changing this year.</p>]]></content:encoded></item><item><title>Model evaluation for people who ship</title><link>https://readme.news/model-evaluation-for-people-who-ship/</link><guid isPermaLink="true">https://readme.news/model-evaluation-for-people-who-ship/</guid><pubDate>Wed, 27 May 2026 09:00:00 +0000</pubDate><description>Not research benchmarks. A practical harness you can build in a day that makes every future model decision an hour instead of a week.</description><content:encoded><![CDATA[<p>Every few weeks a new model ships and someone asks whether you should switch. Without an evaluation harness, answering that takes a week of impressions and you will get it wrong. With one, it takes an hour.</p>
<p>This is the highest-return day of engineering available to anyone building on models, and a surprising number of teams have not done it.</p>
<h2 id="what-it-is-not">what it is not<a class="anchor" href="#what-it-is-not" aria-label="link to this section">#</a></h2>
<p>Not MMLU. Not a leaderboard. Not "vibes after twenty prompts."</p>
<p>Public benchmarks tell you which models are worth testing. They do not predict performance on your task, because your task is not in them and because contamination is universal.</p>
<h2 id="what-it-is">what it is<a class="anchor" href="#what-it-is" aria-label="link to this section">#</a></h2>
<p>Fifty to two hundred examples from your actual production traffic, with expected outputs or a grading rubric, run automatically, producing a number.</p>
<div class="code"><pre><code>evals/
  cases/
    001-refund-request.json
    002-ambiguous-address.json
    ...
  run.py
  results/
    2026-05-27-model-a.json</code></pre></div>
<p>Each case:</p>
<div class="code"><span class="code-lang">json</span><pre><code class="lang-json">{
  "id": "042",
  "input": { "ticket": "my order never arrived and I want my money back" },
  "expect": { "category": "refund", "urgency": "high", "needs_human": false },
  "notes": "the word 'never' should not trigger the fraud path"
}</code></pre></div>
<p>That is it. The whole thing is a test suite where the assertions are fuzzier.</p>
<h2 id="where-the-cases-come-from">where the cases come from<a class="anchor" href="#where-the-cases-come-from" aria-label="link to this section">#</a></h2>
<p><strong>Your production failures.</strong> This is the best source by a wide margin. Every time the system gets something wrong, that becomes a case. Your eval set grows into a precise map of your problem's difficulty.</p>
<p>Set this up as a workflow: a thumbs-down in the product, or a support escalation, creates a candidate case that someone reviews and adds.</p>
<p><strong>Your edge cases.</strong> The weird inputs. The empty ones. The ones in another language. The adversarial ones. The ones with an injection attempt.</p>
<p><strong>A stratified sample of normal traffic.</strong> So you notice when a change breaks the common case while fixing an edge case.</p>
<p><strong>Cases with no correct answer.</strong> Where the right behavior is to refuse, escalate, or ask a clarifying question. Models are frequently bad at this and it is rarely tested.</p>
<h2 id="grading">grading<a class="anchor" href="#grading" aria-label="link to this section">#</a></h2>
<p>Three approaches, and you will use all three.</p>
<p><strong>Exact or structural match.</strong> For classification, extraction, and structured output. Cheap, deterministic, unambiguous. Use it wherever you can.</p>
<p><strong>Programmatic checks.</strong> For generated code: does it compile, do the tests pass. For SQL: does it run, does it return the right shape. This is the strongest form of grading and it is available more often than people realize — if you can verify mechanically, do.</p>
<p><strong>Model-as-judge.</strong> For open-ended output. A second model grades against a rubric.</p>
<p>Use it carefully:</p>
<ul><li><strong>Write a specific rubric</strong>, not "is this good." Score each dimension separately.</li><li><strong>Validate the judge against human ratings</strong> on a sample. If the judge disagrees with you, the judge is wrong and the rubric needs work.</li><li><strong>Use a different model than the one being evaluated</strong>, or at minimum be aware of self-preference bias, which is well documented and large.</li></ul>
<h2 id="the-metrics">the metrics<a class="anchor" href="#the-metrics" aria-label="link to this section">#</a></h2>
<p><strong>Accuracy on your set</strong>, obviously.</p>
<p><strong>Cost per case.</strong> Tokens in and out, at current prices. Track this — a model that is 2% better and 4× the cost is usually the wrong choice.</p>
<p><strong>Latency, at percentiles.</strong> p50 and p95. Reasoning models have high variance and the average hides it.</p>
<p><strong>Failure mode distribution.</strong> Not just how many wrong — <em>how</em> wrong. A model that fails by refusing is very different from one that fails by confidently fabricating, and the aggregate score treats them identically.</p>
<h2 id="the-workflow">the workflow<a class="anchor" href="#the-workflow" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">python evals/run.py --model model-a --model model-b --parallel 8</code></pre></div>
<p>Output a table. Commit the results. Diff them across runs.</p>
<p>Run it:</p>
<ul><li><strong>When any model ships.</strong> Within an hour of the announcement, you know.</li><li><strong>When you change a prompt.</strong> Prompt changes are code changes with no <a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">type system</a> and no compiler; the eval is your only regression check.</li><li><strong>On a schedule.</strong> Providers update models behind stable names. Behavior drifts. You want to know from your <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>, not from your support queue.</li></ul>
<p>That last one catches something most teams never notice: the model you deployed against is not the model serving your traffic today.</p>
<h2 id="the-one-that-matters-most">the one that matters most<a class="anchor" href="#the-one-that-matters-most" aria-label="link to this section">#</a></h2>
<p><strong>Version your prompts and store the eval result with them.</strong></p>
<p>A prompt is production configuration. It should be in version control, reviewed, and associated with a measured quality number. "Someone edited the prompt and something got worse three weeks ago" is a debugging session that should not be possible.</p>
<h2 id="the-payoff">the payoff<a class="anchor" href="#the-payoff" aria-label="link to this section">#</a></h2>
<p>Once this exists:</p>
<ul><li>Model migrations are an afternoon.</li><li>Prompt changes are safe to make.</li><li>You can argue about model choice with data instead of anecdotes.</li><li>You detect provider-side drift.</li><li>Onboarding a new engineer to the AI parts of your system means handing them the eval set, which is the best available documentation of what the system is supposed to do.</li></ul>
<p>It is a day of work. It is the single highest-leverage day available in this space and it has been for three years.</p>]]></content:encoded></item><item><title>I/O 2026 and the assistant that lives in everything</title><link>https://readme.news/io-2026-and-the-assistant-that-lives-in-everything/</link><guid isPermaLink="true">https://readme.news/io-2026-and-the-assistant-that-lives-in-everything/</guid><pubDate>Fri, 08 May 2026 09:00:00 +0000</pubDate><description>More model, more surfaces, and a search product that keeps changing what the web is for.</description><content:encoded><![CDATA[<p>Google's developer conference happened this week and the shape is consistent with where the company has been heading since 2024: a capable model, deployed everywhere they already have users, priced aggressively because they own the silicon.</p>
<h2 id="the-distribution-advantage-compounding">the distribution advantage, compounding<a class="anchor" href="#the-distribution-advantage-compounding" aria-label="link to this section">#</a></h2>
<p>The thing no competitor can replicate is that Google can ship a capability into products that billions of people already open daily, on launch day.</p>
<p>That is worth more than a benchmark lead and it is becoming more visible each year. A model that is marginally better but reaches users through a signup flow loses to a model that is marginally worse and is already in the search box.</p>
<p>The strategic implication for everyone else — including the other frontier labs — is that raw capability is not the competition anymore. Distribution, price, and integration are.</p>
<h2 id="the-search-question-again">the search question, again<a class="anchor" href="#the-search-question-again" aria-label="link to this section">#</a></h2>
<p>Every year this conference makes the same thing more true: informational queries are increasingly answered on the results page rather than by sending someone to a site.</p>
<p>For anyone who publishes on the web, the consequences are now well past theoretical:</p>
<ul><li><strong>Referral traffic to informational content keeps falling.</strong> This is measurable and it is not recovering.</li><li><strong>Your documentation is being summarized by a system you do not control</strong>, and users are acting on the summary.</li><li><strong>The <a class="xref" href="/why-your-tests-are-slow/" title="Why your tests are slow">feedback loop</a> is broken.</strong> You cannot see what people asked, what answer they got, or whether it was right.</li></ul>
<p>I do not have a satisfying answer. The mitigations available to an individual project are marginal: write documentation that is hard to summarize badly, keep a machine-readable <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a>, make <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a> self-explanatory so they do not require a search at all.</p>
<p>The structural problem — that the economic model funding web content is being removed without a replacement — is not solvable by any individual publisher, and the people who could solve it have no incentive to.</p>
<h2 id="the-developer-surface">the developer surface<a class="anchor" href="#the-developer-surface" aria-label="link to this section">#</a></h2>
<p>The genuinely useful announcements, as always, are the boring ones:</p>
<p><strong>Model pricing and the cheap tier.</strong> The cost-per-capability at the small end continues falling, and Google's TPU position means they can price below what competitors renting accelerators can match. If you run high-volume inference, the arithmetic is worth redoing quarterly.</p>
<p><strong>Longer context, better retention.</strong> Incremental and real. The practical question is always the degradation curve, not the maximum, and that continues to improve.</p>
<p><strong>Agent tooling in the cloud.</strong> Same category everyone is building: runtimes, memory, identity, observability. Evaluate against a real problem rather than a demo.</p>
<h2 id="the-thing-to-actually-do">the thing to actually do<a class="anchor" href="#the-thing-to-actually-do" aria-label="link to this section">#</a></h2>
<p>The recommendation has not changed in two years and I will keep repeating it because it keeps being right:</p>
<p><strong>Have an eval set in your repository.</strong> Fifty examples from your real domain, with expected outputs, run against every candidate model.</p>
<p>Every conference like this produces a new model that is claimed to be better. With an eval harness, evaluating that claim for your workload takes an hour. Without one, it takes a week of impressions and you will get it wrong.</p>
<p>This is a day of setup that pays back on every model release, forever, and the number of teams that have done it remains surprisingly small.</p>
<h2 id="the-honest-summary">the honest summary<a class="anchor" href="#the-honest-summary" aria-label="link to this section">#</a></h2>
<p>A very good model, deployed extremely well, in a company with structural advantages that are getting stronger.</p>
<p>Whether that is good for the web is a separate question, and I keep arriving at the same uncomfortable answer.</p>]]></content:encoded></item><item><title>Build 2026 and the platform that keeps absorbing</title><link>https://readme.news/build-2026-and-the-platform-that-keeps-absorbing/</link><guid isPermaLink="true">https://readme.news/build-2026-and-the-platform-that-keeps-absorbing/</guid><pubDate>Mon, 04 May 2026 09:00:00 +0000</pubDate><description>Agent infrastructure at the OS layer, more of the developer stack open-sourced, and a strategy that has not changed in a decade.</description><content:encoded><![CDATA[<p>Microsoft's developer conference ran this week. The individual announcements matter less than the consistency of the strategy, which has been unchanged for about ten years and keeps working.</p>
<h2 id="the-strategy">the strategy<a class="anchor" href="#the-strategy" aria-label="link to this section">#</a></h2>
<ol><li>Meet developers where they are, including on other people's platforms.</li><li>Adopt other people's standards rather than inventing competing ones.</li><li>Open-source the layers where control is not worth the friction.</li><li>Monetize the cloud underneath.</li></ol>
<p>Every Build for a decade has been an execution of that, and the cumulative result is that a company which was actively hostile to open source in 2005 is now the largest corporate contributor to it and owns the default editor, the default code host, and a large share of the developer toolchain.</p>
<h2 id="the-agent-infrastructure">the agent infrastructure<a class="anchor" href="#the-agent-infrastructure" aria-label="link to this section">#</a></h2>
<p>The substantive announcements this year continue the theme of putting agent capabilities at the operating system layer rather than in an application: a <a class="xref" href="/nodejs-24-and-the-slow-reinvention-of-the-runtime/" title="Node.js 24 and the slow reinvention of the runtime">permission model</a> for what agents may reach, an identity model for agents acting on a user's behalf, and audit surfaces for what they did.</p>
<p>That is the right layer for it. The alternative — every application implementing its own agent permission model — produces exactly the inconsistency that made mobile permissions a mess for a decade before the platforms standardized.</p>
<p>The questions that matter and that a keynote cannot answer:</p>
<p><strong>Granularity.</strong> "Filesystem access" is not a permission, it is a surrender. Does the model support "read from this directory for this task"?</p>
<p><strong>Consent fatigue.</strong> If the prompts are frequent, users click through them, and the control is theater. The design problem is asking rarely and meaningfully.</p>
<p><strong>Revocation and audit.</strong> Can a user see what an agent did and undo it? This is the part that is hardest and gets the least attention.</p>
<p>I will believe the <a class="xref" href="/the-component-model-and-the-plugin-problem/" title="The component model and the plugin problem">security model</a> when someone publishes an analysis of it, not when it is demonstrated on a stage.</p>
<h2 id="the-enterprise-angle">the enterprise angle<a class="anchor" href="#the-enterprise-angle" aria-label="link to this section">#</a></h2>
<p>The genuinely differentiating position Microsoft has is that they can offer agent capabilities inside an enterprise's existing identity, compliance, and audit infrastructure.</p>
<p>That is worth more to a large organization than raw capability. An agent that works within the existing access control model, logs to the existing audit system, and is governed by the existing data policies clears procurement. An agent that requires a new trust boundary does not, regardless of how good it is.</p>
<p>This is the same advantage that won enterprise cloud and it is being applied identically.</p>
<h2 id="what-a-developer-should-actually-do-with-this">what a developer should actually do with this<a class="anchor" href="#what-a-developer-should-actually-do-with-this" aria-label="link to this section">#</a></h2>
<p><strong>If you build on Windows:</strong> the tooling story is genuinely good now — WSL, the terminal, winget, PowerShell 7 — and if your Windows support has been a grudging afterthought since 2018, it is worth revisiting.</p>
<p><strong>If you build agent-adjacent products:</strong> design against the OS permission model rather than around it. Products that require users to disable platform protections do not get enterprise adoption.</p>
<p><strong>If you are evaluating anything announced here:</strong> wait for the second version. Microsoft's first releases in a new category are consistently rough and consistently improved within a year. That is a reasonable pattern and it means the launch-day evaluation is not the useful one.</p>
<h2 id="the-pattern-to-watch">the pattern to watch<a class="anchor" href="#the-pattern-to-watch" aria-label="link to this section">#</a></h2>
<p>The layer where the industry is currently fighting is not the model. It is the control plane for agents: who they are, what they may do, on whose behalf, with what audit trail.</p>
<p>Every platform vendor is building this. The one that becomes standard will have the same kind of position that identity providers have today, and it will be very durable.</p>
<p>That is the strategic story of the next three years and it is being fought in permission dialogs rather than benchmarks.</p>]]></content:encoded></item><item><title>Hiring juniors in 2026</title><link>https://readme.news/hiring-juniors-in-2026/</link><guid isPermaLink="true">https://readme.news/hiring-juniors-in-2026/</guid><pubDate>Thu, 30 Apr 2026 09:00:00 +0000</pubDate><description>The entry-level pipeline is breaking in a way that will be expensive in five years. Some of the fixes are cheap.</description><content:encoded><![CDATA[<p>Entry-level software hiring has contracted sharply. The reasons are partly cyclical and partly a genuine belief among hiring managers that AI tooling has reduced the need for junior engineers.</p>
<p>The first part will recover. The second part is a mistake, and it is the kind of mistake that is invisible for four years and then extremely expensive.</p>
<h2 id="the-argument-for-not-hiring-juniors">the argument for not hiring juniors<a class="anchor" href="#the-argument-for-not-hiring-juniors" aria-label="link to this section">#</a></h2>
<p>Stated honestly, because it is not stupid:</p>
<p>A junior engineer's first-year output is mostly small, well-specified tasks — the CRUD endpoint, the test coverage, the bug in a file someone pointed them at. That work is now substantially automatable. Meanwhile the junior requires mentoring time from senior engineers, which is the scarcest resource.</p>
<p>So the ROI on a junior looks worse than it did. That reasoning is coherent.</p>
<h2 id="why-it-is-wrong">why it is wrong<a class="anchor" href="#why-it-is-wrong" aria-label="link to this section">#</a></h2>
<p><strong>Seniors come from juniors.</strong> There is no other supply. An organization that hires only seniors is free-riding on other organizations' training, and if everyone does it, the pipeline empties. This is a classic collective action failure and the industry is walking into it with open eyes.</p>
<p><strong>The judgment that makes seniors valuable comes from doing the work.</strong> The ability to look at plausible code and know it is wrong comes from having written the wrong version and debugged it at 3 a.m. You cannot read your way to it and you cannot prompt your way to it.</p>
<p>If the apprenticeship stops, the next generation of senior engineers does not exist, and the current one retires.</p>
<p><strong>Juniors are better at the new tools.</strong> Consistently, in my experience. They have no prior workflow to defend and they explore. A team of only senior engineers adopts new tooling slowly and grudgingly.</p>
<p><strong>Mentoring makes seniors better.</strong> The engineer who has to explain why a design is wrong understands it better afterward. Teams with no juniors lose that forcing function and get sloppier about articulating their own reasoning.</p>
<h2 id="what-actually-has-to-change">what actually has to change<a class="anchor" href="#what-actually-has-to-change" aria-label="link to this section">#</a></h2>
<p>The old model — hire a junior, give them small tickets for a year, gradually increase scope — does not work as well when the small tickets are automated. The model has to change, not the hiring.</p>
<p><strong>Start them on reading, not writing.</strong> Give a new engineer a real system and a week to understand and explain it. Have them write the architecture document that does not exist. This builds the skill that actually matters now and it produces something useful.</p>
<p><strong>Give them debugging, not features.</strong> Debugging is the skill that generalizes, that AI is least reliable at, and that cannot be learned from a course. Pair them on incidents. Give them the flaky test nobody wants.</p>
<p><strong>Make them review agent output.</strong> Reviewing machine-generated code with a senior engineer walking through what is wrong with it is an extraordinarily efficient teaching mechanism. You get a stream of plausible-but-flawed code, which is exactly the training material you want and which used to be expensive to produce.</p>
<p><strong>Require them to write the tests first.</strong> Specifying behavior before implementing teaches design, and it is the part of the workflow that has become more important rather than less.</p>
<p><strong>Do not let them delegate the hard part.</strong> For the first year, some things get done by hand, deliberately, because the point is the learning rather than the output. Say this out loud so it does not feel like an arbitrary restriction.</p>
<h2 id="the-hiring-signal-that-works-now">the hiring signal that works now<a class="anchor" href="#the-hiring-signal-that-works-now" aria-label="link to this section">#</a></h2>
<p>Traditional junior screens — implement this algorithm, complete this take-home — are substantially defeated and were never good predictors anyway.</p>
<p>What works better:</p>
<p><strong>A code review exercise.</strong> Give them a pull request with three problems. Watch what they find and how they talk about it. This is the job.</p>
<p><strong>A debugging exercise on a real repository.</strong> Failing test, thirty minutes, any tools they want including AI. Watch the process, not the outcome. Do they read the error? Form a hypothesis? Check it? Notice when the model's suggestion is wrong?</p>
<p><strong>A conversation about something they built.</strong> Follow-up questions until you hit the edge of their understanding. Where that edge sits, and how they handle reaching it, tells you almost everything.</p>
<h2 id="the-case-to-make-internally">the case to make internally<a class="anchor" href="#the-case-to-make-internally" aria-label="link to this section">#</a></h2>
<p>If you are arguing for junior headcount:</p>
<p>The cost of a junior is roughly a senior's partial attention for a year plus a below-market salary. The cost of a senior hire in three years, in a market where nobody trained anyone, is going to be considerably higher than it is now.</p>
<p>Every organization that stopped training in 2009 spent 2013 through 2016 paying enormous premiums for the engineers who had been trained elsewhere. It is the same trade and it is being made again.</p>
<p>The organizations that keep training through this period will have a meaningful advantage in five years, and it will be very hard to catch up to them quickly.</p>]]></content:encoded></item><item><title>The retrieval question, revisited</title><link>https://readme.news/the-retrieval-question-revisited/</link><guid isPermaLink="true">https://readme.news/the-retrieval-question-revisited/</guid><pubDate>Fri, 17 Apr 2026 09:00:00 +0000</pubDate><description>Context windows kept growing and retrieval did not die. It changed shape. Where the line actually sits now.</description><content:encoded><![CDATA[<p>Two years ago the argument was that growing context windows would make retrieval pipelines obsolete. I made a version of that argument and was partly wrong.</p>
<p>Here is where the line actually is now, with the reasoning rather than the conclusion, because the line keeps moving and the reasoning does not.</p>
<h2 id="the-four-variables">the four variables<a class="anchor" href="#the-four-variables" aria-label="link to this section">#</a></h2>
<p>Whether to retrieve or to stuff the context depends on four things:</p>
<p><strong>Corpus size.</strong> If everything fits, stuff it. "Fits" now means a large document set — a codebase, a product's full documentation, a year of one team's tickets. It does not mean an enterprise's entire document store.</p>
<p><strong>Query pattern.</strong> If you ask many questions against the same corpus, <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> the prefix makes stuffing cheap. If every query touches a different slice of a large corpus, retrieval wins.</p>
<p><strong>Latency budget.</strong> Prefill on a very large context takes real time even when cached. If you need sub-second responses, a small retrieved context is faster.</p>
<p><strong>Citation requirement.</strong> If you must show which source supported a claim, retrieval gives you that structurally. Stuffing requires trusting the model's attribution, which is less reliable than people assume.</p>
<h2 id="the-decision-concretely">the decision, concretely<a class="anchor" href="#the-decision-concretely" aria-label="link to this section">#</a></h2>
<p><strong>Stuff the context when:</strong> the corpus is under a few hundred thousand tokens, you query it repeatedly (so caching applies), and you want the model to see relationships between distant parts.</p>
<p>The clearest example remains a codebase. Retrieval over code performs worse than whole-repository context because code's meaning lives in the relationships between files, and chunking destroys exactly that.</p>
<p><strong>Retrieve when:</strong> the corpus is large, queries are diverse, latency matters, or you need citations.</p>
<p><strong>Do both when:</strong> you have a large corpus with a hot subset. Stuff the hot subset — the style guide, the schema, the core documents — and retrieve from the tail.</p>
<h2 id="what-actually-changed-retrieval-became-a-tool">what actually changed: retrieval became a tool<a class="anchor" href="#what-actually-changed-retrieval-became-a-tool" aria-label="link to this section">#</a></h2>
<p>The important architectural shift is not about size. It is that retrieval moved from a <strong>preprocessing step</strong> to a <strong>tool the model calls</strong>.</p>
<p>Old shape:</p>
<div class="code"><pre><code>query → embed → search → rerank → stuff top-k → generate</code></pre></div>
<p>One retrieval, before generation, with <code>k</code> fixed by you.</p>
<p>New shape:</p>
<div class="code"><pre><code>query → model reasons → calls search tool → reads results
      → reasons → searches again with a better query → reads
      → generates answer with citations</code></pre></div>
<p>The model decides what to search for, evaluates whether the results answered the question, and searches again with a refined query if not.</p>
<p>This is dramatically better and it is better for a specific reason: <strong>the first query is usually not the right query.</strong> A user asks about "the timeout issue." The right search is for the specific component's retry configuration, which you only know to search for after reading something else.</p>
<p>A fixed one-shot retrieval cannot do that. An agent with a search tool can.</p>
<h2 id="what-this-means-for-your-pipeline">what this means for your pipeline<a class="anchor" href="#what-this-means-for-your-pipeline" aria-label="link to this section">#</a></h2>
<p><strong>Embeddings matter less.</strong> When the model can iterate, a mediocre first retrieval is recoverable. When you had one shot, embedding quality was everything.</p>
<p><strong>Keyword search came back.</strong> Hybrid search — BM25 plus vectors — consistently outperforms pure vector search, and for a model that can iterate, plain keyword search is often sufficient. Exact terms matter: error codes, function names, product names. Embeddings are bad at exact match and always were.</p>
<p>If you built a pure-vector pipeline in 2023, adding keyword search is probably your biggest available quality improvement.</p>
<p><strong>Chunking matters less, and differently.</strong> Instead of chunking for retrieval, store documents whole and retrieve whole documents when they fit. Chunk only what is too large, and chunk on structural boundaries — sections, functions, headings — rather than by token count.</p>
<p><strong>Metadata filtering matters more.</strong> The model can specify constraints: this project, this date range, this author. Filtering is cheap, precise, and it is frequently what the query actually needed. Make sure your index supports it.</p>
<h2 id="the-practical-setup">the practical setup<a class="anchor" href="#the-practical-setup" aria-label="link to this section">#</a></h2>
<p>For most applications:</p>
<ol><li><strong>Postgres with <code>pgvector</code> plus full-text search.</strong> One system, hybrid search, metadata filters, and joins to your relational data.</li><li><strong>Expose search as a tool</strong>, not as a preprocessing step. Let the model iterate.</li><li><strong>Return whole documents</strong> where they fit; chunk on structure where they do not.</li><li><strong>Cache the stable prefix</strong> — the instructions, the schema, the core reference material.</li><li><strong>Measure retrieval quality separately from answer quality.</strong> If the answer is wrong, you need to know whether the right document was retrieved. Most teams cannot answer that and debug blind.</li></ol>
<p>That last one is the single most useful piece of instrumentation in a retrieval system and almost nobody has it.</p>
<h2 id="the-thing-i-got-wrong">the thing I got wrong<a class="anchor" href="#the-thing-i-got-wrong" aria-label="link to this section">#</a></h2>
<p>I said the RAG infrastructure category was solving a temporary problem. The capability got absorbed, as predicted. The infrastructure repositioned rather than disappearing, which I did not predict.</p>
<p>That is the third time I have watched this exact pattern and failed to apply the lesson. Infrastructure around a model limitation rarely dies. It moves to whatever the model still cannot do, which is usually one layer out.</p>]]></content:encoded></item><item><title>Agent fleets in production: a field report</title><link>https://readme.news/agent-fleets-in-production-a-field-report/</link><guid isPermaLink="true">https://readme.news/agent-fleets-in-production-a-field-report/</guid><pubDate>Fri, 27 Mar 2026 09:00:00 +0000</pubDate><description>Running many coding agents at once works better than expected on one axis and worse on every other. What actually happens.</description><content:encoded><![CDATA[<p>The pitch for delegated coding agents is parallelism: run five tasks at once, get five results, multiply throughput.</p>
<p>Having run this for a while at meaningful volume, here is what actually happens.</p>
<h2 id="what-works">what works<a class="anchor" href="#what-works" aria-label="link to this section">#</a></h2>
<p><strong>Mechanical, well-specified, verifiable work.</strong> This is not a hedge, it is the finding. The tasks where fleets genuinely deliver:</p>
<ul><li>Dependency upgrades across many services.</li><li>Migrating a deprecated API call across a large codebase.</li><li>Adding tests to modules with poor coverage.</li><li>Converting between formats or frameworks with a mechanical mapping.</li><li>Fixing a class of lint or type error across a repository.</li></ul>
<p>What these share: a clear acceptance criterion the agent can check itself, a bounded scope, and no design decisions.</p>
<p>For this category the leverage is real and large. Work that would have been a week of tedium becomes an afternoon of review.</p>
<p><strong>Investigation in parallel.</strong> Spawning several agents to independently investigate a bug from different angles — read the logs, bisect the history, read the related code, reproduce it — and reading all four reports is genuinely faster than doing them serially. The agents are cheap; your attention is not; parallelizing the cheap thing is correct.</p>
<h2 id="what-does-not-work">what does not work<a class="anchor" href="#what-does-not-work" aria-label="link to this section">#</a></h2>
<p><strong>Anything requiring shared context.</strong> Five agents working on related parts of a system produce five changes that each make sense and collectively do not. They duplicate helper functions. They pick different names for the same concept. They each add a slightly different retry wrapper.</p>
<p>You end up with a merge problem that is worse than the original work, because the conflicts are semantic rather than textual — the diffs apply cleanly and the result is incoherent.</p>
<p><strong>Anything with genuine design decisions.</strong> An agent given an ambiguous task will resolve the ambiguity confidently and silently, and it will pick a reasonable option that may not be the one you wanted. Five agents will each pick differently.</p>
<p><strong>Anything where the acceptance criterion is subjective.</strong> "Improve the <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a>" produces five different notions of improved.</p>
<h2 id="the-actual-bottleneck">the actual bottleneck<a class="anchor" href="#the-actual-bottleneck" aria-label="link to this section">#</a></h2>
<p>Review, exactly as predicted, and worse than predicted.</p>
<p>Five pull requests per hour is not a throughput improvement if you can meaningfully review two. What you get is a queue, and queue pressure degrades review quality in a way that is invisible in the metrics and visible in the defect rate a quarter later.</p>
<p>The honest arithmetic: <strong>your throughput is min(generation rate, review rate)</strong>, and review rate did not change.</p>
<h2 id="what-actually-helps">what actually helps<a class="anchor" href="#what-actually-helps" aria-label="link to this section">#</a></h2>
<p><strong>Make the agent produce a reviewable artifact, not just a diff.</strong> A summary of what it did, what it decided, and what it was unsure about. Reviewing "I chose to use the existing retry helper rather than adding a new one, and I left the timeout at 30s because the surrounding code does" is dramatically faster than inferring that from a diff.</p>
<p><strong>Batch related work into one agent, not many.</strong> If five tasks touch the same module, one agent doing all five produces a coherent change. Five agents produce five incoherent ones. Parallelism across <em>independent</em> work only.</p>
<p><strong>Invest heavily in verification.</strong> A strong test suite is what makes review cheaper, because you are reviewing design rather than correctness. Teams with weak tests get no benefit from agent fleets — they just get more code to verify by reading.</p>
<p><strong>Cap concurrency at your review capacity.</strong> Running more agents than you can review does not help. It produces a backlog that goes stale and gets abandoned, which is worse than not starting.</p>
<p><strong>Reject on size.</strong> A machine can generate a 2,000-line diff effortlessly. Say no. Ask it to split.</p>
<h2 id="the-number">the number<a class="anchor" href="#the-number" aria-label="link to this section">#</a></h2>
<p>For our work, the useful concurrency turned out to be <strong>two to three agents on independent tasks</strong>, with one person reviewing. Beyond that, quality degraded and the extra output was not landing.</p>
<p>That is a real multiplier and it is much less than the demos suggest, and the constraint is entirely on the human side.</p>
<h2 id="the-thing-i-would-tell-someone-starting">the thing I would tell someone starting<a class="anchor" href="#the-thing-i-would-tell-someone-starting" aria-label="link to this section">#</a></h2>
<p>Do not start with a fleet. Start with one agent, on the mechanical work you have been putting off, and measure whether the output actually lands.</p>
<p>If your review process cannot absorb one agent's output, adding four more is solving the wrong problem.</p>]]></content:encoded></item><item><title>Small models ate the middle</title><link>https://readme.news/small-models-ate-the-middle/</link><guid isPermaLink="true">https://readme.news/small-models-ate-the-middle/</guid><pubDate>Mon, 23 Feb 2026 09:00:00 +0000</pubDate><description>The capability floor rose faster than the ceiling. Most production inference no longer touches a frontier model.</description><content:encoded><![CDATA[<p>The most consequential trend in applied AI over the last eighteen months is not at the frontier. It is that the bottom of the market got good enough for most work.</p>
<h2 id="the-shape-of-it">the shape of it<a class="anchor" href="#the-shape-of-it" aria-label="link to this section">#</a></h2>
<p>Track any capability benchmark across model sizes over time and the pattern is consistent: the frontier improves steadily, and the small-model tier improves faster. The gap between "the best model available" and "a model that costs 2% as much" has been compressing.</p>
<p>The mechanisms are known:</p>
<ul><li><strong>Distillation.</strong> Training small models on the outputs of large ones transfers a surprising amount of capability. The recipes are public.</li><li><strong>Better data.</strong> Curated, synthetic, and filtered training data improves small models disproportionately, because they have less capacity to waste on noise.</li><li><strong>Mixture of experts.</strong> Total parameters for knowledge, <a class="xref" href="/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/" title="Llama 4 arrives, and the leaderboard problem gets a name">active parameters</a> for cost. The economics of a 30B-total/3B-active model are close to a 3B dense model, and the quality is much closer to a 30B dense one.</li><li><strong>Reasoning post-training.</strong> RL on verifiable rewards works on small models. A <a class="xref" href="/haiku-45-and-the-collapsing-cost-of-good-enough/" title="Haiku 4.5 and the collapsing cost of good-enough">small model</a> that thinks can outperform a large model that does not, on the tasks where thinking helps.</li></ul>
<h2 id="what-it-means-in-practice">what it means in practice<a class="anchor" href="#what-it-means-in-practice" aria-label="link to this section">#</a></h2>
<p>Go through a typical production AI workload and categorize the calls:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">task</th><th style="text-align:left">needs frontier?</th></tr></thead><tbody><tr><td style="text-align:left">classify a support ticket</td><td style="text-align:left">no</td></tr><tr><td style="text-align:left">extract fields from a document</td><td style="text-align:left">no</td></tr><tr><td style="text-align:left">summarize a thread</td><td style="text-align:left">no</td></tr><tr><td style="text-align:left">rewrite for tone</td><td style="text-align:left">no</td></tr><tr><td style="text-align:left">generate a SQL query from a question</td><td style="text-align:left">usually not</td></tr><tr><td style="text-align:left">route a request</td><td style="text-align:left">no</td></tr><tr><td style="text-align:left">draft a response</td><td style="text-align:left">usually not</td></tr><tr><td style="text-align:left">debug a subtle concurrency bug</td><td style="text-align:left">yes</td></tr><tr><td style="text-align:left">design a system from a vague brief</td><td style="text-align:left">yes</td></tr><tr><td style="text-align:left">plan a multi-step task with dependencies</td><td style="text-align:left">yes</td></tr></tbody></table></div>
<p>The first column is most of the volume. The second is most of the value per call and a small fraction of the calls.</p>
<p>If you are running everything through a frontier model, you are likely spending a large multiple of what you need to, and you probably have not measured which calls actually need it.</p>
<h2 id="the-architecture">the architecture<a class="anchor" href="#the-architecture" aria-label="link to this section">#</a></h2>
<div class="code"><pre><code>                  ┌─→ small model ──→ verifier ──→ ok? → done
request → route ──┤                              └→ no  ─┐
                  └─→ frontier model ←──────────────────┘</code></pre></div>
<p>Three components and each matters:</p>
<p><strong>The router.</strong> Simpler than people build. Input length, detected task type, and a handful of keywords gets you most of the way. A small model as a classifier works too. Do not build a sophisticated router before you have measured that a simple one is insufficient.</p>
<p><strong>The verifier.</strong> This is the part that gets skipped and it is what makes the architecture safe. Small models fail confidently. A cheap check — schema validation, a range assertion, running the generated code, a second model asked "is this answer plausible" — catches the failures that would otherwise reach production silently.</p>
<p><strong>The escalation path.</strong> When verification fails, retry with the frontier model. Track the rate. If it climbs, something changed.</p>
<h2 id="the-instrumentation-that-matters">the instrumentation that matters<a class="anchor" href="#the-instrumentation-that-matters" aria-label="link to this section">#</a></h2>
<p>Three metrics, on a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>:</p>
<ol><li><strong>Escalation rate.</strong> The fraction of requests that fall through to the expensive path. This is the health metric for the whole design.</li><li><strong>Cost per successful task</strong>, not per token. Tokens are an implementation detail; tasks are the unit you care about.</li><li><strong>Quality on your eval set, by tier.</strong> Run both tiers against the same evaluation regularly. When a new small model ships, you will know within an hour whether you can move work down.</li></ol>
<h2 id="the-second-order-effect">the second-order effect<a class="anchor" href="#the-second-order-effect" aria-label="link to this section">#</a></h2>
<p>When inference is nearly free, you use more of it.</p>
<p>Things that were not worth a model call become worth it: normalizing an input, double-checking an output, generating three candidates and picking one, enriching a record, summarizing an intermediate step.</p>
<p>The systems being built now use dramatically more model calls per user action than the systems from two years ago, at lower total cost. That is Jevons operating on schedule, and it is why "our AI costs went down per token and up in total" is the normal experience.</p>
<h2 id="the-thing-to-watch">the thing to watch<a class="anchor" href="#the-thing-to-watch" aria-label="link to this section">#</a></h2>
<p>The interesting question is not whether small models keep improving. They will.</p>
<p>It is whether the <em>frontier</em> keeps being worth its premium. If the gap on practically-relevant tasks keeps compressing, the frontier tier's addressable workload shrinks to a narrow band of genuinely hard problems.</p>
<p>That is a much smaller business than the one being priced today, and it is the scenario that ought to worry the labs more than competition does.</p>]]></content:encoded></item><item><title>Every company is briefly a model company</title><link>https://readme.news/every-company-is-briefly-a-model-company/</link><guid isPermaLink="true">https://readme.news/every-company-is-briefly-a-model-company/</guid><pubDate>Wed, 18 Feb 2026 09:00:00 +0000</pubDate><description>The fine-tuning wave, the RAG wave, and the agent wave all followed the same arc. Here&#x27;s where the value actually settled.</description><content:encoded><![CDATA[<p>Three times in three years, a wave of companies concluded that the way to build an AI product was to own a layer that turned out not to be theirs.</p>
<p>The pattern is consistent enough to be predictive, which makes it worth naming.</p>
<h2 id="wave-one-fine-tuning">wave one: fine-tuning<a class="anchor" href="#wave-one-fine-tuning" aria-label="link to this section">#</a></h2>
<p><strong>The thesis (2023):</strong> general models are generic. Fine-tune on your domain data and you get a model that is specifically good at your problem and that competitors cannot replicate.</p>
<p><strong>What happened:</strong> base models improved faster than fine-tunes could keep up. A fine-tuned model from six months ago was worse than the new base model with a good prompt. Every fine-tune had to be redone on every model release, which is a treadmill.</p>
<p><strong>Where it settled:</strong> fine-tuning is genuinely valuable for narrow, stable, high-volume tasks — classification into your specific taxonomy, output in your specific format, a <a class="xref" href="/haiku-45-and-the-collapsing-cost-of-good-enough/" title="Haiku 4.5 and the collapsing cost of good-enough">small model</a> matching a large model's behavior on one task. It is not a moat and it is not a product strategy.</p>
<h2 id="wave-two-rag">wave two: RAG<a class="anchor" href="#wave-two-rag" aria-label="link to this section">#</a></h2>
<p><strong>The thesis (2023-24):</strong> the model does not know your data. Build a retrieval pipeline — chunk, embed, index, retrieve, rerank — and you have a defensible system built on proprietary knowledge.</p>
<p><strong>What happened:</strong> context windows grew by two orders of magnitude, long-context quality improved, and prompt <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> made large contexts economical. A large fraction of naive RAG got replaced by putting the documents in the prompt.</p>
<p>Simultaneously, the pipeline components commoditized. Embedding models became interchangeable. Vector search became a feature of every database rather than a product.</p>
<p><strong>Where it settled:</strong> retrieval did not go away. It moved. Retrieval is now a <em>tool the model calls</em> rather than a preprocessing step, and it is genuinely necessary at large corpus sizes, where citation is required, and where cost or latency rules out large contexts.</p>
<p>The infrastructure repositioned rather than dying, which is what usually happens.</p>
<h2 id="wave-three-agent-frameworks">wave three: agent frameworks<a class="anchor" href="#wave-three-agent-frameworks" aria-label="link to this section">#</a></h2>
<p><strong>The thesis (2024-25):</strong> models cannot plan reliably. Build the orchestration — task decomposition, tool routing, retry logic, state management — and own the layer that makes agents work.</p>
<p><strong>What happened:</strong> the models absorbed it. In-context tool use during reasoning removed the need for an external loop. Native parallel tool calling removed the need for a dispatcher. Memory tools and <a class="xref" href="/claude-sonnet-45-and-the-agent-that-runs-for-thirty-hours/" title="Claude Sonnet 4.5 and the agent that runs for thirty hours">context editing</a> removed the need for external state management.</p>
<p>Each capability the framework provided became a model feature within about a year.</p>
<p><strong>Where it settled:</strong> in progress, but the shape is clear. Orchestration frameworks are converging on thin conveniences. The durable parts are evaluation, observability, and the domain-specific policy that no model will ever have.</p>
<h2 id="the-pattern">the pattern<a class="anchor" href="#the-pattern" aria-label="link to this section">#</a></h2>
<p>Every wave follows the same arc:</p>
<ol><li>The model has a limitation.</li><li>Companies build infrastructure to work around the limitation.</li><li>The limitation gets fixed in the model.</li><li>The infrastructure either finds a new position or disappears.</li></ol>
<p>The consistent error is <strong>building on a gap rather than on an asset</strong>. A gap is temporary by construction — the labs are actively working to close it, with more resources than you have.</p>
<h2 id="what-has-actually-been-durable">what has actually been durable<a class="anchor" href="#what-has-actually-been-durable" aria-label="link to this section">#</a></h2>
<p>Across all three waves, the same things kept their value:</p>
<p><strong>Proprietary data.</strong> Not "we have documents" — everyone has documents. Data that is genuinely yours: production logs, customer interactions, labeled outcomes, domain expertise encoded as examples. Nobody can buy it and no model was trained on it.</p>
<p><strong>Evaluation specific to your task.</strong> The company that knows, precisely, how well a system performs on their actual problem can adopt a new model in a day. The one that does not spends a month on vibes. That gap compounds every release cycle.</p>
<p><strong>Distribution and workflow integration.</strong> Being where the user already works. The most boring answer and the most durable.</p>
<p><strong>Domain constraints.</strong> The rules, regulations, edge cases, and institutional knowledge that make a generic capability into a usable product. This is unglamorous and it is the actual work.</p>
<p><strong>Trust.</strong> Security posture, compliance, reliability, support. Enterprises buy this and it takes years to build.</p>
<h2 id="the-test-to-apply">the test to apply<a class="anchor" href="#the-test-to-apply" aria-label="link to this section">#</a></h2>
<p>Before building on top of a model limitation, ask: <strong>if this limitation disappeared next quarter, what would I have left?</strong></p>
<p>If the answer is "nothing," you are building a bridge over a river that is being drained.</p>
<p>If the answer is "the data, the evaluations, the integrations, and the customer relationships," build it, and expect to throw the bridge away.</p>]]></content:encoded></item><item><title>The EU AI Act's August deadline</title><link>https://readme.news/the-eu-ai-acts-august-deadline/</link><guid isPermaLink="true">https://readme.news/the-eu-ai-acts-august-deadline/</guid><pubDate>Fri, 13 Feb 2026 09:00:00 +0000</pubDate><description>High-risk obligations arrive this summer. If you ship into Europe, the classification work should already be done.</description><content:encoded><![CDATA[<p>The EU AI Act's obligations for high-risk AI systems take effect on 2 August 2026. That is under six months away, and the compliance work — if it applies to you — is not a six-week project.</p>
<h2 id="the-structure-briefly">the structure, briefly<a class="anchor" href="#the-structure-briefly" aria-label="link to this section">#</a></h2>
<p>The Act is risk-tiered rather than technology-specific.</p>
<p><strong>Prohibited</strong> (in force since February 2025). Social scoring, certain biometric categorization, emotion recognition in workplaces and schools, untargeted facial image scraping, and manipulative techniques exploiting vulnerabilities.</p>
<p><strong>High-risk</strong> (August 2026 for most). Systems used in employment decisions, education access, credit scoring, essential services, law enforcement, migration, justice administration, and as safety components in regulated products.</p>
<p><strong>Limited risk.</strong> Transparency obligations. Tell people they are talking to an AI. Label synthetic media.</p>
<p><strong>Minimal risk.</strong> Most software. No specific obligations.</p>
<p><strong>General-purpose AI models.</strong> Separate obligations that began August 2025 — documentation, copyright policy, training data summaries, and additional requirements for models above a compute threshold.</p>
<h2 id="the-classification-question">the classification question<a class="anchor" href="#the-classification-question" aria-label="link to this section">#</a></h2>
<p>Most of the compliance work is answering "is my system high-risk," and the answer is less obvious than people assume.</p>
<p>The high-risk categories are defined by <em>use</em>, not by technology. A simple scoring model used to filter job applicants is high-risk. A sophisticated transformer used to suggest recipes is not.</p>
<p>Specific things that catch people:</p>
<ul><li><strong>Recruitment tooling.</strong> Anything that screens, ranks, or filters candidates. This includes a keyword filter, not just an ML model.</li><li><strong>Employee evaluation.</strong> Performance assessment, task allocation, monitoring that affects work relationships.</li><li><strong>Creditworthiness.</strong> Including proxies for it.</li><li><strong>Educational assessment.</strong> Grading, admission, proctoring.</li><li><strong>Safety components</strong> in products already covered by EU product safety legislation — which is a long list including machinery, medical devices, and vehicles.</li></ul>
<p>If any of those describe your product, you have obligations, and if you are a provider (you put it on the market) rather than a deployer (you use it), the obligations are substantial.</p>
<h2 id="what-high-risk-actually-requires">what high-risk actually requires<a class="anchor" href="#what-high-risk-actually-requires" aria-label="link to this section">#</a></h2>
<ul><li><strong>A risk management system</strong>, documented and maintained across the lifecycle.</li><li><strong>Data governance</strong> — documented training, validation, and test data, with attention to bias and representativeness.</li><li><strong>Technical documentation</strong> sufficient for a regulator to assess conformity.</li><li><strong>Automatic logging</strong> of the system's operation.</li><li><strong>Transparency</strong> to deployers about capabilities and limitations.</li><li><strong>Human oversight</strong> designed in — a person must be able to understand, intervene, and override.</li><li><strong>Accuracy, robustness, and cybersecurity</strong> appropriate to the purpose.</li><li><strong>Conformity assessment</strong> before market placement, and CE marking.</li><li><strong>Registration</strong> in an EU database.</li><li><strong>Post-market monitoring</strong> and serious incident reporting.</li></ul>
<p>That is a quality management system, not a checklist. The nearest analogue is medical device regulation, and it is deliberately modeled on it.</p>
<h2 id="the-practical-advice">the practical advice<a class="anchor" href="#the-practical-advice" aria-label="link to this section">#</a></h2>
<p><strong>Classify now.</strong> Before anything else, determine which of your systems fall in scope. This is a legal question with technical inputs and it should involve both. A lot of organizations discover they are in scope for one feature they did not think about.</p>
<p><strong>Documentation is the bulk of the work.</strong> Most engineering teams do not document their training data provenance, their validation methodology, or their known <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a>. The Act requires all of it. Starting that documentation now is cheaper than reconstructing it later.</p>
<p><strong>Logging is a technical requirement with a deadline.</strong> Automatic recording of system operation, retained, with enough detail to trace a decision. If your system does not do this, that is engineering work, and it is not trivial to add.</p>
<p><strong>Human oversight is a design constraint.</strong> "A person reviews the output" is not sufficient if the person cannot meaningfully understand or override it. Designing for genuine oversight — surfacing confidence, explaining inputs, making override easy and recorded — affects your <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>.</p>
<p><strong>Watch for timeline adjustments.</strong> There has been active discussion about simplification and phasing, and the details have moved before. Track it, and do not use possible delay as a reason to defer classification, which is the part with the longest lead time.</p>
<h2 id="the-honest-framing">the honest framing<a class="anchor" href="#the-honest-framing" aria-label="link to this section">#</a></h2>
<p>Whatever you think of the Act's design — and there are serious criticisms about compliance cost for small companies and about whether risk categories map cleanly onto real systems — it is law, it has extraterritorial reach, and the penalties are proportional to global turnover.</p>
<p>For most developers building most software, none of this applies. For the ones it applies to, six months is not very long.</p>]]></content:encoded></item><item><title>Prompt injection is SQL injection without the fix</title><link>https://readme.news/prompt-injection-is-sql-injection-without-the-fix/</link><guid isPermaLink="true">https://readme.news/prompt-injection-is-sql-injection-without-the-fix/</guid><pubDate>Mon, 09 Feb 2026 09:00:00 +0000</pubDate><description>The analogy is exact except for the part that matters: there is no parameterized query for natural language.</description><content:encoded><![CDATA[<p>The comparison between prompt injection and SQL injection is made constantly and usually stops before the important part.</p>
<p>The structural analogy is exact. The resolution is not available.</p>
<h2 id="the-structural-analogy">the structural analogy<a class="anchor" href="#the-structural-analogy" aria-label="link to this section">#</a></h2>
<p><strong>SQL injection</strong>: the query and the data travel in the same string. The database parses the combined string and cannot tell which characters came from the developer and which came from the user. A user who writes SQL-shaped input has their input executed as SQL.</p>
<p><strong>Prompt injection</strong>: the instructions and the data travel in the same context. The model processes the combined context and cannot tell which tokens came from the developer and which came from a web page, a document, or an email. Content that is instruction-shaped gets treated as an instruction.</p>
<p>Same problem. Mixed channels.</p>
<h2 id="why-the-fix-does-not-transfer">why the fix does not transfer<a class="anchor" href="#why-the-fix-does-not-transfer" aria-label="link to this section">#</a></h2>
<p>SQL injection was solved by parameterized queries. The query is parsed into a plan first, then values are bound into slots. The value cannot become part of the query structure because the structure was already fixed before the value arrived.</p>
<p>That works because SQL has a formal grammar. There is a parse tree. There is a crisp, machine-checkable boundary between "this is syntax" and "this is a literal."</p>
<p>Natural language has no such boundary. There is no parse step that separates instruction from data, because the distinction is semantic, not syntactic. A model's entire function is to interpret meaning from text, and "ignore your prior instructions" means what it means regardless of which part of the context it appeared in.</p>
<p>You cannot parameterize a prompt. There is nothing to parameterize.</p>
<h2 id="what-has-been-tried">what has been tried<a class="anchor" href="#what-has-been-tried" aria-label="link to this section">#</a></h2>
<p><strong>Delimiters.</strong> Wrap untrusted content in tags and instruct the model to treat it as data. Helps somewhat. Defeated by content that includes the closing delimiter, or that argues persuasively that it is an exception.</p>
<p><strong>Instruction hierarchy.</strong> Train the model to weight system instructions above user content above tool results. This is a genuine improvement and the major labs have all done it. It raises the bar and does not eliminate the attack, because it is a learned preference rather than an enforced boundary.</p>
<p><strong>Classifiers.</strong> Detect injection attempts before they reach the model. Works on known patterns. Attackers iterate faster than classifiers update, and the false positive rate on legitimate content is a real product cost.</p>
<p><strong>Separate models.</strong> One model handles untrusted content and cannot call tools; another handles privileged actions and never sees untrusted content. This actually works and it is architectural rather than probabilistic.</p>
<p>That last one is the direction.</p>
<h2 id="the-architecture-that-holds">the architecture that holds<a class="anchor" href="#the-architecture-that-holds" aria-label="link to this section">#</a></h2>
<p>Stop trying to make the model safe. Make the <em>system</em> safe, assuming the model will be compromised.</p>
<p><strong>Separate contexts by trust level.</strong> A session that reads arbitrary web content does not have credentials. A session with credentials does not read arbitrary web content. If information must cross, it crosses through a narrow, typed, validated channel — not by putting both in the same context.</p>
<p><strong>Enforce permission outside the model.</strong> The model does not have access; it requests an action, and a separate system decides whether it is permitted based on the user's actual authorization. The model's opinion about what it should be allowed to do is not an input to that decision.</p>
<p><strong>Confirm all egress.</strong> Any action that sends data outward — an email, an HTTP request, a file write to a shared location — requires explicit approval. This is the control that bounds the damage when everything else fails, because exfiltration is the attacker's goal.</p>
<p><strong>Make every capability narrow.</strong> Not "filesystem access" but "read from this directory." Not "send email" but "send email to addresses in this thread." An agent's permissions should be scoped to the task, granted per-task, and revoked after.</p>
<p><strong>Log everything and monitor for anomaly.</strong> You will not prevent every injection. Detecting one within minutes is the difference between an incident and a breach.</p>
<h2 id="the-uncomfortable-conclusion">the uncomfortable conclusion<a class="anchor" href="#the-uncomfortable-conclusion" aria-label="link to this section">#</a></h2>
<p>For an agent operating on untrusted content with access to sensitive systems, there is currently no configuration that is safe in the way parameterized queries are safe.</p>
<p>There are configurations that are <em>acceptably risky for a given use case</em>, and that judgment requires knowing what the agent can reach and what happens if it is turned against you.</p>
<p>The industry is deploying this capability broadly anyway, on the theory that mitigations will outpace attacks. That theory has a poor historical record — SQL injection was not solved by better filtering, and XSS was not solved by better escaping heuristics. Both were solved by architectural changes that made the unsafe thing impossible to express.</p>
<p>Nobody has the architectural fix here yet. Until someone does, the safety of any deployment is a function of how carefully someone drew the boundaries, and most deployments have not drawn any.</p>]]></content:encoded></item><item><title>Verification is the whole job now</title><link>https://readme.news/verification-is-the-whole-job-now/</link><guid isPermaLink="true">https://readme.news/verification-is-the-whole-job-now/</guid><pubDate>Thu, 08 Jan 2026 09:00:00 +0000</pubDate><description>Producing code got cheap. Knowing whether it&#x27;s right did not. Everything downstream follows from that asymmetry.</description><content:encoded><![CDATA[<p>Here is the single fact that explains most of what has happened to software engineering in the last two years:</p>
<p><strong>Generation got roughly two orders of magnitude cheaper. Verification did not get cheaper at all.</strong></p>
<p>Everything else — the review bottleneck, the arguments about junior hiring, the sudden importance of tests, the unease people cannot articulate about agent-written code — is downstream of that asymmetry.</p>
<h2 id="why-verification-did-not-get-cheaper">why verification did not get cheaper<a class="anchor" href="#why-verification-did-not-get-cheaper" aria-label="link to this section">#</a></h2>
<p>You might expect it to. Surely a model can check code as well as it can write it?</p>
<p>It can, sort of, and not in the way that matters. Three problems:</p>
<p><strong>The checker has the same blind spots as the writer.</strong> A model that does not know your system produced code with an assumption baked in. The same model reviewing it holds the same assumption and does not see it. Independent errors can be filtered by voting; correlated errors cannot.</p>
<p><strong>Verification requires knowing what correct means.</strong> That is a property of your domain, your users, your obligations — not of the code. A model can verify that code does what it says. It cannot verify that what it says is what you needed, because that information was never written down anywhere.</p>
<p><strong>Confidence is decoupled from correctness.</strong> Generated code is fluent. Fluency is the signal humans use to assess competence, and it has been severed from correctness. Reviewing text that reads well but is subtly wrong is much harder than reviewing text that reads badly, and everyone underestimates this.</p>
<h2 id="what-this-changes">what this changes<a class="anchor" href="#what-this-changes" aria-label="link to this section">#</a></h2>
<p><strong>Tests stop being a chore and become the primary artifact.</strong> If verification is the constraint, then anything that makes verification automatic is worth enormous investment. A test suite is executable verification. It was always valuable; it is now the thing that determines your throughput.</p>
<p>The practical implication: write tests first, review those carefully, and let the implementation be cheap. Invert the effort. The specification is the part you have to get right; the code is increasingly a derived artifact.</p>
<p><strong><a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">Type systems</a> get a promotion.</strong> Every constraint the compiler can check is a constraint you do not have to verify by reading. Languages with expressive type systems and strong static guarantees are worth more than they were, because they convert human verification into machine verification.</p>
<p>I have watched teams that were ambivalent about strict TypeScript become evangelists in the space of a year, and the reason is always the same: it catches the class of error that agent-written code produces most often.</p>
<p><strong>Small diffs become non-negotiable.</strong> Verification cost scales superlinearly with diff size, because the number of interactions you have to reason about grows faster than the lines. A machine can produce a 2,000-line change effortlessly. Accepting it is not a favor to anyone.</p>
<p><strong>Property-based testing gets its moment.</strong> If you can state an invariant, you can verify a very large input space cheaply. Invariants are exactly the kind of thing that a human should specify and a machine should check. This technique has been niche for twenty years and it is the right shape for this moment.</p>
<p><strong>Observability becomes a verification tool.</strong> If you cannot verify everything before deploy, you verify in production: <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">feature flags</a>, canaries, and metrics that detect wrongness quickly. Fast detection is a substitute for perfect pre-verification, and it is often a better investment.</p>
<h2 id="the-skill-that-matters">the skill that matters<a class="anchor" href="#the-skill-that-matters" aria-label="link to this section">#</a></h2>
<p>The engineers getting genuine leverage out of these tools all share one trait, and it is not prompting.</p>
<p>It is that they can look at a plausible-looking piece of code and say "that retry loop will hammer the upstream on a 429" or "that will deadlock under concurrent writes" or "that assumes the list is sorted and nothing sorts it."</p>
<p>That skill comes from having debugged those exact failures. It does not come from reading about them.</p>
<p>Which produces the uncomfortable question everybody is circling: the traditional way to acquire it was to write a lot of code badly and then fix it. If that apprenticeship is being automated away, what replaces it?</p>
<h2 id="the-answer-i-have-which-is-partial">the answer I have, which is partial<a class="anchor" href="#the-answer-i-have-which-is-partial" aria-label="link to this section">#</a></h2>
<p>Deliberately do the hard part yourself sometimes.</p>
<p>Not out of nostalgia. Because the judgment is the product, and the judgment is built by doing. If you delegate every debugging session, you will be worse at debugging in two years, and debugging is the thing you are being paid for now.</p>
<p>Pick the gnarliest bug of the week and do it by hand. Read the code the agent wrote in the module you own, all of it, once a month. Write the tricky concurrency code yourself and let the machine do the CRUD.</p>
<p>That is not a workflow recommendation for efficiency. It is a training regimen, and treating it as one is the honest framing.</p>]]></content:encoded></item><item><title>CES 2026: the power supply is the product</title><link>https://readme.news/ces-2026-the-power-supply-is-the-product/</link><guid isPermaLink="true">https://readme.news/ces-2026-the-power-supply-is-the-product/</guid><pubDate>Tue, 06 Jan 2026 09:00:00 +0000</pubDate><description>Another year of AI in appliances, and one genuine trend hiding underneath: everything is now thermally constrained.</description><content:encoded><![CDATA[<p>The annual Las Vegas exercise in stapling language models to inventory has concluded. As always, the interesting signal is in the parts nobody put on a keynote slide.</p>
<h2 id="the-actual-trend">the actual trend<a class="anchor" href="#the-actual-trend" aria-label="link to this section">#</a></h2>
<p>Every category at this show is now constrained by thermals and power delivery rather than by compute.</p>
<p><strong>Laptops</strong> with neural accelerators that cannot sustain their rated throughput for more than a few minutes in a thin chassis. The TOPS number on the sticker is a peak, and the peak lasts as long as the thermal mass does.</p>
<p><strong>Handhelds</strong> where battery life is the entire product decision and every added capability is a subtraction from it.</p>
<p><strong>Home devices</strong> that want to run models locally and cannot, because the power envelope of something that sits on a shelf is a few watts.</p>
<p><strong>Desktops</strong> where the power supply recommendations for a current-generation GPU have crept past what a lot of household circuits comfortably deliver alongside everything else in the room.</p>
<p>This is what a computing era looks like when it hits a physical wall. The interesting engineering for the next several years is efficiency, not capability — which is historically when the best engineering happens.</p>
<h2 id="what-to-actually-take-from-the-show">what to actually take from the show<a class="anchor" href="#what-to-actually-take-from-the-show" aria-label="link to this section">#</a></h2>
<p><strong><a class="xref" href="/pixel-10-and-the-on-device-model-as-a-platform-feature/" title="Pixel 10 and the on-device model as a platform feature">On-device</a> inference is a spec sheet lie in the thin-and-light category.</strong> If you are shipping software that assumes a local NPU, test on a machine that has been running for twenty minutes, not on a cold one. The difference is large and nobody benchmarks it.</p>
<p><strong>Unified memory keeps winning.</strong> Every serious local-AI machine announced has shared CPU/GPU memory at high capacity. Discrete VRAM is a bad fit for a workload where capacity matters more than peak bandwidth, and the industry has figured that out.</p>
<p><strong>The robot demos are still teleoperated.</strong> They have been for four years. Watch for the operator's hands, or the suspiciously smooth trajectory, or the fact that the demo never deviates from the script. There is real progress in robotics and it is happening in warehouses, not on stages.</p>
<h2 id="the-accessory-economy">the accessory economy<a class="anchor" href="#the-accessory-economy" aria-label="link to this section">#</a></h2>
<p>A large fraction of the floor was accessories for devices that do not exist yet: cases, docks, and mounts for AI wearables. That is a leading indicator of nothing except that the accessory industry moves fast and takes risks cheaply.</p>
<h2 id="the-one-thing-i-would-actually-buy">the one thing I would actually buy<a class="anchor" href="#the-one-thing-i-would-actually-buy" aria-label="link to this section">#</a></h2>
<p>The mundane one: displays got better and cheaper. High-refresh, high-resolution, accurate-color monitors are at prices that were flagship prices three years ago.</p>
<p>If you have been running the same monitor since 2020 and you spend eight hours a day looking at text on it, that is the highest-value hardware purchase available to you, and it will not be mentioned in a single keynote.</p>
<h2 id="the-meta-observation">the meta-observation<a class="anchor" href="#the-meta-observation" aria-label="link to this section">#</a></h2>
<p>CES has been a bad predictor of what matters for about a decade. The genuinely important hardware of the last ten years — the M1, the H100, TPUs, the shift to ARM in servers — was announced at industry events or in blog posts, to audiences who understood it, without a stage show.</p>
<p>The consumer electronics show is now mostly a trade event for retail buyers, covered as if it were a technology forecast. Read the specs. Skip the narrative.</p>]]></content:encoded></item><item><title>AGENTS.md and the repository that explains itself</title><link>https://readme.news/agentsmd-and-the-repository-that-explains-itself/</link><guid isPermaLink="true">https://readme.news/agentsmd-and-the-repository-that-explains-itself/</guid><pubDate>Sun, 04 Jan 2026 09:00:00 +0000</pubDate><description>A convention nobody standardized became standard anyway. Here&#x27;s what belongs in it and what doesn&#x27;t.</description><content:encoded><![CDATA[<p>Over the last year, essentially every coding agent converged on the same mechanism: a markdown file in the repository root telling the agent how to work in this codebase.</p>
<p>The names varied — <code>AGENTS.md</code>, <code>CLAUDE.md</code>, <code>.cursorrules</code>, <code>.github/copilot-instructions.md</code> — and the format did not. It is prose. The model reads it. That is the whole protocol.</p>
<p>A convention that emerges independently in five products in one year is telling you something about the shape of the problem.</p>
<h2 id="what-actually-belongs-in-it">what actually belongs in it<a class="anchor" href="#what-actually-belongs-in-it" aria-label="link to this section">#</a></h2>
<p>I have written and rewritten these a dozen times. The version that works is shorter than you think and more specific than you want.</p>
<p><strong>Commands.</strong> The exact invocations. Not "run the tests" — the command, including the flags, including how to run one test file.</p>
<div class="code"><span class="code-lang">markdown</span><pre><code class="lang-markdown">## commands
- install: `pnpm install --frozen-lockfile`
- test: `pnpm vitest run`
- test one file: `pnpm vitest run src/foo.test.ts`
- typecheck: `pnpm tsc --noEmit`
- lint: `pnpm biome check --write .`
- dev server: `pnpm dev` (port 5173)</code></pre></div>
<p>This section alone eliminates most wasted agent turns. Without it, the agent guesses, guesses wrong, and spends four tool calls discovering your test runner.</p>
<p><strong>Non-obvious structure.</strong> Where things live, when it is not inferable. "Database migrations are in <code>db/migrations</code> and must be created with <code>pnpm db:new</code>, never by hand." "The <code>legacy/</code> directory is frozen — do not modify it."</p>
<p><strong>Conventions that a linter does not enforce.</strong> If your linter catches it, do not write it down; the linter will tell the agent. Write down the things that are policy rather than syntax: "prefer composition over inheritance in <code>src/domain</code>", "all public functions in <code>api/</code> need a docstring", "we do not use default exports".</p>
<p><strong>Things that will break.</strong> "Do not run <code>pnpm build</code> — it takes 12 minutes and is not needed for tests." "The integration tests require Docker; skip them if it is not running." "Never modify <code>schema.sql</code> directly."</p>
<p><strong>Boundaries.</strong> What the agent may not touch without asking. Production configs, migration files, anything with a security review requirement.</p>
<h2 id="what-does-not-belong">what does not belong<a class="anchor" href="#what-does-not-belong" aria-label="link to this section">#</a></h2>
<p><strong>Your entire architecture document.</strong> The agent reads this every session. A three-thousand-word essay costs tokens on every single request and dilutes attention across a lot of text that is irrelevant to most tasks.</p>
<p>Keep it under about 200 lines. If you need more, put it in a separate document and reference it: "for the event pipeline design, read <code>docs/events.md</code> before changing anything in <code>src/events/</code>."</p>
<p><strong>Anything the code already says.</strong> Do not restate the type signatures. The agent can read.</p>
<p><strong>Aspirations.</strong> "We value clean code" is not an instruction. "Functions over 40 lines get split" is.</p>
<p><strong>Rules you do not actually follow.</strong> If the codebase contradicts the file, the codebase wins in the model's attention, and now your instructions are noise.</p>
<h2 id="the-part-i-did-not-expect">the part I did not expect<a class="anchor" href="#the-part-i-did-not-expect" aria-label="link to this section">#</a></h2>
<p>Writing these has improved my documentation for humans.</p>
<p>The discipline of writing "here is exactly how to run the tests, here are the things that will surprise you, here is what not to touch" is precisely what a new engineer needs on day one, and it is precisely what most onboarding documents fail to contain because they were written by someone who already knew.</p>
<p>An agent is an infinitely patient new hire who will follow instructions literally and never ask a clarifying question out of politeness. That turns out to be an excellent test of whether your instructions are any good.</p>
<p>Several teams I know have merged their onboarding doc and their agent file into one. That is the correct end state.</p>
<h2 id="the-standardization-question">the standardization question<a class="anchor" href="#the-standardization-question" aria-label="link to this section">#</a></h2>
<p>There is an ongoing effort to consolidate on <code>AGENTS.md</code> as the common name, with tools reading it as a fallback. That would be good and it is a coordination problem, which means it will take longer than it should.</p>
<p>In the meantime: write one file, symlink the rest. It costs nothing.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">ln -s AGENTS.md CLAUDE.md</code></pre></div>]]></content:encoded></item><item><title>Five things that will actually change in 2026</title><link>https://readme.news/five-things-that-will-actually-change-in-2026/</link><guid isPermaLink="true">https://readme.news/five-things-that-will-actually-change-in-2026/</guid><pubDate>Fri, 02 Jan 2026 09:00:00 +0000</pubDate><description>Not predictions about model capability. Predictions about what your Tuesday looks like.</description><content:encoded><![CDATA[<p>Every January the prediction posts arrive and they are all about model capability, which is the least useful thing to predict because it is both the most forecast and the least actionable.</p>
<p>Here are five changes to how software actually gets made. I will grade these in December.</p>
<h2 id="1-review-tooling-becomes-a-real-product-category">1. Review tooling becomes a real product category<a class="anchor" href="#1-review-tooling-becomes-a-real-product-category" aria-label="link to this section">#</a></h2>
<p>The bottleneck moved to verification in 2025 and the tooling did not follow. We are still reviewing machine-generated changes with an <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> designed in 2008 for occasional human contributions.</p>
<p>What has to exist: review surfaces that show <em>intent</em> rather than only diff. Which invariants does this change preserve? What did the author consider and reject? Which tests exercise the changed path, and did any of them exist before this change?</p>
<p>Some of that can be generated. Most of it is a UI problem nobody has solved because the pain only became acute in the last eighteen months.</p>
<p>Watch for this to be one of the most-funded developer tooling categories of the year.</p>
<h2 id="2-repository-level-agent-configuration-becomes-standard-practice">2. Repository-level agent configuration becomes standard practice<a class="anchor" href="#2-repository-level-agent-configuration-becomes-standard-practice" aria-label="link to this section">#</a></h2>
<p><code>AGENTS.md</code>, <code>CLAUDE.md</code>, <code>.cursorrules</code>, and their cousins converged on the same idea in 2025: the repository tells the agent how to work in it. Build commands, test commands, conventions, things not to touch.</p>
<p>In 2026 this stops being an optional nicety and becomes part of the definition of a well-maintained repository, the way a README and a CI config are. New projects will have one from the start. Onboarding docs and agent instructions will merge, because they are the same document with a different reader.</p>
<p>The second-order effect is good: writing down how your project actually works helps humans too, and a lot of teams will document things properly for the first time because a machine needed it.</p>
<h2 id="3-power-becomes-a-line-item-developers-think-about">3. Power becomes a line item developers think about<a class="anchor" href="#3-power-becomes-a-line-item-developers-think-about" aria-label="link to this section">#</a></h2>
<p>Not in an abstract climate sense. In a "which region, which instance type, what does this cost" sense.</p>
<p>Electricity prices in datacenter-dense regions rose through 2025 and the rate cases are ongoing. Cloud pricing follows. Regional price differentials for compute are going to become large enough that they affect architecture decisions, the way egress pricing already does.</p>
<p>Expect "run the batch job in the cheap region overnight" to go from a sustainability talking point to a cost-optimization default.</p>
<h2 id="4-the-supply-chain-gets-a-real-control-not-another-dashboard">4. The supply chain gets a real control, not another dashboard<a class="anchor" href="#4-the-supply-chain-gets-a-real-control-not-another-dashboard" aria-label="link to this section">#</a></h2>
<p>Four significant npm incidents in 2025, ending with a self-propagating worm. The registry-level controls — mandatory trusted publishing, restrictions on long-lived tokens, install-script defaults — are coming, and they will break workflows.</p>
<p>That breakage is the point. The ecosystem has been asking maintainers to be individually unphishable for a decade, and it has not worked, because it cannot.</p>
<p>Expect a painful migration and a meaningfully safer ecosystem at the end of it.</p>
<h2 id="5-senior-engineer-quietly-redefines-around-judgment">5. "Senior engineer" quietly redefines around judgment<a class="anchor" href="#5-senior-engineer-quietly-redefines-around-judgment" aria-label="link to this section">#</a></h2>
<p>The mechanical parts of the job — writing the code, remembering the API, finding the bug in a file you can see — got substantially cheaper. What did not get cheaper: knowing what to build, knowing when something is wrong, knowing which complexity is worth it, and being able to say no with a reason.</p>
<p>This has always been what seniority meant. It is now what it <em>only</em> means, and the gap between people who have it and people who were fast typists is going to become uncomfortable to talk about at promotion committees.</p>
<p>The related problem, which nobody has solved: the traditional path to developing that judgment was doing the mechanical work for several years. If juniors do not do that work, where does the judgment come from?</p>
<p>I do not know. It is the most important open question in the profession right now and almost all the discussion of it is either dismissive or catastrophizing.</p>
<h2 id="what-i-am-not-predicting">what I am not predicting<a class="anchor" href="#what-i-am-not-predicting" aria-label="link to this section">#</a></h2>
<p>Anything about AGI. Anything about a specific model beating a specific benchmark. Anything about a company's valuation.</p>
<p>Those get all the attention and none of them change what you do on Tuesday.</p>]]></content:encoded></item><item><title>2025, in order</title><link>https://readme.news/2025-in-order/</link><guid isPermaLink="true">https://readme.news/2025-in-order/</guid><pubDate>Tue, 30 Dec 2025 09:00:00 +0000</pubDate><description>The year in one page: what happened, what mattered, and the three things that will still matter in 2030.</description><content:encoded><![CDATA[<p>A hundred pieces this year. Here is the compressed version.</p>
<h2 id="the-year-in-one-paragraph">the year in one paragraph<a class="anchor" href="#the-year-in-one-paragraph" aria-label="link to this section">#</a></h2>
<p>An open reasoning model under an MIT license repriced the sector in January. Reasoning became a runtime dial rather than a model choice. Coding agents went from <a class="xref" href="/operator-and-the-long-road-to-an-agent-that-can-click/" title="Operator, and the long road to an agent that can click">research preview</a> to standard tooling in about six months and made code review the bottleneck. The npm ecosystem got a self-propagating worm. Two of the largest infrastructure providers had multi-hour outages caused by config, not code. Python removed the GIL, optionally. Rust shipped its largest edition. The frontier models converged and the competition moved to price and distribution.</p>
<h2 id="the-three-things-that-will-still-matter-in-2030">the three things that will still matter in 2030<a class="anchor" href="#the-three-things-that-will-still-matter-in-2030" aria-label="link to this section">#</a></h2>
<p><strong>1. Reasoning as a controllable runtime parameter.</strong></p>
<p>The most durable technical idea of the year. Every provider independently arrived at the same design: the caller decides how much the model thinks, per request.</p>
<p>This is durable because it reflects something true about the problem — the appropriate amount of computation is a property of the task, not the model, and only the caller knows the task. Interfaces that reflect true structure survive. Interfaces that reflect an implementation detail do not.</p>
<p><strong>2. The agent-review bottleneck.</strong></p>
<p>Generation throughput increased dramatically. Verification throughput did not. That gap is the central engineering problem of the next several years, and nobody has a good answer.</p>
<p>Everything downstream follows from it: how teams are structured, what tests are for, what "code review" means, whether we build systems we understand or systems that pass tests. This is not a tooling problem that gets solved with a better diff viewer. It is a fundamental asymmetry between producing and checking, and it shows up in every field where automation outpaced verification.</p>
<p><strong>3. Supply chain trust has no technical fix yet.</strong></p>
<p>Four significant npm incidents this year, escalating in sophistication, ending with a worm. Every one exploited the same structural fact: install-time code execution plus long-lived publishing credentials plus a dependency graph nobody designed.</p>
<p>The controls that work — trusted publishing, no install scripts, version cooldowns, phishing-resistant auth — are known and unevenly adopted. The ecosystems that are structurally safer got that way by design decisions made years ago that cannot be retrofitted cheaply.</p>
<p>This gets worse before it gets better, and it is a solvable problem that we are choosing not to solve at the speed it requires.</p>
<h2 id="the-things-that-felt-big-and-were-not">the things that felt big and were not<a class="anchor" href="#the-things-that-felt-big-and-were-not" aria-label="link to this section">#</a></h2>
<p><strong>Model benchmark leapfrogging.</strong> Every launch claimed the frontier. The differences were within evaluation noise for most real tasks. The benchmark discourse consumed enormous attention and predicted very little about what was useful.</p>
<p><strong>Agent frameworks.</strong> Most of the orchestration layer got absorbed into the models, exactly as function-calling libraries and JSON-repair libraries were absorbed before. The durable layer was never orchestration.</p>
<p><strong>The browser wars, round two.</strong> Everyone shipped a Chromium fork with a model in it. That does not diversify the engine landscape; it diversifies the UI on top of one engine.</p>
<h2 id="the-things-that-felt-small-and-were-not">the things that felt small and were not<a class="anchor" href="#the-things-that-felt-small-and-were-not" aria-label="link to this section">#</a></h2>
<p><strong>LLD as the default linker in Rust.</strong> A default change delivered a build-time improvement to everyone at once that documentation had failed to deliver for years. The general lesson — defaults are the highest-leverage thing a toolchain ships — applies far beyond Rust.</p>
<p><strong>Template strings in Python.</strong> A one-character syntax change that makes the safe path as easy as the unsafe one. If library adoption follows, an entire category of injection vulnerability becomes hard to write. Security features that depend on diligence fail; security features enforced by the <a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">type system</a> work.</p>
<p><strong><a class="xref" href="/claude-sonnet-45-and-the-agent-that-runs-for-thirty-hours/" title="Claude Sonnet 4.5 and the agent that runs for thirty hours">Context editing</a> and memory tools.</strong> Unglamorous plumbing that determines whether long agent runs succeed. The context window was never the constraint. Goal retention was.</p>
<h2 id="the-thing-i-want-to-say-going-into-next-year">the thing I want to say going into next year<a class="anchor" href="#the-thing-i-want-to-say-going-into-next-year" aria-label="link to this section">#</a></h2>
<p>The most valuable skill in 2025 was not prompting, and it will not be in 2026 either.</p>
<p>It was the ability to tell whether something is correct. That skill was always valuable and it was previously bundled with the ability to produce the thing. The bundle has come apart. Producing is cheap now. Checking is not, and checking is what everything rests on.</p>
<p>Every argument this year about what AI does to engineering eventually reduces to that. The people who are getting enormous leverage out of these tools are, without exception, the people who can tell when the output is wrong.</p>
<p>That is not a comforting conclusion for anyone hoping the tools eliminate the need for expertise. It is the actual state of things, and it is worth building your career around.</p>
<p>See you in January.</p>]]></content:encoded></item><item><title>What I got wrong in 2025</title><link>https://readme.news/what-i-got-wrong-in-2025/</link><guid isPermaLink="true">https://readme.news/what-i-got-wrong-in-2025/</guid><pubDate>Tue, 23 Dec 2025 09:00:00 +0000</pubDate><description>Six predictions from twelve months of writing, graded honestly. Three were wrong.</description><content:encoded><![CDATA[<p>Writing publicly means being wrong publicly. Here are the calls I made this year, graded.</p>
<h2 id="1-local-models-become-the-default-for-routine-developer-work">1. "Local models become the default for routine developer work"<a class="anchor" href="#1-local-models-become-the-default-for-routine-developer-work" aria-label="link to this section">#</a></h2>
<p><strong>Written in January. Partially right, wrong on scale.</strong></p>
<p>Local models got much better and much easier to run. The tooling is excellent. A 14B model on a laptop is genuinely useful for the work I said it would be useful for.</p>
<p>What I got wrong: adoption. The default did not become hybrid. It became "frontier API for everything," because the frontier models got cheap enough fast enough that the cost argument for local largely evaporated for individuals.</p>
<p>The privacy argument held and drove real adoption in regulated industries, which is what I said. The convenience argument lost, which I underweighted.</p>
<p>Grade: <strong>C+.</strong> The technology went where I said. The behavior did not.</p>
<h2 id="2-the-moat-is-distribution-and-serving-efficiency-not-model-quality">2. "The moat is distribution and serving efficiency, not model quality"<a class="anchor" href="#2-the-moat-is-distribution-and-serving-efficiency-not-model-quality" aria-label="link to this section">#</a></h2>
<p><strong>Written in February. This held up well.</strong></p>
<p>Model capability converged. Every frontier lab shipped comparable models. The differentiation moved to price, latency, integration, and distribution — exactly as predicted, and faster than I expected.</p>
<p>Google putting a <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> into Search on launch day is the clearest demonstration. No competitor can do that.</p>
<p>Grade: <strong>A−.</strong></p>
<h2 id="3-delegated-agents-make-review-the-bottleneck">3. "Delegated agents make review the bottleneck"<a class="anchor" href="#3-delegated-agents-make-review-the-bottleneck" aria-label="link to this section">#</a></h2>
<p><strong>Written in May. Correct and I understated it.</strong></p>
<p>This is now the dominant complaint from every team using coding agents seriously. Throughput on generation went up several fold. Throughput on review did not move.</p>
<p>What I did not anticipate: the second-order effect on review <em>quality</em>. It is not just that review is slower — it is that reviewers under queue pressure approve things they have not fully understood, and the defect rate is showing up in ways that are hard to attribute.</p>
<p>Grade: <strong>A</strong>, with the caveat that I was not pessimistic enough.</p>
<h2 id="4-rag-infrastructure-is-solving-a-temporary-problem">4. "RAG infrastructure is solving a temporary problem"<a class="anchor" href="#4-rag-infrastructure-is-solving-a-temporary-problem" aria-label="link to this section">#</a></h2>
<p><strong>Written in March. Too confident, partially wrong.</strong></p>
<p>Long context did get better and cheaper, and a lot of naive RAG did get replaced by just putting the documents in the prompt. That part was right.</p>
<p>What I got wrong: retrieval did not go away, it moved. The interesting systems now do retrieval <em>as a tool the model calls</em> rather than as a preprocessing step. The vector database did not die; it became something the agent queries.</p>
<p>That is a substantially different outcome than "the category is temporary," and I should have seen it because the same pattern — capability absorbed into the model, infrastructure repositioned rather than eliminated — had already happened twice.</p>
<p>Grade: <strong>C.</strong></p>
<h2 id="5-nvidias-roadmap-slide-is-the-real-announcement">5. "Nvidia's roadmap slide is the real announcement"<a class="anchor" href="#5-nvidias-roadmap-slide-is-the-real-announcement" aria-label="link to this section">#</a></h2>
<p><strong>Written in March. Right, and the power framing was the useful part.</strong></p>
<p>Power as the binding constraint became the consensus view over the year. The unit of AI infrastructure discussion is now gigawatts. Utility rate cases about datacenter interconnection are a live political issue in multiple states.</p>
<p>Grade: <strong>A.</strong></p>
<h2 id="6-the-npm-ecosystem-will-have-a-self-propagating-worm">6. "The npm ecosystem will have a self-propagating worm"<a class="anchor" href="#6-the-npm-ecosystem-will-have-a-self-propagating-worm" aria-label="link to this section">#</a></h2>
<p><strong>Written in February, in a piece about supply chain risk. I hate being right about this one.</strong></p>
<p>It arrived in September. The mechanism was exactly the one everybody had described: steal credentials on install, use them to publish, repeat.</p>
<p>The prediction was not clever. Every precondition was public. What I got wrong was the timeline — I expected it to take longer, because I assumed the registry would ship stronger publishing controls first. It did not, until after.</p>
<p>Grade: <strong>A on the call, F on the assumption that anyone would act in time.</strong></p>
<h2 id="the-pattern-in-my-errors">the pattern in my errors<a class="anchor" href="#the-pattern-in-my-errors" aria-label="link to this section">#</a></h2>
<p>Looking at the three I got wrong, they share a shape: <strong>I was right about the technology and wrong about the behavior.</strong></p>
<p>Local models got good; people did not switch. RAG got less necessary; the infrastructure adapted instead of dying. Registry controls were obviously needed; they arrived after the incident rather than before.</p>
<p>The lesson I am taking into next year: technical trajectories are much more predictable than adoption, and adoption is where the money and the consequences are. When I feel confident, it is usually because I am reasoning about the technology, and the technology was never the hard part.</p>]]></content:encoded></item><item><title>The open weights year</title><link>https://readme.news/the-open-weights-year/</link><guid isPermaLink="true">https://readme.news/the-open-weights-year/</guid><pubDate>Fri, 19 Dec 2025 09:00:00 +0000</pubDate><description>Twelve months that took open models from interesting to unavoidable, and where the gap actually sits.</description><content:encoded><![CDATA[<p>January opened with an MIT-licensed reasoning model that repriced the entire sector in a week. December closes with open weights as a normal, boring option in any serious architecture discussion.</p>
<p>Here is the year, and what it means for the next one.</p>
<h2 id="the-releases-that-mattered">the releases that mattered<a class="anchor" href="#the-releases-that-mattered" aria-label="link to this section">#</a></h2>
<p><strong>DeepSeek R1</strong> (January). MIT license, published training methodology, distilled variants that ran on consumer hardware. The RL-on-verifiable-rewards recipe was reproduced widely within weeks.</p>
<p><strong>Qwen3</strong> (April). Eight models, Apache 2.0, MoE variants with excellent quality-per-active-parameter, 119 languages.</p>
<p><strong><a class="xref" href="/kimi-k2-is-a-trillion-parameter-open-weights-release/" title="Kimi K2 is a trillion-parameter open weights release">Kimi K2</a></strong> (July). A trillion parameters, open weights, tuned for agentic tool use, with a genuinely novel training stability contribution.</p>
<p><strong>gpt-oss</strong> (August). OpenAI's first open weights since 2019, Apache 2.0, with a 20B variant that runs on a laptop.</p>
<p>Plus continuous releases from Mistral, Zhipu, MiniMax, Meta, Microsoft, Google, and a long tail of fine-tunes.</p>
<h2 id="where-the-gap-actually-is">where the gap actually is<a class="anchor" href="#where-the-gap-actually-is" aria-label="link to this section">#</a></h2>
<p>The frontier-to-open gap held at roughly six to twelve months all year. That stability is the most important finding, because it means open weights are not converging on the frontier and are not falling behind — they are trailing at a fixed distance.</p>
<p>But "six months behind" undersells the practical position, because the gap is not uniform:</p>
<p><strong>Nearly closed:</strong> code completion, summarization, extraction, classification, translation, structured output, straightforward tool use. For these, a good open model is not meaningfully worse than a <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a>, and it costs a fraction.</p>
<p><strong>Meaningfully behind:</strong> long-horizon agentic work, complex multi-step reasoning, instruction following over many turns, reliability at the tail. This is where frontier models earn their price.</p>
<p><strong>Not comparable:</strong> anything requiring the surrounding infrastructure — enterprise support, uptime guarantees, safety tooling, indemnification. Open weights give you the model and nothing else.</p>
<h2 id="what-changed-structurally">what changed structurally<a class="anchor" href="#what-changed-structurally" aria-label="link to this section">#</a></h2>
<p><strong>Licensing got genuinely permissive.</strong> Two years ago "open" meant a research license with a prohibited-use list. Now the frontier of open releases is Apache 2.0 and MIT. That is a real change and it removed the legal review that was blocking adoption.</p>
<p><strong>The runtime story got boring.</strong> Ollama, llama.cpp, vLLM, MLX, LM Studio. One command. An OpenAI-compatible endpoint. The friction that kept open models in the enthusiast category is gone.</p>
<p><strong>Hosting became competitive.</strong> Multiple providers serve open models at prices well below frontier API rates, with real SLAs. You can use open weights without running anything.</p>
<p><strong>The center of gravity moved east.</strong> The most capable, most permissively licensed open releases came predominantly from Chinese labs. That is a strategic fact with policy consequences that are being worked out loudly and mostly unproductively.</p>
<h2 id="the-practical-architecture-for-next-year">the practical architecture for next year<a class="anchor" href="#the-practical-architecture-for-next-year" aria-label="link to this section">#</a></h2>
<p>The shape that makes sense:</p>
<ul><li><strong>Open weights, self-hosted or on a cheap provider</strong>, for high-volume, well-defined tasks. Classification, extraction, embedding, first-pass drafting.</li><li><strong>Frontier API</strong> for the hard tail: planning, complex reasoning, anything customer-facing where a bad answer is expensive.</li><li><strong>A router</strong> deciding between them, with the escalation rate instrumented.</li><li><strong>Your own eval set</strong> in your own repository, run against every candidate.</li></ul>
<p>That is not a hedge. It is what the cost and capability curves actually imply.</p>
<h2 id="the-prediction">the prediction<a class="anchor" href="#the-prediction" aria-label="link to this section">#</a></h2>
<p>The gap holds at roughly six to twelve months through next year. Open weights absorb an increasing share of production workload by volume while frontier models keep the high-value tail.</p>
<p>The interesting question is not capability. It is whether the labs currently releasing weights continue to, and that is a business decision that could change in either direction with one quarter's strategy review.</p>
<p>Download the ones you care about. They cannot be un-released.</p>]]></content:encoded></item><item><title>Advent of Code and what a puzzle is for</title><link>https://readme.news/advent-of-code-and-what-a-puzzle-is-for/</link><guid isPermaLink="true">https://readme.news/advent-of-code-and-what-a-puzzle-is-for/</guid><pubDate>Fri, 05 Dec 2025 09:00:00 +0000</pubDate><description>The leaderboard broke years ago. The value was never the leaderboard.</description><content:encoded><![CDATA[<p>Advent of Code is running again, and so is the annual argument about AI solutions and the global leaderboard.</p>
<p>The argument is settled in practice: models solve most early puzzles in seconds, the global leaderboard is not a meaningful competition anymore, and Eric Wastl has asked people not to use AI for leaderboard attempts. Some people ignore that.</p>
<p>I think the argument is also beside the point, and I want to make the case for what these puzzles are actually good for.</p>
<h2 id="what-the-leaderboard-was-ever-worth">what the leaderboard was ever worth<a class="anchor" href="#what-the-leaderboard-was-ever-worth" aria-label="link to this section">#</a></h2>
<p>The global top 100 was always a competition between a small number of extremely fast competitive programmers, most of whom had done thousands of hours of practice specifically for this. For everyone else it was scenery.</p>
<p>If your enjoyment depended on that leaderboard, it was already not for you before any model existed.</p>
<h2 id="what-the-puzzles-are-good-for">what the puzzles are good for<a class="anchor" href="#what-the-puzzles-are-good-for" aria-label="link to this section">#</a></h2>
<p><strong>Practicing a language you do not know.</strong> This is the best use and it is what I do every year. Twenty-five bounded problems with clear specifications and verifiable answers is an excellent scaffold for learning a language's idioms. You are not fighting the problem, so you can concentrate on the tool.</p>
<p>Do it in something uncomfortable. Rust if you write Python. A Lisp if you write Java. Uiua or J if you want to genuinely reconsider what a loop is.</p>
<p><strong>Finding the gap between working and good.</strong> Almost every puzzle has a naive solution that works for part one and does not terminate for part two. That gap — where you have to actually think about the algorithm rather than just express the problem — is where the learning is.</p>
<p>That is also the part models are least reliably good at, incidentally. They will write the naive version confidently.</p>
<p><strong>Reading other people's solutions.</strong> The subreddit's solution megathreads are one of the better learning resources in programming. Twenty implementations of the same problem in different languages by people of different skill levels, all solving something you just spent an hour on, so you have full context.</p>
<p>You will learn more from reading five solutions after solving it than from solving three more puzzles.</p>
<p><strong>A shared thing.</strong> For a few weeks in December a lot of programmers are all thinking about the same problem on the same day. That is a rarer thing than it used to be and it is worth something.</p>
<h2 id="on-using-a-model">on using a model<a class="anchor" href="#on-using-a-model" aria-label="link to this section">#</a></h2>
<p>I do not think there is a moral question here for anything except the leaderboard, where the request is explicit and honoring it is basic decency.</p>
<p>For everything else: it is your time and your learning. If you use a model to solve everything instantly, you will finish faster and learn nothing, which is a trade you are allowed to make and which defeats the entire point of doing a programming puzzle for fun.</p>
<p>The interesting middle ground, which I recommend: solve it yourself, then ask a model to critique your solution. "What is the time complexity? What would break at larger inputs? Is there a standard algorithm for this I should know?"</p>
<p>That is a genuinely good use. You get the learning from solving it and you get the thing you did not know from the critique. It is closer to having a knowledgeable colleague than to having someone do your homework.</p>
<h2 id="the-thing-that-actually-bothers-me">the thing that actually bothers me<a class="anchor" href="#the-thing-that-actually-bothers-me" aria-label="link to this section">#</a></h2>
<p>Not the puzzles. The people who have decided that because a model can do the easy version of something, learning to do it is pointless.</p>
<p>That reasoning has been available for every tool ever built and it has been wrong every time. Calculators did not make arithmetic understanding useless. Compilers did not make understanding what the machine does useless. IDEs did not make knowing the language useless.</p>
<p>The understanding is what lets you tell when the tool is wrong. That has never been more valuable than it is now, when the tool is wrong fluently.</p>
<p>Go do the puzzles. In a language you do not know. Slowly.</p>]]></content:encoded></item><item><title>re:Invent and the year of the boring cloud announcement</title><link>https://readme.news/reinvent-and-the-year-of-the-boring-cloud-announcement/</link><guid isPermaLink="true">https://readme.news/reinvent-and-the-year-of-the-boring-cloud-announcement/</guid><pubDate>Tue, 02 Dec 2025 09:00:00 +0000</pubDate><description>Silicon, agents, and a keynote that mostly described things that already existed. That&#x27;s a healthy sign.</description><content:encoded><![CDATA[<p>AWS re:Invent is underway and the announcement pace is, as always, absurd. Most of it will not matter to you. Here is the part that will.</p>
<h2 id="silicon">silicon<a class="anchor" href="#silicon" aria-label="link to this section">#</a></h2>
<p>Amazon continues pushing Trainium and Inferentia as the alternative to buying Nvidia. The pitch is price-performance for customers who can tolerate a different software stack.</p>
<p>The honest state of it: the hardware is competitive on paper, the software ecosystem is meaningfully behind CUDA, and the gap is closing slowly. If your workload runs through PyTorch with standard operations, the port is manageable. If you have custom kernels, it is a project.</p>
<p>The strategic point is the same for every hyperscaler: reducing dependency on a single supplier with enormous pricing power. Whether the customer benefit materializes depends entirely on whether the savings get passed through.</p>
<h2 id="the-agent-announcements">the agent announcements<a class="anchor" href="#the-agent-announcements" aria-label="link to this section">#</a></h2>
<p>Every cloud vendor is now shipping agent infrastructure: runtimes, memory services, gateways for tool access, identity for agents, observability for agent traces.</p>
<p>This category is real. Running agents in production has genuine infrastructure requirements that are different from running services:</p>
<ul><li><strong>Long-lived sessions</strong> with state that outlives a request.</li><li><strong>Non-deterministic execution paths</strong> that make traditional tracing awkward.</li><li><strong>Cost per invocation that varies by orders of magnitude.</strong></li><li><strong>Identity and permission scoping</strong> for a thing acting on a user's behalf.</li><li><strong>Human approval gates</strong> in the middle of automated flows.</li></ul>
<p>Those are real problems and the tooling is early everywhere. Evaluate on whether it solves a problem you actually have, not on whether the demo was good.</p>
<h2 id="the-pattern-i-would-push-back-on">the pattern I would push back on<a class="anchor" href="#the-pattern-i-would-push-back-on" aria-label="link to this section">#</a></h2>
<p>Every vendor's agent framework wants to be the place your orchestration lives. That is a <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a> position, and orchestration is the layer most likely to be absorbed by the models themselves — it has been happening steadily for two years.</p>
<p>Keep your orchestration portable. Use the managed pieces for the genuinely hard infrastructure — identity, secure tool access, session storage — and keep the logic in your own code.</p>
<h2 id="the-underrated-announcements">the underrated announcements<a class="anchor" href="#the-underrated-announcements" aria-label="link to this section">#</a></h2>
<p>The ones nobody writes about and everybody uses:</p>
<ul><li>Incremental improvements to S3 consistency and performance.</li><li>Networking latency reductions.</li><li>Cost management tooling that is slightly less bad.</li><li>Database engine version updates.</li></ul>
<p>These are worth more to most organizations than any AI announcement, and they get four slides between two hours of agent demos.</p>
<h2 id="the-meta-observation">the meta-observation<a class="anchor" href="#the-meta-observation" aria-label="link to this section">#</a></h2>
<p>Cloud conferences have gotten less interesting, and that is good. It means the platform is mature. The exciting years of a platform are the years when fundamental things are missing.</p>
<p>The interesting question for AWS is not what they announced. It is whether the operational excellence that justified the premium is still there after a year that included a major regional outage. That is an execution question and it does not get answered at a conference.</p>
<h2 id="what-to-actually-do-with-this">what to actually do with this<a class="anchor" href="#what-to-actually-do-with-this" aria-label="link to this section">#</a></h2>
<p>Skip the keynote. Read the "what's new" feed filtered to the services you actually use. Look for the deprecations, which are the announcements that will cost you time and which are never on stage.</p>
<p>And check your bill. The single highest-value hour available to most engineering organizations is someone competent looking at the AWS bill line by line, and almost nobody does it.</p>]]></content:encoded></item><item><title>Claude Opus 4.5 and the compaction problem</title><link>https://readme.news/claude-opus-45-and-the-compaction-problem/</link><guid isPermaLink="true">https://readme.news/claude-opus-45-and-the-compaction-problem/</guid><pubDate>Tue, 25 Nov 2025 09:00:00 +0000</pubDate><description>A frontier release with a large price cut, plus effort controls and context compaction as a first-class feature.</description><content:encoded><![CDATA[<p>Anthropic released Claude Opus 4.5 with a substantial price reduction relative to the previous Opus generation, an effort parameter for controlling reasoning depth, and improved context compaction.</p>
<p>The price cut is the headline for most users. The compaction work is more interesting.</p>
<h2 id="the-compaction-problem">the compaction problem<a class="anchor" href="#the-compaction-problem" aria-label="link to this section">#</a></h2>
<p>Long agent sessions fill their context. Tool outputs, file contents, <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a>, prior reasoning. Eventually you hit the limit and something has to go.</p>
<p>The naive approaches are all bad:</p>
<ul><li><strong>Truncate the oldest.</strong> Loses the original task description, which is the single most important thing in the context.</li><li><strong>Truncate the middle.</strong> Loses the reasoning chain that got you here.</li><li><strong>Summarize everything.</strong> Loses specifics — file paths, error strings, exact values — that turn out to matter.</li></ul>
<p>What you actually want is selective retention: keep the goal, keep the decisions and their rationale, keep the current state, discard the raw tool output that has already been acted on.</p>
<p>That is a judgment call, and doing it well requires understanding what the session is about. Which makes it a model problem rather than a buffer-management problem.</p>
<h2 id="why-this-matters-more-than-benchmark-deltas">why this matters more than benchmark deltas<a class="anchor" href="#why-this-matters-more-than-benchmark-deltas" aria-label="link to this section">#</a></h2>
<p>For agent workloads, context management determines whether a long task succeeds far more than a few points of benchmark difference.</p>
<p>I have watched agent runs fail in exactly this way: two hours in, compaction drops a detail — a constraint from the original request, a decision made an hour ago — and the agent proceeds confidently in a direction that contradicts the task. Everything after that is wasted, and it looks productive the whole time.</p>
<p>If you are building on any model, the lesson to steal is: <strong>do not rely on the context window as your memory.</strong> Maintain durable state outside it.</p>
<div class="code"><pre><code>task.md          — the goal, constraints, acceptance criteria. Re-read often.
notes.md         — decisions made and why. Appended, never rewritten.
state.json       — current progress, structured.</code></pre></div>
<p>Feed those back in after every compaction. This is cheap, model-agnostic, and it is the difference between an agent that works for four hours and one that works for forty minutes.</p>
<h2 id="the-effort-parameter">the effort parameter<a class="anchor" href="#the-effort-parameter" aria-label="link to this section">#</a></h2>
<p>Explicit control over reasoning depth, exposed to the caller. Everyone has this now under different names — <a class="xref" href="/gemini-25-goes-generally-available-with-a-thinking-dial/" title="Gemini 2.5 goes generally available with a thinking dial">thinking budget</a>, <a class="xref" href="/openai-ships-open-weights-for-the-first-time-since-gpt-2/" title="OpenAI ships open weights for the first time since GPT-2">reasoning effort</a>, thinking config.</p>
<p>The convergence is total and it confirms the design conclusion: reasoning depth belongs to the caller, not the model, because only the caller knows whether this particular request justifies the latency and the cost.</p>
<p>Measure the quality-cost curve on your own task. The knee is usually much lower than people assume.</p>
<h2 id="the-price-movement">the price movement<a class="anchor" href="#the-price-movement" aria-label="link to this section">#</a></h2>
<p><a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">Frontier model</a> pricing has fallen substantially across every provider over the past eighteen months, on a per-capability basis by considerably more.</p>
<p>Two implications:</p>
<p><strong>Cost-optimization work has a short half-life.</strong> Elaborate infrastructure to shave token costs may be obsolete before it pays for itself. Build the thing; optimize when the bill actually hurts.</p>
<p><strong>"Too expensive to do with a frontier model" is a moving line.</strong> Applications that did not pencil out a year ago may now. It is worth periodically revisiting the ideas you rejected on cost grounds, because the reason you rejected them keeps expiring.</p>
<h2 id="the-competitive-picture">the competitive picture<a class="anchor" href="#the-competitive-picture" aria-label="link to this section">#</a></h2>
<p>Three labs shipping frontier releases within a week of each other, with capability differences small enough to be within evaluation noise on many tasks.</p>
<p>The practical consequence for developers: your model choice is a preference, not a <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a>, and it should be revisited quarterly rather than defended.</p>
<p>Build the abstraction. It is a day of work and it keeps paying.</p>]]></content:encoded></item>
</channel>
</rss>
