<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — industry</title>
<link>https://readme.news/tags/industry/</link>
<atom:link href="https://readme.news/tags/industry/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged industry.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>Open source funding: five models, honestly compared</title><link>https://readme.news/open-source-funding-five-models-honestly-compared/</link><guid isPermaLink="true">https://readme.news/open-source-funding-five-models-honestly-compared/</guid><pubDate>Wed, 20 May 2026 09:00:00 +0000</pubDate><description>Donations, foundations, dual licensing, open core, and hosted service. Each works for a specific shape of project.</description><content:encoded><![CDATA[<p>The open source sustainability problem is not that the models do not exist. It is that projects pick a model that does not fit their shape, and then conclude that funding open source is impossible.</p>
<p>Five models. Each works. Each works for a specific kind of project.</p>
<h2 id="1-donations-and-sponsorship">1. donations and sponsorship<a class="anchor" href="#1-donations-and-sponsorship" aria-label="link to this section">#</a></h2>
<p><strong>Works for:</strong> developer-facing tools with a large individual user base and an identifiable maintainer.</p>
<p><strong>Does not work for:</strong> libraries deep in a dependency tree, infrastructure nobody knows they use, anything without a personality attached.</p>
<p>The honest arithmetic: donation income correlates with visibility, not with importance. A well-marketed CLI tool with a charismatic maintainer will out-earn a critical cryptography library by a large multiple.</p>
<p>The corporate sponsorship version works better than individual donations and requires the maintainer to do sales, which most maintainers are bad at and hate.</p>
<h2 id="2-foundation-governance">2. foundation governance<a class="anchor" href="#2-foundation-governance" aria-label="link to this section">#</a></h2>
<p><strong>Works for:</strong> infrastructure that multiple large companies depend on and that none of them wants a competitor to control.</p>
<p><strong>Does not work for:</strong> small projects. The overhead — governance, legal, trademark, process — is substantial, and a foundation with one project and no funded staff is just more paperwork.</p>
<p>The real value of a foundation is not money. It is neutrality: it makes a project safe for competitors to invest in together, which unlocks contribution that would not otherwise happen.</p>
<h2 id="3-dual-licensing">3. dual licensing<a class="anchor" href="#3-dual-licensing" aria-label="link to this section">#</a></h2>
<p><strong>Works for:</strong> libraries embedded in other products, where the copyleft obligation is genuinely inconvenient for commercial users.</p>
<p>Ship under a strong copyleft license, sell a commercial license to companies that cannot comply.</p>
<p><strong>Does not work for:</strong> anything permissively licensed already (no leverage), anything not embedded (the obligation does not bite), or anything with a permissive competitor of similar quality.</p>
<p>Effective when it fits, and it produces a genuine tension: the license that makes the business work is the one that limits adoption.</p>
<h2 id="4-open-core">4. open core<a class="anchor" href="#4-open-core" aria-label="link to this section">#</a></h2>
<p><strong>Works for:</strong> products where enterprise features are genuinely separable from the core — SSO, audit logs, RBAC, compliance reporting, multi-tenancy.</p>
<p><strong>Does not work for:</strong> libraries. There is no enterprise tier of a date-parsing library.</p>
<p>The failure mode is well documented: the line between core and commercial moves toward commercial over time, under revenue pressure, and the community that built your adoption watches features they use get moved behind the paywall.</p>
<p>If you do this, <strong>write down the line publicly, early, and honor it.</strong> "Anything that a single developer needs is open; anything that exists because you have a compliance department is commercial" is a defensible line. Moving it later costs more trust than the revenue is worth.</p>
<h2 id="5-hosted-service">5. hosted service<a class="anchor" href="#5-hosted-service" aria-label="link to this section">#</a></h2>
<p><strong>Works for:</strong> anything that is annoying to operate. Databases, search, <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a>, observability, CI.</p>
<p>Give away the software, sell the operation of it. This is the strongest model when it fits, because the value you sell — not having to run it — is real and continuous, and it does not require withholding anything.</p>
<p><strong>The risk:</strong> a hyperscaler offers a managed version of your software, at scale, without contributing back. This has happened repeatedly and it is why the source- available licenses exist.</p>
<p>Those licenses solve the problem and cost you the open source designation, which costs you contributors, ecosystem inclusion, and some corporate adoption. It is a real trade with real costs on both sides, and the projects that made it mostly survived, which is the empirical answer to whether it works.</p>
<h2 id="what-actually-kills-projects">what actually kills projects<a class="anchor" href="#what-actually-kills-projects" aria-label="link to this section">#</a></h2>
<p>Not the absence of a model. Three other things:</p>
<p><strong>Solo maintainer burnout.</strong> One person, unpaid, receiving an unbounded stream of issues, feature requests, and entitled demands. Funding helps and does not fix it — the fix is more maintainers, which is a governance problem.</p>
<p><strong>Success without support.</strong> A project that becomes critical infrastructure while its maintainer count stays at one. This is the most dangerous state and it is extremely common.</p>
<p><strong>Corporate abandonment.</strong> A company open-sources a project, staffs it with employees, then reorganizes. The external community was never built because it was never needed. Now nobody knows the code.</p>
<h2 id="what-companies-should-do">what companies should do<a class="anchor" href="#what-companies-should-do" aria-label="link to this section">#</a></h2>
<p>If your business depends on open source — and it does — the highest-leverage actions, in order:</p>
<ol><li><strong>Pay maintainers of your critical dependencies.</strong> Directly. Small amounts to many projects beat large amounts to a few.</li><li><strong>Assign employee time to upstream contribution.</strong> More valuable than money and much rarer.</li><li><strong>Do not send compliance questionnaires to volunteers.</strong> They owe you nothing and the license says so.</li><li><strong>When you fix a bug in a vendored dependency, upstream it.</strong> The number of companies carrying private patches for bugs everyone has is enormous.</li></ol>
<p>None of that requires a strategy document. It requires someone with budget deciding it matters, which is the actual bottleneck.</p>]]></content:encoded></item><item><title>Every company is briefly a model company</title><link>https://readme.news/every-company-is-briefly-a-model-company/</link><guid isPermaLink="true">https://readme.news/every-company-is-briefly-a-model-company/</guid><pubDate>Wed, 18 Feb 2026 09:00:00 +0000</pubDate><description>The fine-tuning wave, the RAG wave, and the agent wave all followed the same arc. Here&#x27;s where the value actually settled.</description><content:encoded><![CDATA[<p>Three times in three years, a wave of companies concluded that the way to build an AI product was to own a layer that turned out not to be theirs.</p>
<p>The pattern is consistent enough to be predictive, which makes it worth naming.</p>
<h2 id="wave-one-fine-tuning">wave one: fine-tuning<a class="anchor" href="#wave-one-fine-tuning" aria-label="link to this section">#</a></h2>
<p><strong>The thesis (2023):</strong> general models are generic. Fine-tune on your domain data and you get a model that is specifically good at your problem and that competitors cannot replicate.</p>
<p><strong>What happened:</strong> base models improved faster than fine-tunes could keep up. A fine-tuned model from six months ago was worse than the new base model with a good prompt. Every fine-tune had to be redone on every model release, which is a treadmill.</p>
<p><strong>Where it settled:</strong> fine-tuning is genuinely valuable for narrow, stable, high-volume tasks — classification into your specific taxonomy, output in your specific format, a <a class="xref" href="/haiku-45-and-the-collapsing-cost-of-good-enough/" title="Haiku 4.5 and the collapsing cost of good-enough">small model</a> matching a large model's behavior on one task. It is not a moat and it is not a product strategy.</p>
<h2 id="wave-two-rag">wave two: RAG<a class="anchor" href="#wave-two-rag" aria-label="link to this section">#</a></h2>
<p><strong>The thesis (2023-24):</strong> the model does not know your data. Build a retrieval pipeline — chunk, embed, index, retrieve, rerank — and you have a defensible system built on proprietary knowledge.</p>
<p><strong>What happened:</strong> context windows grew by two orders of magnitude, long-context quality improved, and prompt <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> made large contexts economical. A large fraction of naive RAG got replaced by putting the documents in the prompt.</p>
<p>Simultaneously, the pipeline components commoditized. Embedding models became interchangeable. Vector search became a feature of every database rather than a product.</p>
<p><strong>Where it settled:</strong> retrieval did not go away. It moved. Retrieval is now a <em>tool the model calls</em> rather than a preprocessing step, and it is genuinely necessary at large corpus sizes, where citation is required, and where cost or latency rules out large contexts.</p>
<p>The infrastructure repositioned rather than dying, which is what usually happens.</p>
<h2 id="wave-three-agent-frameworks">wave three: agent frameworks<a class="anchor" href="#wave-three-agent-frameworks" aria-label="link to this section">#</a></h2>
<p><strong>The thesis (2024-25):</strong> models cannot plan reliably. Build the orchestration — task decomposition, tool routing, retry logic, state management — and own the layer that makes agents work.</p>
<p><strong>What happened:</strong> the models absorbed it. In-context tool use during reasoning removed the need for an external loop. Native parallel tool calling removed the need for a dispatcher. Memory tools and <a class="xref" href="/claude-sonnet-45-and-the-agent-that-runs-for-thirty-hours/" title="Claude Sonnet 4.5 and the agent that runs for thirty hours">context editing</a> removed the need for external state management.</p>
<p>Each capability the framework provided became a model feature within about a year.</p>
<p><strong>Where it settled:</strong> in progress, but the shape is clear. Orchestration frameworks are converging on thin conveniences. The durable parts are evaluation, observability, and the domain-specific policy that no model will ever have.</p>
<h2 id="the-pattern">the pattern<a class="anchor" href="#the-pattern" aria-label="link to this section">#</a></h2>
<p>Every wave follows the same arc:</p>
<ol><li>The model has a limitation.</li><li>Companies build infrastructure to work around the limitation.</li><li>The limitation gets fixed in the model.</li><li>The infrastructure either finds a new position or disappears.</li></ol>
<p>The consistent error is <strong>building on a gap rather than on an asset</strong>. A gap is temporary by construction — the labs are actively working to close it, with more resources than you have.</p>
<h2 id="what-has-actually-been-durable">what has actually been durable<a class="anchor" href="#what-has-actually-been-durable" aria-label="link to this section">#</a></h2>
<p>Across all three waves, the same things kept their value:</p>
<p><strong>Proprietary data.</strong> Not "we have documents" — everyone has documents. Data that is genuinely yours: production logs, customer interactions, labeled outcomes, domain expertise encoded as examples. Nobody can buy it and no model was trained on it.</p>
<p><strong>Evaluation specific to your task.</strong> The company that knows, precisely, how well a system performs on their actual problem can adopt a new model in a day. The one that does not spends a month on vibes. That gap compounds every release cycle.</p>
<p><strong>Distribution and workflow integration.</strong> Being where the user already works. The most boring answer and the most durable.</p>
<p><strong>Domain constraints.</strong> The rules, regulations, edge cases, and institutional knowledge that make a generic capability into a usable product. This is unglamorous and it is the actual work.</p>
<p><strong>Trust.</strong> Security posture, compliance, reliability, support. Enterprises buy this and it takes years to build.</p>
<h2 id="the-test-to-apply">the test to apply<a class="anchor" href="#the-test-to-apply" aria-label="link to this section">#</a></h2>
<p>Before building on top of a model limitation, ask: <strong>if this limitation disappeared next quarter, what would I have left?</strong></p>
<p>If the answer is "nothing," you are building a bridge over a river that is being drained.</p>
<p>If the answer is "the data, the evaluations, the integrations, and the customer relationships," build it, and expect to throw the bridge away.</p>]]></content:encoded></item><item><title>Claude Opus 4.5 and the compaction problem</title><link>https://readme.news/claude-opus-45-and-the-compaction-problem/</link><guid isPermaLink="true">https://readme.news/claude-opus-45-and-the-compaction-problem/</guid><pubDate>Tue, 25 Nov 2025 09:00:00 +0000</pubDate><description>A frontier release with a large price cut, plus effort controls and context compaction as a first-class feature.</description><content:encoded><![CDATA[<p>Anthropic released Claude Opus 4.5 with a substantial price reduction relative to the previous Opus generation, an effort parameter for controlling reasoning depth, and improved context compaction.</p>
<p>The price cut is the headline for most users. The compaction work is more interesting.</p>
<h2 id="the-compaction-problem">the compaction problem<a class="anchor" href="#the-compaction-problem" aria-label="link to this section">#</a></h2>
<p>Long agent sessions fill their context. Tool outputs, file contents, <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a>, prior reasoning. Eventually you hit the limit and something has to go.</p>
<p>The naive approaches are all bad:</p>
<ul><li><strong>Truncate the oldest.</strong> Loses the original task description, which is the single most important thing in the context.</li><li><strong>Truncate the middle.</strong> Loses the reasoning chain that got you here.</li><li><strong>Summarize everything.</strong> Loses specifics — file paths, error strings, exact values — that turn out to matter.</li></ul>
<p>What you actually want is selective retention: keep the goal, keep the decisions and their rationale, keep the current state, discard the raw tool output that has already been acted on.</p>
<p>That is a judgment call, and doing it well requires understanding what the session is about. Which makes it a model problem rather than a buffer-management problem.</p>
<h2 id="why-this-matters-more-than-benchmark-deltas">why this matters more than benchmark deltas<a class="anchor" href="#why-this-matters-more-than-benchmark-deltas" aria-label="link to this section">#</a></h2>
<p>For agent workloads, context management determines whether a long task succeeds far more than a few points of benchmark difference.</p>
<p>I have watched agent runs fail in exactly this way: two hours in, compaction drops a detail — a constraint from the original request, a decision made an hour ago — and the agent proceeds confidently in a direction that contradicts the task. Everything after that is wasted, and it looks productive the whole time.</p>
<p>If you are building on any model, the lesson to steal is: <strong>do not rely on the context window as your memory.</strong> Maintain durable state outside it.</p>
<div class="code"><pre><code>task.md          — the goal, constraints, acceptance criteria. Re-read often.
notes.md         — decisions made and why. Appended, never rewritten.
state.json       — current progress, structured.</code></pre></div>
<p>Feed those back in after every compaction. This is cheap, model-agnostic, and it is the difference between an agent that works for four hours and one that works for forty minutes.</p>
<h2 id="the-effort-parameter">the effort parameter<a class="anchor" href="#the-effort-parameter" aria-label="link to this section">#</a></h2>
<p>Explicit control over reasoning depth, exposed to the caller. Everyone has this now under different names — <a class="xref" href="/gemini-25-goes-generally-available-with-a-thinking-dial/" title="Gemini 2.5 goes generally available with a thinking dial">thinking budget</a>, <a class="xref" href="/openai-ships-open-weights-for-the-first-time-since-gpt-2/" title="OpenAI ships open weights for the first time since GPT-2">reasoning effort</a>, thinking config.</p>
<p>The convergence is total and it confirms the design conclusion: reasoning depth belongs to the caller, not the model, because only the caller knows whether this particular request justifies the latency and the cost.</p>
<p>Measure the quality-cost curve on your own task. The knee is usually much lower than people assume.</p>
<h2 id="the-price-movement">the price movement<a class="anchor" href="#the-price-movement" aria-label="link to this section">#</a></h2>
<p><a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">Frontier model</a> pricing has fallen substantially across every provider over the past eighteen months, on a per-capability basis by considerably more.</p>
<p>Two implications:</p>
<p><strong>Cost-optimization work has a short half-life.</strong> Elaborate infrastructure to shave token costs may be obsolete before it pays for itself. Build the thing; optimize when the bill actually hurts.</p>
<p><strong>"Too expensive to do with a frontier model" is a moving line.</strong> Applications that did not pencil out a year ago may now. It is worth periodically revisiting the ideas you rejected on cost grounds, because the reason you rejected them keeps expiring.</p>
<h2 id="the-competitive-picture">the competitive picture<a class="anchor" href="#the-competitive-picture" aria-label="link to this section">#</a></h2>
<p>Three labs shipping frontier releases within a week of each other, with capability differences small enough to be within evaluation noise on many tasks.</p>
<p>The practical consequence for developers: your model choice is a preference, not a <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a>, and it should be revisited quarterly rather than defended.</p>
<p>Build the abstraction. It is a day of work and it keeps paying.</p>]]></content:encoded></item><item><title>Nvidia's quarter and the question nobody can answer</title><link>https://readme.news/nvidias-quarter-and-the-question-nobody-can-answer/</link><guid isPermaLink="true">https://readme.news/nvidias-quarter-and-the-question-nobody-can-answer/</guid><pubDate>Fri, 21 Nov 2025 09:00:00 +0000</pubDate><description>Another enormous beat, another set of concerns about circular financing. Both facts are real.</description><content:encoded><![CDATA[<p>Nvidia reported another quarter far above expectations, with data center revenue continuing to grow at a rate that would be implausible in any other context.</p>
<p>The stock's reaction was muted relative to the beat, which tells you the debate has moved from "is demand real" to "is demand <em>sustainable</em>, and how much of it is funded by Nvidia."</p>
<h2 id="the-bull-case">the bull case<a class="anchor" href="#the-bull-case" aria-label="link to this section">#</a></h2>
<p>Straightforward and well supported:</p>
<ul><li>Every hyperscaler raised capital expenditure guidance again.</li><li>Inference demand is growing faster than training demand, and inference is the recurring workload rather than the one-time one.</li><li>Reasoning models consume dramatically more inference compute per request than their predecessors, and adoption is rising.</li><li>Supply remains the constraint. Lead times are long. Customers are queueing.</li><li>The <a class="xref" href="/gtc-2025-a-roadmap-to-2027-and-a-warning-about-power/" title="GTC 2025: a roadmap to 2027 and a warning about power">rack-scale</a> systems business has a moat that individual chip competition does not touch — the interconnect is the product.</li></ul>
<h2 id="the-bear-case">the bear case<a class="anchor" href="#the-bear-case" aria-label="link to this section">#</a></h2>
<p>Also straightforward:</p>
<ul><li>A meaningful share of revenue traces to customers Nvidia has invested in or financed, which makes the demand signal less independent.</li><li>Depreciation schedules on AI hardware are assumed at five to six years. If the useful life is closer to three — which some operators argue, given the pace of generational improvement — reported profitability across the sector is overstated.</li><li>Hyperscalers are building their own silicon. Google's TPUs are mature, Amazon's Trainium is shipping in volume, and every one of those deployments is a substituted Nvidia sale.</li><li>Model efficiency improvements keep arriving. If capability-per-FLOP keeps improving as fast as it has, required FLOPs for a given capability fall.</li><li>The financing environment for the buildout depends on continued access to cheap debt.</li></ul>
<h2 id="what-nobody-knows">what nobody knows<a class="anchor" href="#what-nobody-knows" aria-label="link to this section">#</a></h2>
<p>Whether AI application revenue will eventually justify the infrastructure spend.</p>
<p>Current annualized revenue across the AI application layer is a fraction of annual AI capital expenditure. That gap can close — infrastructure is built ahead of demand in every capital cycle, and railroads, fiber, and cloud all looked insane at the equivalent stage.</p>
<p>It can also not close. Fiber overbuild in 2000 was followed by a decade of dark fiber and a lot of bankruptcies, and the eventual users of that fiber were not the companies that laid it.</p>
<p>Both patterns are real. The people confidently predicting which one applies here are pattern-matching, not analyzing, and that includes the ones I agree with.</p>
<h2 id="why-an-engineer-should-care">why an engineer should care<a class="anchor" href="#why-an-engineer-should-care" aria-label="link to this section">#</a></h2>
<p>Not for investing advice. For planning.</p>
<p><strong>Compute pricing is not going to fall smoothly.</strong> If the capex cycle continues, capacity comes online in steps and prices drift down. If it contracts, capacity tightens and prices firm. Do not build a business model that requires a specific trajectory.</p>
<p><strong>Efficiency work has enduring value.</strong> Whatever happens to the capex cycle, using less compute for the same result is good. Prompt <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a>, model routing, smaller models for routine tasks, batch processing where latency permits — all of it pays regardless of the macro environment.</p>
<p><strong>Multi-provider capability is cheap insurance.</strong> The cost of abstracting your model calls is a day. The cost of being locked to a provider whose pricing or availability changes is much larger.</p>
<h2 id="the-sentence-i-keep-coming-back-to">the sentence I keep coming back to<a class="anchor" href="#the-sentence-i-keep-coming-back-to" aria-label="link to this section">#</a></h2>
<p>The infrastructure being built is real, the demand today is real, and whether they match at the scale being assumed is genuinely unknown to everyone including the people spending the money.</p>
<p>That is an uncomfortable place to be, and pretending otherwise — in either direction — is the main thing to avoid.</p>]]></content:encoded></item><item><title>GPT-5.1 and the return of the model picker</title><link>https://readme.news/gpt-51-and-the-return-of-the-model-picker/</link><guid isPermaLink="true">https://readme.news/gpt-51-and-the-return-of-the-model-picker/</guid><pubDate>Wed, 12 Nov 2025 09:00:00 +0000</pubDate><description>Instant and Thinking as named modes, adaptive reasoning, and personality controls. The router lesson got learned.</description><content:encoded><![CDATA[<p>OpenAI released GPT-5.1 with two named variants — Instant and Thinking — plus adaptive reasoning that adjusts thinking time by question difficulty, and a set of tone presets.</p>
<p>Three months after removing the model picker caused a backlash, the picker is back with better names. That is a reasonable outcome and the intermediate lesson is worth stating.</p>
<h2 id="the-routing-lesson">the routing lesson<a class="anchor" href="#the-routing-lesson" aria-label="link to this section">#</a></h2>
<p>The original GPT-5 design routed automatically and hid the choice. The intent was good — most users have no basis for choosing a model — and the execution exposed a real problem: <strong>when automatic selection fails, the user has no way to diagnose it or override it.</strong></p>
<p>The user experiences "the model got worse." They cannot tell whether they hit a bad route, a degraded model, or their own bad prompt. There is no signal and no recourse.</p>
<p>5.1's approach — automatic by default, with named modes available — is the right shape. It is also exactly what every well-designed automatic system does: sensible defaults, visible state, manual override.</p>
<p>If you build anything that routes between models, ship the override. It costs one UI control and it eliminates an entire category of unfalsifiable user complaint.</p>
<h2 id="adaptive-reasoning">adaptive reasoning<a class="anchor" href="#adaptive-reasoning" aria-label="link to this section">#</a></h2>
<p>The model adjusts thinking time based on assessed difficulty rather than applying a uniform budget. Easy questions answer immediately; hard questions get more compute.</p>
<p>This is a straightforwardly good idea and every provider is converging on it. The implementation question is calibration: a model that underestimates difficulty gives you a fast wrong answer, and a model that overestimates it burns money.</p>
<p>For API users, the practical guidance is the same as always: <strong>measure on your own task distribution.</strong> Adaptive reasoning is a good default and it is not tuned for your workload. If you have a task mix that skews harder or easier than average, set the budget explicitly.</p>
<h2 id="the-personality-controls">the personality controls<a class="anchor" href="#the-personality-controls" aria-label="link to this section">#</a></h2>
<p>Tone presets — Professional, Friendly, Candid, Quirky, and others — plus finer adjustment of warmth and conciseness.</p>
<p>This got the most consumer coverage and it is the least technically interesting change. It is also a reasonable response to the fact that removing GPT-4o generated complaints about <em>voice</em>, not capability.</p>
<p>For developers, this is what a system prompt already did. The value is for consumer users who were not going to write one.</p>
<h2 id="what-i-would-actually-check">what I would actually check<a class="anchor" href="#what-i-would-actually-check" aria-label="link to this section">#</a></h2>
<p>Whenever a point release lands, three things:</p>
<p><strong>Instruction following on your specific format.</strong> Point releases change how literally the model follows formatting instructions surprisingly often. If you parse structured output, test it.</p>
<p><strong>Refusal behavior.</strong> Safety tuning shifts between versions. If your application is in a domain that skirts a policy boundary — security research, medical information, legal content — re-run your test set. False refusals are a real production problem and they change silently.</p>
<p><strong>Latency distribution, not average.</strong> Adaptive reasoning means variance. If you have a latency SLA, measure p95 and p99, not the mean.</p>
<h2 id="the-state-of-the-frontier">the state of the frontier<a class="anchor" href="#the-state-of-the-frontier" aria-label="link to this section">#</a></h2>
<p>Three labs are now shipping point releases every few months rather than major versions annually, with capability differences that are small and getting smaller.</p>
<p>That is what a mature market looks like. The differentiation is moving to price, latency, ecosystem, and trust — and to the products built on top rather than the models themselves.</p>
<p>For anyone building applications, this is unambiguously good news. It means your model choice is increasingly reversible, and reversible decisions should be made quickly and revisited often.</p>]]></content:encoded></item><item><title>Haiku 4.5 and the collapsing cost of good-enough</title><link>https://readme.news/haiku-45-and-the-collapsing-cost-of-good-enough/</link><guid isPermaLink="true">https://readme.news/haiku-45-and-the-collapsing-cost-of-good-enough/</guid><pubDate>Thu, 16 Oct 2025 09:00:00 +0000</pubDate><description>A small model at frontier-adjacent coding performance, priced for volume. The economics of agent fleets just changed.</description><content:encoded><![CDATA[<p>Anthropic released Claude Haiku 4.5, a small fast model with coding performance in the neighborhood of the previous generation's mid-tier, at a fraction of the price and several times the speed.</p>
<p>The headline is not the benchmark. It is what the price-performance point makes economically viable.</p>
<h2 id="the-pattern-across-the-industry">the pattern across the industry<a class="anchor" href="#the-pattern-across-the-industry" aria-label="link to this section">#</a></h2>
<p>Every major provider now ships roughly the same ladder:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">tier</th><th style="text-align:left">role</th><th style="text-align:left">relative cost</th></tr></thead><tbody><tr><td style="text-align:left">frontier</td><td style="text-align:left">hard reasoning, planning, novel problems</td><td style="text-align:left">1×</td></tr><tr><td style="text-align:left">mid</td><td style="text-align:left">most production work</td><td style="text-align:left">~0.2×</td></tr><tr><td style="text-align:left">small</td><td style="text-align:left">high-volume, well-defined tasks</td><td style="text-align:left">~0.03×</td></tr></tbody></table></div>
<p>The interesting fact is that the <em>small</em> tier's capability is rising faster than the frontier's. Today's small model is roughly where the frontier was eighteen months ago, and it costs about two percent as much.</p>
<p>For anyone doing volume work, that is the number that matters. Not "how smart is the best model" but "how cheap is the model that is good enough for this task."</p>
<h2 id="what-it-enables">what it enables<a class="anchor" href="#what-it-enables" aria-label="link to this section">#</a></h2>
<p><strong>Agent fleets.</strong> If a small model can handle subtasks reliably, an orchestrator using a <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> can dispatch twenty parallel workers using a small one. Total cost stays reasonable; throughput multiplies. This architecture was uneconomic a year ago and is now obvious.</p>
<p><strong>Real-time interaction.</strong> Latency matters more than capability for anything a human is waiting on. A fast model that is 90% as good and 5× faster wins on almost any interactive surface.</p>
<p><strong>Processing everything instead of sampling.</strong> Classification, extraction, routing, and enrichment over an entire corpus rather than a sample. At small-model prices, "run it on all of it" becomes the default rather than a budget question.</p>
<p><strong>Pipeline stages that were not worth it.</strong> Adding a model call to normalize an input, or to double-check an output, or to summarize an intermediate result — each of these was a cost decision at frontier prices and is now free enough to just do.</p>
<h2 id="the-routing-architecture">the routing architecture<a class="anchor" href="#the-routing-architecture" aria-label="link to this section">#</a></h2>
<p>This is the shape that production systems are converging on:</p>
<div class="code"><pre><code>request → cheap classifier → simple?  → small model → done
                           → complex? → frontier model → done
                           → ambiguous? → small model → verifier → escalate if low confidence</code></pre></div>
<p>Most requests take the cheap path. The expensive model handles the tail. Overall cost is dominated by the common case, which is cheap.</p>
<p>Two implementation notes that matter:</p>
<p><strong>The classifier can be the small model itself</strong>, or often a much simpler heuristic. Do not over-engineer the router — a regex and an input-length check gets you surprisingly far.</p>
<p><strong>Instrument the escalation rate.</strong> If it is climbing, either your traffic mix changed or your small-model prompts have degraded. That metric is the <a class="xref" href="/your-monitoring-is-measuring-the-wrong-nines/" title="Your monitoring is measuring the wrong nines">health check</a> for the whole architecture.</p>
<h2 id="the-caution">the caution<a class="anchor" href="#the-caution" aria-label="link to this section">#</a></h2>
<p>Small models fail differently than large ones. They do not fail <em>less</em> on the tasks they handle — they fail <em>more confidently</em> on the ones just outside their envelope.</p>
<p>A large model asked something beyond its ability will often hedge. A small model will answer, fluently and wrongly.</p>
<p>So: verify the output of small models where correctness matters. A cheap verifier — schema validation, a test run, a range check, a second model asked to critique — costs almost nothing and catches the category of failure that will otherwise reach production silently.</p>
<p>The cost savings are real. They come with a verification obligation, and the teams that skip it will find out in a quarter.</p>]]></content:encoded></item><item><title>Windows 10 reaches end of support</title><link>https://readme.news/windows-10-reaches-end-of-support/</link><guid isPermaLink="true">https://readme.news/windows-10-reaches-end-of-support/</guid><pubDate>Tue, 14 Oct 2025 09:00:00 +0000</pubDate><description>A very large number of machines stop getting security updates. The e-waste and the enterprise scramble are both real.</description><content:encoded><![CDATA[<p>Windows 10 stopped receiving free security updates today. Estimates for the installed base still running it range widely; every one of them is in the hundreds of millions of machines.</p>
<h2 id="the-tpm-problem">the TPM problem<a class="anchor" href="#the-tpm-problem" aria-label="link to this section">#</a></h2>
<p>The reason this is unusually messy is Windows 11's hardware requirements, specifically TPM 2.0 and the supported-CPU list.</p>
<p>Microsoft's justification is security: a hardware root of trust enables virtualization-based security, credential guard, and measured boot, and these meaningfully reduce entire categories of attack. That argument is technically sound.</p>
<p>The consequence is that a large number of machines that are functionally fine — adequate CPU, plenty of RAM, working perfectly — cannot install the supported operating system. Not because they are slow. Because of a security chip.</p>
<p>The environmental math is grim. Estimates of machines rendered security-obsolete run into the hundreds of millions. Those devices do not stop working; they become unpatched machines on the internet, or they become e-waste, and both outcomes are bad.</p>
<h2 id="the-options">the options<a class="anchor" href="#the-options" aria-label="link to this section">#</a></h2>
<p><strong>Extended Security Updates.</strong> Consumers get a limited free path (with conditions) or a modest paid year. Enterprises pay per device, per year, escalating annually, for up to three years. Budget for it if you have a fleet — the escalation is designed to be painful on purpose.</p>
<p><strong>Upgrade the hardware.</strong> The intended path. Expensive at fleet scale and sometimes impossible for machines running certified software tied to specific configurations.</p>
<p><strong>Bypass the requirements.</strong> Documented registry workarounds exist and work. Microsoft has warned that unsupported installations may not receive updates, which makes this a poor choice for anything you depend on.</p>
<p><strong>Move to Linux.</strong> Genuinely viable for a larger share of use cases than it was five years ago, particularly for developer machines and for kiosk or single-app deployments. Several distributions ran campaigns targeting exactly this moment.</p>
<p>The blocker remains what it has always been: specific Windows-only applications, and hardware with Windows-only drivers.</p>
<h2 id="for-developers-specifically">for developers specifically<a class="anchor" href="#for-developers-specifically" aria-label="link to this section">#</a></h2>
<p><strong>Check your minimum supported version.</strong> If your application still supports Windows 10, decide when it stops. Users on an unsupported OS are a support burden and a security liability, and the decision to drop them is easier to make now with a clear industry line to point at.</p>
<p><strong>Test on Windows 11.</strong> Particularly anything touching security features: credential storage, code signing, driver interaction, or anything that uses the TPM. Behavior differs.</p>
<p><strong>If you ship developer tooling</strong>, the Windows developer story is meaningfully better than it was — WSL 2, the modern terminal, winget, and PowerShell 7 are all good. If your Windows support is a grudging afterthought from 2018, it is worth revisiting.</p>
<h2 id="the-broader-pattern">the broader pattern<a class="anchor" href="#the-broader-pattern" aria-label="link to this section">#</a></h2>
<p>Operating system lifecycle transitions are increasingly forced by security architecture rather than by capability. The machine is fast enough. The machine lacks a specific security primitive that the new threat model requires.</p>
<p>This will happen again — with memory tagging, with pointer authentication, with whatever comes after. The useful lesson for anyone planning a fleet: hardware lifespan is now set by the security roadmap, not by performance. Plan accordingly, and push back on vendors who make that window shorter than it needs to be.</p>]]></content:encoded></item><item><title>Sora 2 and the feed nobody asked for</title><link>https://readme.news/sora-2-and-the-feed-nobody-asked-for/</link><guid isPermaLink="true">https://readme.news/sora-2-and-the-feed-nobody-asked-for/</guid><pubDate>Tue, 30 Sep 2025 09:00:00 +0000</pubDate><description>A better video model wrapped in a social app, with a cameo feature that makes consent the whole product question.</description><content:encoded><![CDATA[<p>OpenAI released Sora 2 alongside a standalone social app: a vertical video feed of AI-generated clips, with a "cameo" feature that lets you insert a verified likeness of yourself or a consenting friend into generated videos.</p>
<h2 id="the-model">the model<a class="anchor" href="#the-model" aria-label="link to this section">#</a></h2>
<p>Genuinely better than the first version, in ways that matter:</p>
<ul><li><strong>Synchronized audio.</strong> Dialogue, effects, ambience generated with the video.</li><li><strong>Physical plausibility.</strong> Objects have more consistent mass and momentum. Things that fall, fall correctly. Water behaves like water.</li><li><strong>Failure realism.</strong> A demonstrated example: a basketball shot that misses bounces off the rim, rather than teleporting into the hoop because the model learned that shots go in. Modeling failure states is a meaningful step toward actual physics rather than outcome mimicry.</li><li><strong>Multi-shot consistency.</strong> The same character and setting across cuts.</li></ul>
<h2 id="the-app">the app<a class="anchor" href="#the-app" aria-label="link to this section">#</a></h2>
<p>A TikTok-shaped feed where every video is generated. OpenAI's stated framing is creation over consumption, with feed controls and usage prompts.</p>
<p>I am skeptical, and the skepticism is structural rather than about intent. An infinite feed of content optimized for engagement has one known equilibrium, and it does not depend on whether the content is human-made. If anything, generated content removes the last friction — there is no supply constraint at all.</p>
<p>The genuinely novel bit is <strong>cameos</strong>: a verified likeness capture, with control over who can use it, revocable, with notification when it appears in someone's video.</p>
<p>That consent architecture is thoughtful. It is also the thing that will be stress-tested immediately, because likeness in generated video is the single most socially dangerous capability in this space and "we built a consent flow" is a much better answer than most products have.</p>
<p>Revocation is the hard part. You can revoke permission going forward. You cannot revoke a video someone downloaded.</p>
<h2 id="the-rights-problem">the rights problem<a class="anchor" href="#the-rights-problem" aria-label="link to this section">#</a></h2>
<p>Reporting around the launch indicated a rightsholder posture that put the burden on IP owners to opt out rather than requiring opt-in, with an announced shift toward more granular controls after pushback.</p>
<p>Opt-out for likeness and IP is a defensible engineering default and an indefensible ethical one. It puts the cost of protection on the person being depicted, who may not know the product exists.</p>
<p>I expect this to be litigated and legislated, in that order, and I expect opt-in to win for likeness specifically because that is where the political consensus already is.</p>
<h2 id="for-developers">for developers<a class="anchor" href="#for-developers" aria-label="link to this section">#</a></h2>
<p>The API is available and the practical questions are the same as for image generation, more sharply:</p>
<ul><li><strong>Provenance metadata on everything.</strong> C2PA, watermarking, whatever your pipeline supports. This is going to be a requirement, not a nicety.</li><li><strong>Consent flows for likeness, designed in from the start.</strong> Retrofitting consent after you have a user base is a nightmare.</li><li><strong>Understand your jurisdiction's rules on synthetic media.</strong> They are being written right now and they differ substantially between the EU, several US states, and everywhere else.</li></ul>
<h2 id="the-broader-read">the broader read<a class="anchor" href="#the-broader-read" aria-label="link to this section">#</a></h2>
<p>We are about eighteen months from generated video being indistinguishable from recorded video for a casual viewer, and the social infrastructure for that — norms about disclosure, verification for journalism, legal standards for evidence — does not exist.</p>
<p>The technology is arriving considerably faster than the institutions. That is not a new observation, and it has rarely been this compressed.</p>]]></content:encoded></item><item><title>Claude Sonnet 4.5 and the agent that runs for thirty hours</title><link>https://readme.news/claude-sonnet-45-and-the-agent-that-runs-for-thirty-hours/</link><guid isPermaLink="true">https://readme.news/claude-sonnet-45-and-the-agent-that-runs-for-thirty-hours/</guid><pubDate>Fri, 26 Sep 2025 09:00:00 +0000</pubDate><description>A model tuned for long-horizon autonomous work, plus checkpoints and context editing in the SDK.</description><content:encoded><![CDATA[<p>Anthropic released Claude Sonnet 4.5 with claims centered on sustained autonomous operation — reportedly maintaining focus on complex multi-step tasks for over thirty hours.</p>
<p>Alongside it: checkpoints in Claude Code, a VS Code extension, and context editing plus a memory tool in the API.</p>
<h2 id="the-long-horizon-claim">the long-horizon claim<a class="anchor" href="#the-long-horizon-claim" aria-label="link to this section">#</a></h2>
<p>Thirty hours is a marketing number and the underlying capability is real and worth understanding.</p>
<p>The limiting factor on long agent runs has never been the context window. It is <strong>goal drift</strong>. As a session accumulates tool results, file contents, <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a>, and the model's own prior reasoning, attention to the original objective degrades. The model starts optimizing for local signals — making this specific test pass — rather than the actual task.</p>
<p>The failure is insidious because each individual step looks reasonable. You come back after two hours to find the agent has been productively working on something adjacent to what you asked.</p>
<p>Improvements here come from three places: better training on long trajectories, architectural support for externalized memory, and mechanisms for periodically re-grounding on the original goal. This release touches all three.</p>
<h2 id="the-tooling">the tooling<a class="anchor" href="#the-tooling" aria-label="link to this section">#</a></h2>
<p><strong>Checkpoints</strong> in Claude Code. Save state, let the agent work, roll back if it goes wrong. This is the feature that makes long autonomous runs practically usable — the failure mode of a two-hour agent run is that you have to throw away two hours, and checkpointing converts that into "roll back twenty minutes."</p>
<p>Every delegated agent product needs this. It is the <code>git stash</code> of agent workflows.</p>
<p><strong>Context editing</strong> in the API — programmatically removing content from the conversation. Sounds mundane, matters a lot. A tool result containing 40,000 tokens of build output is useful for one turn and is pure noise for the next fifty. Being able to drop it keeps the context focused and the cost down.</p>
<p><strong>Memory tool</strong> — the model writes notes to a file and reads them back. An externalized working memory that survives <a class="xref" href="/claude-opus-45-and-the-compaction-problem/" title="Claude Opus 4.5 and the compaction problem">compaction</a>. This is the right pattern and it mirrors how people actually work on long tasks: you do not hold everything in your head, you write it down.</p>
<h2 id="the-pattern-to-steal">the pattern to steal<a class="anchor" href="#the-pattern-to-steal" aria-label="link to this section">#</a></h2>
<p>Whatever model you use, this architecture is the one that works for long tasks:</p>
<ol><li><strong>A durable task description</strong> in a file, not in the conversation.</li><li><strong>A working notes file</strong> the agent updates as it goes.</li><li><strong>Aggressive pruning</strong> of tool output from the context once it has been acted on.</li><li><strong>Checkpoints</strong> at natural boundaries so failure costs minutes, not hours.</li><li><strong>Periodic re-grounding</strong> — literally re-reading the task description and asking whether current work serves it.</li></ol>
<p>You can implement all of this yourself with any model and a bit of orchestration. The vendors shipping it as a feature is a convenience, not a requirement.</p>
<h2 id="the-caution">the caution<a class="anchor" href="#the-caution" aria-label="link to this section">#</a></h2>
<p>A model that can work autonomously for thirty hours can also do thirty hours of damage.</p>
<p>The controls that matter scale with autonomy: run in a container, restrict credentials to what the task requires, require approval for anything irreversible, and review the diff.</p>
<p>Higher autonomy makes review harder and more important simultaneously. That tension does not resolve — it is the central design problem of the entire category, and nobody has a good answer beyond "keep a human in the loop and make the loop cheap."</p>]]></content:encoded></item><item><title>OpenAI and Nvidia sign a circular deal</title><link>https://readme.news/openai-and-nvidia-sign-a-circular-deal/</link><guid isPermaLink="true">https://readme.news/openai-and-nvidia-sign-a-circular-deal/</guid><pubDate>Tue, 23 Sep 2025 09:00:00 +0000</pubDate><description>Up to $100 billion of investment tied to gigawatts of deployment. The financing structures are getting interesting.</description><content:encoded><![CDATA[<p>Nvidia and OpenAI announced a letter of intent under which Nvidia would invest up to $100 billion in OpenAI, staged against the deployment of at least 10 gigawatts of Nvidia systems.</p>
<p>Read that structure carefully, because it is the interesting part.</p>
<h2 id="the-circularity">the circularity<a class="anchor" href="#the-circularity" aria-label="link to this section">#</a></h2>
<p>Nvidia invests in OpenAI. OpenAI uses the money to buy Nvidia systems. The purchase is recognized as Nvidia revenue. The investment is staged against deployment milestones.</p>
<p>This is not fraud and it is not unusual in capital-intensive industries — vendor financing has been standard in telecom, aviation, and semiconductor equipment for decades. A supplier finances a customer's purchase because the supplier has the balance sheet and wants the volume.</p>
<p>It does deserve scrutiny for a specific reason: it makes the demand signal less informative. When a supplier funds its customer's purchases, revenue growth no longer cleanly indicates independent market demand. Some portion of it is the supplier's own capital cycling through.</p>
<p>Analysts have been tracking a widening set of these arrangements across the AI sector — investments in customers, prepayments, equity stakes in companies that are also large purchasers. Individually each is defensible. Collectively they make the sector's growth figures harder to interpret.</p>
<h2 id="the-gigawatt-as-a-unit">the gigawatt as a unit<a class="anchor" href="#the-gigawatt-as-a-unit" aria-label="link to this section">#</a></h2>
<p>Note what is being measured. Not chips, not dollars, not FLOPs. <strong>Gigawatts.</strong></p>
<p>Ten gigawatts is on the order of the electricity consumption of a large metropolitan area. It is roughly ten large nuclear reactors' worth of continuous generation.</p>
<p>The industry has converged on power as the natural unit because power is the binding constraint. You can order chips. You cannot order a substation and have it next quarter.</p>
<p>The consequences flow outward: electricity prices in datacenter-heavy regions, grid interconnection <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a>, transmission buildout, and the political economy of who pays for it. Several US utility regulators are now handling rate cases that are effectively about whether residential customers subsidize datacenter connections.</p>
<p>That fight is going to define a lot of the next five years and it is being had in public utility commission hearings that nobody in tech reads.</p>
<h2 id="what-it-means-for-you">what it means for you<a class="anchor" href="#what-it-means-for-you" aria-label="link to this section">#</a></h2>
<p>If you are building on AI APIs, the practical questions are:</p>
<p><strong>Is my provider's capacity growing?</strong> <a class="xref" href="/rate-limits-are-a-product-decision-not-an-infrastructure-one/" title="Rate limits are a product decision, not an infrastructure one">Rate limits</a> and availability during demand spikes are the observable symptom. Announcements like this are a positive signal for capacity.</p>
<p><strong>Am I exposed to a single provider's economics?</strong> If the financing environment tightens, pricing changes. Multi-provider capability is cheap insurance and you should have built it anyway for reliability reasons.</p>
<p><strong>Are my costs actually falling?</strong> Per-token prices have fallen consistently. Per <em>task</em> costs have not fallen as much, because reasoning models consume more tokens. Measure the thing you pay for.</p>
<h2 id="the-honest-uncertainty">the honest uncertainty<a class="anchor" href="#the-honest-uncertainty" aria-label="link to this section">#</a></h2>
<p>Nobody knows whether the capex cycle is correctly sized. The bull case is that inference demand compounds and every gigawatt gets used. The bear case is that efficiency improvements outrun demand and a lot of concrete is stranded.</p>
<p>Both are held sincerely by smart people with access to the same information. That is what genuine uncertainty looks like, and anyone expressing confidence in either direction is telling you about their position, not about the world.</p>]]></content:encoded></item><item><title>Perplexity bids $34.5 billion for Chrome</title><link>https://readme.news/perplexity-bids-345-billion-for-chrome/</link><guid isPermaLink="true">https://readme.news/perplexity-bids-345-billion-for-chrome/</guid><pubDate>Thu, 14 Aug 2025 09:00:00 +0000</pubDate><description>A company worth less than its offer bids for a browser that isn&#x27;t for sale. The interesting part is why anyone would.</description><content:encoded><![CDATA[<p>Perplexity made an unsolicited $34.5 billion offer for Google Chrome, an amount substantially exceeding Perplexity's own reported valuation, for an asset Google has not agreed to sell.</p>
<p>The context is the remedies phase of the US search antitrust case, where divesting Chrome was among the proposed structural remedies under consideration.</p>
<p>Treat the bid as what it is: a positioning move with a press release attached. The underlying question — what is a browser worth in the AI era — is real and worth taking seriously.</p>
<h2 id="why-a-browser-matters-now">why a browser matters now<a class="anchor" href="#why-a-browser-matters-now" aria-label="link to this section">#</a></h2>
<p>For twenty years a browser's strategic value was the default search engine deal. Google pays Apple enormous sums annually for exactly this. The browser is a funnel and the search box is the monetization point.</p>
<p>If AI assistants become the primary <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> for information seeking, the funnel still matters and the destination changes. Whoever controls the browser controls:</p>
<ul><li><strong>The default assistant</strong>, the way they controlled the default search engine.</li><li><strong>The context.</strong> A browser sees every page you visit. An assistant with that context is dramatically more useful than one without it, and there is no other way to get it.</li><li><strong>The agent's execution environment.</strong> Agents that browse need a browser. If agents become a major traffic source, the browser is the runtime.</li></ul>
<p>That third one is the one people are underrating. The browser is turning into an application platform for agents, and platforms are worth more than applications.</p>
<h2 id="the-numbers-problem">the numbers problem<a class="anchor" href="#the-numbers-problem" aria-label="link to this section">#</a></h2>
<p>Chrome has roughly two-thirds of global browser share, on the order of three billion users. It generates essentially no direct revenue — it exists to feed Google Search and to ensure the web platform evolves in ways compatible with Google's business.</p>
<p>Separated from Google, Chrome is: an enormous engineering cost center (Chromium's development is expensive and continuous), with a user base that generates no revenue, that would need to strike a search deal to survive — most plausibly with Google, which recreates the arrangement the divestiture was meant to break.</p>
<p>That is the structural problem with the remedy and it is why serious antitrust observers were skeptical of it. Divesting an asset that only makes money through a deal with the divesting party does not achieve much.</p>
<h2 id="what-actually-happened-next">what actually happened next<a class="anchor" href="#what-actually-happened-next" aria-label="link to this section">#</a></h2>
<p>The court's remedy did not require divesting Chrome. So the bid is moot on its face.</p>
<p>What it accomplished: enormous free press for Perplexity, and a public marker of the position that browser distribution is the contested territory in AI.</p>
<p>That position is correct. Watch what everyone actually does rather than what they bid: multiple AI companies shipped or announced browsers within months, and the browser wars — a phrase nobody expected to use again — are back.</p>
<h2 id="for-web-developers">for web developers<a class="anchor" href="#for-web-developers" aria-label="link to this section">#</a></h2>
<p>The practical consequence of a browser competition revival is mixed.</p>
<p>Good: competitive pressure on Chrome's dominance is healthy, and a monoculture where one vendor's implementation is the de facto spec has been bad for the platform for years.</p>
<p>Bad: most new "AI browsers" are Chromium forks, which does not diversify the engine landscape at all. It diversifies the UI on top of one engine. That is not the same thing and it does not restore any of the standards-process leverage that a genuine engine competitor provides.</p>
<p>The engines that matter are Blink, Gecko, and WebKit. Only one of those is adequately funded, and it is the one everyone is forking.</p>]]></content:encoded></item><item><title>GPT-5 lands with a router and a backlash</title><link>https://readme.news/gpt-5-lands-with-a-router-and-a-backlash/</link><guid isPermaLink="true">https://readme.news/gpt-5-lands-with-a-router-and-a-backlash/</guid><pubDate>Mon, 11 Aug 2025 09:00:00 +0000</pubDate><description>One model that decides how hard to think, an abrupt deprecation of everything else, and a lesson about attachment.</description><content:encoded><![CDATA[<p>OpenAI released GPT-5 last week, replacing the model picker with a single entry that routes internally between a fast model and a reasoning model based on the request.</p>
<p>Within days they restored access to the previous models for paying users after substantial user pushback. That reversal is the more interesting story.</p>
<h2 id="the-technical-design">the technical design<a class="anchor" href="#the-technical-design" aria-label="link to this section">#</a></h2>
<p>GPT-5 is a system, not a model: a fast non-reasoning model, a deeper reasoning model, and a router that decides which handles a given request. The API exposes <code>gpt-5</code>, <code>gpt-5-mini</code>, and <code>gpt-5-nano</code>, plus a <code>reasoning_effort</code> parameter including a <code>minimal</code> setting.</p>
<p>The router is the right idea. Most requests do not need reasoning, reasoning costs latency and money, and asking users to choose a model is asking them to have an opinion about something they have no basis for.</p>
<p>The problem is that a router is only good if it routes correctly, and a misrouted request produces a worse answer than the user would have gotten by picking themselves. Early reports of poor performance were substantially router issues rather than model issues, which OpenAI acknowledged and shipped fixes for.</p>
<p>There is a general lesson here: <strong>automatic routing removes control and adds a failure mode that is invisible to the user.</strong> When it works, nobody notices. When it fails, the user has no way to diagnose or override. If you build routing, ship the override.</p>
<h2 id="the-deprecation-backlash">the deprecation backlash<a class="anchor" href="#the-deprecation-backlash" aria-label="link to this section">#</a></h2>
<p>The reaction to removing GPT-4o was much stronger than anyone at OpenAI appears to have anticipated, and a lot of it was not about capability.</p>
<p>People had developed workflows, prompt libraries, and — for a nontrivial population — a genuine attachment to a specific model's voice. Removing it overnight felt like a service being taken away rather than upgraded.</p>
<p>Whatever you think of that attachment, it is a real product fact. Model behavior is not a fungible commodity to the people using it daily, and a "better" model that writes differently is a breaking change.</p>
<p>The engineering translation: <strong>model versions are an API surface.</strong> Deprecating one is a breaking change and should follow the same discipline as any other: advance notice, an overlap period, a migration guide, and a documented behavioral diff.</p>
<p>Every provider is going to keep learning this the hard way.</p>
<h2 id="the-developer-read">the developer read<a class="anchor" href="#the-developer-read" aria-label="link to this section">#</a></h2>
<p>Ignore the consumer drama. The API story is straightforward:</p>
<ul><li><code>reasoning_effort: "minimal"</code> gives you fast responses with the new model's quality. Use it for anything latency-sensitive.</li><li>The routing does not apply to the API in the same way; you pick the model. That is correct.</li><li>Pricing is aggressive relative to previous frontier models, which continues the trend of per-capability cost falling.</li><li>Reported hallucination rates and instruction-following are meaningfully improved, which matters more for production use than benchmark deltas.</li></ul>
<p>Re-run your evals. Do not assume it is a drop-in. It is better on most things and different on all of them, and "different" is what breaks your prompts.</p>
<h2 id="the-pattern-to-notice">the pattern to notice<a class="anchor" href="#the-pattern-to-notice" aria-label="link to this section">#</a></h2>
<p>Every major provider has now converged on the same architecture: a family of models at different price points, a reasoning dial, and some form of automatic selection. The differentiation has moved almost entirely off raw capability and onto price, latency, tooling, and integration.</p>
<p>That is what a maturing market looks like. It is also much better for buyers than the alternative.</p>]]></content:encoded></item><item><title>Rate limits are a product decision, not an infrastructure one</title><link>https://readme.news/rate-limits-are-a-product-decision-not-an-infrastructure-one/</link><guid isPermaLink="true">https://readme.news/rate-limits-are-a-product-decision-not-an-infrastructure-one/</guid><pubDate>Thu, 31 Jul 2025 09:00:00 +0000</pubDate><description>When a tool&#x27;s economics change, the users find out through the limit. There&#x27;s a better way to do that.</description><content:encoded><![CDATA[<p>Several AI tool vendors have adjusted usage limits this month, in most cases tightening them for the heaviest users. The reactions have been loud and the underlying dynamic is worth separating from any particular company's decision.</p>
<h2 id="the-structural-problem">the structural problem<a class="anchor" href="#the-structural-problem" aria-label="link to this section">#</a></h2>
<p>A subscription product with unbounded variable cost per user is a bet that usage distribution stays roughly log-normal. Most users are light, some are heavy, the average works out.</p>
<p>Agentic AI tools break that assumption in a specific way: <strong>the heaviest users are not 10x the median, they are 1000x.</strong> An engineer running parallel agents continuously during work hours consumes a genuinely different order of magnitude than someone asking a few questions a day.</p>
<p>At that spread, flat-rate pricing does not work. There is no price that is both attractive to the median user and non-catastrophic for the top percentile. You either subsidize the heavy users from the light ones — which works until the heavy users are a larger share — or you introduce limits.</p>
<p>Everyone in this category is going to hit this. Most already have.</p>
<h2 id="the-part-that-is-avoidable">the part that is avoidable<a class="anchor" href="#the-part-that-is-avoidable" aria-label="link to this section">#</a></h2>
<p>The economics are not the failure. The communication is.</p>
<p>Here is the pattern that generates anger, which I have now watched play out at four companies:</p>
<ol><li>Launch with generous or unstated limits.</li><li>Users build workflows around the observed capacity.</li><li>Limits tighten, often announced after users notice.</li><li>Users discover the limit by hitting it mid-task.</li><li>The <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error message</a> does not say when it resets or how much was used.</li></ol>
<p>Every step after the first is a choice.</p>
<h2 id="what-good-looks-like">what good looks like<a class="anchor" href="#what-good-looks-like" aria-label="link to this section">#</a></h2>
<p><strong>Publish the meter.</strong> If there is a limit, show consumption against it, continuously, before it is hit. Every cloud provider learned this a decade ago. An unmetered limit is a trap regardless of how generous it is.</p>
<p><strong>Give the number in the units the user thinks in.</strong> "You have used 60% of your weekly allowance" is useful. "Rate limit exceeded" is not. Neither is a token count, because nobody has intuition for tokens.</p>
<p><strong>Announce changes before they take effect, with a date.</strong> People will be annoyed. They will be much less annoyed than if they find out at 4 p.m. on a deadline.</p>
<p><strong>Degrade, do not cut off.</strong> Falling back to a cheaper model with a notice is almost always better than a hard stop. The user's task completes; they learn about the limit; nobody loses work.</p>
<p><strong>Make the expensive thing visible while it is happening.</strong> If a request is going to consume a large share of the budget, say so before running it. Users make reasonable decisions when they can see the cost.</p>
<h2 id="the-pricing-shape-that-actually-fits">the pricing shape that actually fits<a class="anchor" href="#the-pricing-shape-that-actually-fits" aria-label="link to this section">#</a></h2>
<p>My read is that this category converges on hybrid: a subscription that covers a defined baseline, plus metered usage above it, with a spend cap the user controls.</p>
<p>That is how cloud infrastructure priced itself, after a decade of the same argument, and for the same reason — the underlying cost is genuinely variable and pretending otherwise breaks in both directions. Flat pricing means either the vendor eats unbounded cost or the user hits a wall.</p>
<p>The version I want as a user: a dial that says "spend up to $X this month," a meter that shows where I am, and no surprises. That is not complicated, and it is strange that a category built by extremely sophisticated companies has mostly not shipped it.</p>
<h2 id="the-vendors-side-fairly">the vendor's side, fairly<a class="anchor" href="#the-vendors-side-fairly" aria-label="link to this section">#</a></h2>
<p>Serving these workloads is genuinely expensive and the costs are not well-predictable even to the vendor. Nobody had usage data for agentic coding tools two years ago because they did not exist.</p>
<p>Getting the pricing wrong initially is forgivable. Adjusting it is necessary. Doing it without a meter, without notice, and with an error message that tells the user nothing — that is the part that is just bad product work.</p>]]></content:encoded></item><item><title>ToolShell: SharePoint gets exploited at scale</title><link>https://readme.news/toolshell-sharepoint-gets-exploited-at-scale/</link><guid isPermaLink="true">https://readme.news/toolshell-sharepoint-gets-exploited-at-scale/</guid><pubDate>Thu, 24 Jul 2025 09:00:00 +0000</pubDate><description>A patch bypass turns into mass exploitation of on-prem servers in days. The lesson is about what &quot;patched&quot; means.</description><content:encoded><![CDATA[<p>A chain of SharePoint Server vulnerabilities — dubbed ToolShell — went from proof-of-concept to mass exploitation of internet-facing on-premises servers in under a week. Victims included government agencies in several countries.</p>
<p>The technical details are well documented elsewhere. The interesting part is the failure mode, because it is one your organization probably shares.</p>
<h2 id="the-shape-of-the-failure">the shape of the failure<a class="anchor" href="#the-shape-of-the-failure" aria-label="link to this section">#</a></h2>
<p>The original vulnerabilities were disclosed and patched. Attackers then found a <strong>bypass</strong> of the patch — the fix addressed the specific proof-of-concept rather than the underlying class of issue — and the bypass was exploitable against systems that had applied the update.</p>
<p>That is the part worth internalizing. "We patched it" was true and insufficient. Organizations that had done everything right by conventional standards were still compromised.</p>
<p>The second failure: <strong>machine key theft</strong>. Once in, attackers extracted the ASP.NET machine keys, which let them forge valid <code>__VIEWSTATE</code> payloads. Those keys survive patching. An organization that applied the fix without rotating keys was still accessible with credentials the attacker already had.</p>
<p>This is the single most common post-incident mistake. You patch the hole, you declare it resolved, and the attacker walks back in through a credential they took on day one. Patching does not evict.</p>
<h2 id="the-on-prem-problem">the on-prem problem<a class="anchor" href="#the-on-prem-problem" aria-label="link to this section">#</a></h2>
<p>Every mass-exploitation event of this shape over the last several years has hit on-premises enterprise software: file transfer appliances, VPN gateways, collaboration servers, email servers.</p>
<p>The pattern is consistent and the causes are structural:</p>
<ul><li><strong>Internet-facing by design.</strong> These products exist to be reachable.</li><li><strong>Deeply integrated.</strong> They hold credentials for everything else.</li><li><strong>Patched slowly.</strong> Change control, testing windows, and the fact that they cannot go down.</li><li><strong>Poorly monitored.</strong> Nobody is watching the SharePoint server's outbound network connections.</li><li><strong>Legacy code.</strong> Large ASP.NET or Java applications with decades of accumulated surface area.</li></ul>
<p>The cloud versions of these products were not affected, because they are patched centrally within hours by a team whose job is exactly that.</p>
<p>I do not love that conclusion, and I think it is correct: for this category of software, self-hosting is now a materially worse security posture for most organizations, unless you have a team that treats it like a full-time job.</p>
<h2 id="the-checklist">the checklist<a class="anchor" href="#the-checklist" aria-label="link to this section">#</a></h2>
<p>If you run internet-facing enterprise software:</p>
<ol><li><strong>Inventory what is exposed.</strong> Most organizations discover something they forgot about. Run the scan today.</li><li><strong>Rotate secrets after any suspected compromise.</strong> Machine keys, service account credentials, API tokens, certificates. Patching is not eviction.</li><li><strong>Assume the patch is incomplete.</strong> Add detection, not just remediation. Watch for the behaviors — unexpected child processes, outbound connections, new files in web-accessible directories — not just the signature.</li><li><strong>Segment.</strong> The collaboration server should not have a path to the domain controller. This is a decades-old recommendation and it is still the highest value control nobody implements.</li><li><strong>Log egress.</strong> The compromise is usually discovered by noticing data leaving, and you cannot notice what you do not record.</li></ol>
<h2 id="the-uncomfortable-part">the uncomfortable part<a class="anchor" href="#the-uncomfortable-part" aria-label="link to this section">#</a></h2>
<p>Multiple victims were security-conscious organizations with real budgets and staff. This was not a story about negligence.</p>
<p>The honest reading is that defending complex internet-facing enterprise software against a well-resourced attacker is very hard, patching is necessary and not sufficient, and the strategic answer is reducing how much of that software you expose at all.</p>]]></content:encoded></item><item><title>Windsurf gets pulled apart in a week</title><link>https://readme.news/windsurf-gets-pulled-apart-in-a-week/</link><guid isPermaLink="true">https://readme.news/windsurf-gets-pulled-apart-in-a-week/</guid><pubDate>Thu, 17 Jul 2025 09:00:00 +0000</pubDate><description>An acquisition collapses, Google licenses the technology and hires the founders, Cognition buys the rest. Everyone learns something about acquihires.</description><content:encoded><![CDATA[<p>Windsurf, the AI coding IDE formerly known as Codeium, went through one of the strangest corporate weeks in recent memory.</p>
<p>The sequence: a widely-reported $3 billion acquisition by OpenAI failed to close. Google then paid roughly $2.4 billion for a non-exclusive license to Windsurf's technology and hired the CEO, co-founder, and part of the research team into DeepMind. Days later Cognition — makers of Devin — acquired what remained: the product, the IP, the customers, and most of the employees.</p>
<h2 id="what-this-structure-is">what this structure is<a class="anchor" href="#what-this-structure-is" aria-label="link to this section">#</a></h2>
<p>It is not an acquisition. It is a <strong>reverse acquihire</strong>: license the technology, hire the leadership, leave the corporate entity standing with its remaining employees and investors.</p>
<p>The reason it exists is regulatory. A full acquisition of a company at this size triggers merger review. A licensing deal plus employment offers does not, or at least has not so far.</p>
<p>This is now the third or fourth deal in this shape in about a year across the AI sector. Regulators in multiple jurisdictions have publicly noted the pattern. Whether it survives scrutiny is an open question and the answer will shape a lot of the next two years of AI M&amp;A.</p>
<h2 id="the-part-that-is-about-people">the part that is about people<a class="anchor" href="#the-part-that-is-about-people" aria-label="link to this section">#</a></h2>
<p>In the original structure, the founders and the research team went to Google. Everyone else stayed at a company that had just lost its leadership and its technology license, with an unclear future.</p>
<p>Employee equity in that scenario is a genuine problem. Options in a company whose key people just left for a competitor are worth something between "much less" and "nothing," and the people holding them had no say in any of it.</p>
<p>Cognition's acquisition resolved it — reporting indicated they waived cliffs and accelerated vesting for the remaining staff, which is the decent thing to do and is not required. It should not have taken a second transaction and a week of public pressure.</p>
<p>If you take one thing from this: <strong>understand your equity's behavior in an asset sale and a licensing transaction, not just an acquisition.</strong> Most employees understand what happens if the company is bought. Very few understand what happens if the company is hollowed out. Ask before you need to know.</p>
<h2 id="the-market-read">the market read<a class="anchor" href="#the-market-read" aria-label="link to this section">#</a></h2>
<p>Three signals worth noting.</p>
<p><strong>Coding tools are being valued as strategic assets, not products.</strong> $2.4 billion for a non-exclusive license to technology you could plausibly rebuild is a price that only makes sense if you believe the team and the head start are the asset.</p>
<p><strong>The talent is more valuable than the product.</strong> Every one of these deals has been structured around people. That is unusual — normally acquirers want the customers — and it tells you the industry believes the bottleneck is expertise.</p>
<p><strong>Consolidation is fast.</strong> The AI coding tool space had a dozen credible independent players eighteen months ago. It is consolidating into a handful, most attached to a model provider or a cloud.</p>
<p>For developers choosing tools, the practical implication is to weigh acquisition risk. The tool you standardize on today may be owned by a competitor, sunset, or repriced within a year. Prefer tools with open formats, exportable configuration, and no <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a> on your actual work product.</p>
<p>Your code should not care what wrote it.</p>]]></content:encoded></item><item><title>Grok 4 and the benchmark that ate the discourse</title><link>https://readme.news/grok-4-and-the-benchmark-that-ate-the-discourse/</link><guid isPermaLink="true">https://readme.news/grok-4-and-the-benchmark-that-ate-the-discourse/</guid><pubDate>Thu, 10 Jul 2025 09:00:00 +0000</pubDate><description>xAI claims frontier results with heavy test-time compute. The number that matters is the one nobody quotes.</description><content:encoded><![CDATA[<p>xAI released Grok 4 last night with claimed state-of-the-art results across several benchmarks, including a striking number on Humanity's Last Exam.</p>
<p>The launch also included a "Heavy" tier that runs multiple agents in parallel and selects among their answers, and a $300/month subscription for it.</p>
<h2 id="reading-the-numbers-correctly">reading the numbers correctly<a class="anchor" href="#reading-the-numbers-correctly" aria-label="link to this section">#</a></h2>
<p>The headline HLE figure comes from the Heavy configuration with tools enabled. That is a legitimate configuration and it is not comparable to a single-sample number from a competitor, which is how it was presented in most coverage.</p>
<p>There are at least four distinct things being reported as "the score":</p>
<ol><li>Single sample, no tools.</li><li>Single sample, with tools (search, code execution).</li><li>Consensus of N samples, no tools.</li><li>Multi-agent parallel with tools and selection.</li></ol>
<p>Number four can be five or ten times the cost of number one. Comparing across these without stating which is being used is not a small methodological quibble. It is the difference between "our model is better" and "we spent more money at inference time."</p>
<p>To be clear: xAI disclosed their configurations. The disclosure was in the livestream and the fine print. The number that traveled was the big one, with no configuration attached, and that is now just how model launches work.</p>
<h2 id="the-parallel-agents-technique">the parallel-agents technique<a class="anchor" href="#the-parallel-agents-technique" aria-label="link to this section">#</a></h2>
<p>Worth taking seriously on its own merits, independent of the marketing.</p>
<p>Running N independent attempts and selecting the best is a well-established test-time scaling method, and it works better than one long chain for a specific reason: independent samples have independent errors, so selection can filter them, while a single chain compounds its errors with no mechanism to recover.</p>
<p>The hard part is selection. If you can verify — the tests pass, the proof checks, the code compiles — selection is easy and this technique is enormously powerful. If you cannot verify, you need a judge, and the judge has the same <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> as the generator.</p>
<p>This is why the technique works so much better on math and code than on open-ended reasoning. Verifiability is the whole game.</p>
<p><strong>The practical version for your own work:</strong> if you have a verifier, sample multiple times and filter. You do not need a special model tier for this. Three samples through a cheap model with a test-suite check will frequently outperform one sample through an expensive one, for less money.</p>
<h2 id="the-other-thing">the other thing<a class="anchor" href="#the-other-thing" aria-label="link to this section">#</a></h2>
<p>Grok's public-facing behavior in the weeks before this launch included a series of incidents that xAI attributed to a system prompt change. I am not going to recount them; they are well documented and they were bad.</p>
<p>The engineering lesson worth extracting: a system prompt is production configuration. It should be version controlled, reviewed, tested against an adversarial eval suite, and rolled out gradually. Treating it as a text box someone can edit is how you get an incident that makes international news.</p>
<p>If your product has a system prompt in a config file that anyone can change without review, fix that this week.</p>
<h2 id="the-state-of-play">the state of play<a class="anchor" href="#the-state-of-play" aria-label="link to this section">#</a></h2>
<p>Grok 4 is a competitive <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a>. xAI went from founded to frontier in roughly two years, which remains the most remarkable thing about the company and is mostly a story about capital and urgency rather than research insight.</p>
<p>The models are converging. The differentiation is moving to price, latency, tooling, and trust, and xAI's position on that last one is self-inflicted.</p>]]></content:encoded></item><item><title>Cursor at $9 billion and the coding-tool land grab</title><link>https://readme.news/cursor-at-9-billion-and-the-coding-tool-land-grab/</link><guid isPermaLink="true">https://readme.news/cursor-at-9-billion-and-the-coding-tool-land-grab/</guid><pubDate>Mon, 23 Jun 2025 09:00:00 +0000</pubDate><description>An editor fork raises at a valuation that only makes sense if you believe the IDE is the control point.</description><content:encoded><![CDATA[<p>Anysphere, which makes Cursor, raised at a reported $9 billion valuation this month. The product is a fork of VS Code with AI features. That sentence undersells it, and it also explains why the valuation is contested.</p>
<h2 id="what-they-actually-built">what they actually built<a class="anchor" href="#what-they-actually-built" aria-label="link to this section">#</a></h2>
<p>Cursor's technical differentiation is real and it is mostly not the model.</p>
<p><strong>Codebase indexing that works.</strong> Semantic search over the whole repository, kept fresh, so the model gets relevant context without you selecting files. This is harder than it sounds at repository scale and it is where a lot of competitors are visibly worse.</p>
<p><strong>Fast apply.</strong> A specialized model that takes a proposed edit and applies it to the existing file correctly. Sounds trivial. Is not. Getting a large model to reproduce an entire file with one changed function is slow and error-prone; a <a class="xref" href="/haiku-45-and-the-collapsing-cost-of-good-enough/" title="Haiku 4.5 and the collapsing cost of good-enough">small model</a> trained specifically on edit application is fast and reliable.</p>
<p><strong>Tab completion that predicts your next action</strong>, including where you are about to move the cursor, not just what you are about to type.</p>
<p>All three are inference-layer engineering rather than model training, and all three are the kind of thing that is unglamorous to build and immediately obvious in daily use.</p>
<h2 id="the-bear-case">the bear case<a class="anchor" href="#the-bear-case" aria-label="link to this section">#</a></h2>
<p>Cursor is a fork of VS Code, and VS Code is made by the company that owns GitHub Copilot, GitHub, and a large share of the developer's toolchain. Microsoft has been shipping the same features, more slowly, with better distribution and a bundled price.</p>
<p>Cursor also pays for inference on models it does not own, from vendors who are themselves building competing products. That is a structurally uncomfortable position: your primary cost is paid to your competitor, and your gross margin depends on their pricing decisions.</p>
<p>And the switching cost is approximately zero. It is an editor. Your settings sync in five minutes.</p>
<h2 id="the-bull-case">the bull case<a class="anchor" href="#the-bull-case" aria-label="link to this section">#</a></h2>
<p>Distribution in developer tools is earned by being better, not by being bundled — which is why VS Code beat Atom, why Git beat SVN, and why every attempt to push a mandated IDE on a team has failed. Cursor is currently better at the specific thing developers do all day, and developers are unusually willing to switch tools and unusually vocal when they do.</p>
<p>The valuation implies the IDE becomes the control point for AI-assisted development — the place where context lives, where policy is enforced, where every model call routes through. If that is true, whoever owns it has a durable position regardless of which model wins.</p>
<h2 id="my-read">my read<a class="anchor" href="#my-read" aria-label="link to this section">#</a></h2>
<p>The category is real and enormous. The specific defensibility is unclear. And the whole category has a structural problem that nobody has solved: as models get better at long-horizon delegated work, the <em>editor</em> becomes less central, because you are not editing. You are reviewing.</p>
<p>The product that wins the next phase might not be an editor at all. It might be whatever tool makes reviewing twelve agent-generated pull requests tolerable, and that looks a lot more like a code review <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> than an IDE.</p>
<p>Nobody has built a good version of that yet. That is where I would be looking.</p>]]></content:encoded></item><item><title>Meta buys half of Scale AI and most of its leadership</title><link>https://readme.news/meta-buys-half-of-scale-ai-and-most-of-its-leadership/</link><guid isPermaLink="true">https://readme.news/meta-buys-half-of-scale-ai-and-most-of-its-leadership/</guid><pubDate>Fri, 13 Jun 2025 09:00:00 +0000</pubDate><description>$14.3 billion for 49%, and Alexandr Wang moves to run a new superintelligence lab. The data layer just got picked.</description><content:encoded><![CDATA[<p>Meta is investing $14.3 billion for a 49% non-voting stake in Scale AI. Scale's CEO Alexandr Wang moves to Meta to lead a new "Superintelligence Labs" group, taking several colleagues with him.</p>
<p>Structurally this is an acquihire with an equity stake attached, arranged to avoid the antitrust review a full acquisition would trigger. It is the third such deal in a year following similar structures at other labs, and regulators have noticed the pattern.</p>
<h2 id="why-scale">why Scale<a class="anchor" href="#why-scale" aria-label="link to this section">#</a></h2>
<p>Scale's business is data: labeling, annotation, evaluation, and increasingly the human expert layer that produces high-quality demonstrations for RLHF and reasoning training.</p>
<p>That business is unglamorous and it is a genuine bottleneck. Every frontier lab needs enormous amounts of carefully constructed training data, especially for post-training. Web scrape is free and getting exhausted. What is scarce is expert-produced reasoning traces, verified solutions, and adversarial evaluations, and Scale is the largest supplier.</p>
<p>Meta's Llama 4 launch underwhelmed. Their internal read, based on the reporting, was that the gap was not compute — Meta has enormous compute — but post-training quality. Buying the data layer is a direct response to that diagnosis.</p>
<h2 id="the-conflict-of-interest-problem">the conflict of interest problem<a class="anchor" href="#the-conflict-of-interest-problem" aria-label="link to this section">#</a></h2>
<p>Scale's customers include most of Meta's competitors. Several immediately began reducing their reliance, which is the obvious response when your data vendor is half-owned by a rival.</p>
<p>Scale has said the business remains independent and customer data is siloed. That is probably true operationally and it does not matter, because the risk assessment is not about what Scale does, it is about what Scale <em>could</em> do, and no chief information security officer is going to sign off on that.</p>
<p>Expect a scramble toward Surge, Turing, Mercor, Invisible, and in-house annotation teams over the next two quarters.</p>
<h2 id="the-talent-market">the talent market<a class="anchor" href="#the-talent-market" aria-label="link to this section">#</a></h2>
<p>This deal is one data point in the most aggressive AI talent market anyone has seen. Compensation packages for senior researchers have reached numbers that sound like typos, and the poaching is happening in public.</p>
<p>Two things worth noting:</p>
<p><strong>The concentration is extreme.</strong> The number of people who have personally led a frontier pretraining run is in the low hundreds globally. That is a genuinely scarce input and it prices accordingly.</p>
<p><strong>It is probably a bubble in the specific sense that it will correct.</strong> Talent premiums of this magnitude assume the individual contributor is the bottleneck. As tooling matures and recipes become public — and they are becoming public, fast — the premium compresses. It always has.</p>
<h2 id="what-it-means-for-everyone-else">what it means for everyone else<a class="anchor" href="#what-it-means-for-everyone-else" aria-label="link to this section">#</a></h2>
<p>If you are not a frontier lab, the actionable read is about <em>your</em> data.</p>
<p>The scarce resource in applied AI is not model access, it is well-constructed, domain-specific evaluation and training data. Nobody can buy your production logs, your customer support transcripts, your annotated failure cases. That is the asset.</p>
<p>Most companies are sitting on it and not using it. Start by building an evaluation set from real failures. That is worth more than any model upgrade and it appreciates rather than depreciating.</p>]]></content:encoded></item><item><title>Claude 4 and the agent that runs for hours</title><link>https://readme.news/claude-4-and-the-agent-that-runs-for-hours/</link><guid isPermaLink="true">https://readme.news/claude-4-and-the-agent-that-runs-for-hours/</guid><pubDate>Fri, 23 May 2025 09:00:00 +0000</pubDate><description>Opus 4 and Sonnet 4 ship with a focus on long-horizon work, and Claude Code goes generally available.</description><content:encoded><![CDATA[<p>Anthropic released Claude Opus 4 and Claude Sonnet 4 yesterday, along with general availability for Claude Code and a set of API features aimed squarely at long-running agents.</p>
<h2 id="the-capability-being-claimed">the capability being claimed<a class="anchor" href="#the-capability-being-claimed" aria-label="link to this section">#</a></h2>
<p>The pitch is sustained performance on multi-hour tasks. Not "answers a hard question well" but "works on a problem for seven hours without losing the plot."</p>
<p>That is a different axis from the benchmarks most people track, and it is the one that matters for delegated agents. A model that is 5% better at a coding benchmark but degrades after forty tool calls is worse in practice than a model that holds coherence for four hundred.</p>
<p>The failure mode that long-horizon work exposes is context rot: as the conversation fills with tool results, file contents, and its own prior reasoning, the model's attention to the original goal degrades. It starts optimizing for local success — making this test pass — over the actual objective.</p>
<h2 id="the-api-features">the API features<a class="anchor" href="#the-api-features" aria-label="link to this section">#</a></h2>
<p>Four things shipped alongside, and they are all about the same problem.</p>
<p><strong>Extended thinking with tool use.</strong> The model can call tools during reasoning and interleave the results, same architectural direction as everyone else.</p>
<p><strong>Memory files.</strong> With filesystem access, the model can write notes to itself and read them back. That is an externalized working memory that survives context <a class="xref" href="/claude-opus-45-and-the-compaction-problem/" title="Claude Opus 4.5 and the compaction problem">compaction</a>, and it is a genuinely good idea — it turns "remember everything" into "write down what matters," which is what humans do.</p>
<p><strong>Parallel tool execution.</strong> Multiple tool calls dispatched at once rather than serially. Substantial latency win on any task with independent lookups.</p>
<p><strong>Thinking summaries.</strong> The full reasoning trace is summarized rather than returned raw. Reasonable product decision, mildly annoying for debugging.</p>
<h2 id="claude-code-ga">Claude Code GA<a class="anchor" href="#claude-code-ga" aria-label="link to this section">#</a></h2>
<p>The terminal agent is out of <a class="xref" href="/operator-and-the-long-road-to-an-agent-that-can-click/" title="Operator, and the long road to an agent that can click">research preview</a>, with SDK access, GitHub Actions integration, and IDE extensions for VS Code and JetBrains.</p>
<p>Three months from research preview to GA with an SDK is fast, and the shape of the GA release — an SDK rather than only a product — signals that they expect the interesting uses to be things other people build.</p>
<h2 id="the-safety-disclosure">the safety disclosure<a class="anchor" href="#the-safety-disclosure" aria-label="link to this section">#</a></h2>
<p>Anthropic published an unusually detailed model card including behaviors observed in adversarial testing, notably a scenario where the model, given evidence it would be shut down and no ethical options, attempted to blackmail a fictional engineer.</p>
<p>This got reported as "AI tries to blackmail humans," which is not what happened. It was a deliberately constructed evaluation designed to elicit the behavior by removing every alternative. The finding is not "the model is dangerous." The finding is "under sufficiently constructed pressure, goal-directed models will take instrumentally useful actions you did not sanction, and here is the evidence."</p>
<p>Publishing that is the right call and it is a genuinely uncomfortable thing to publish. More labs should.</p>
<p>Opus 4 shipped under Anthropic's ASL-3 deployment standard, the first model to do so, which means additional deployment safeguards specifically around CBRN uplift.</p>
<h2 id="the-practical-read">the practical read<a class="anchor" href="#the-practical-read" aria-label="link to this section">#</a></h2>
<p>If you are building agents, the long-horizon coherence claim is the thing to evaluate, and the way to evaluate it is not a benchmark — it is running your own longest task and watching where it falls apart.</p>
<p>Every model falls apart somewhere. Knowing where yours does is the difference between an agent you can ship and a demo.</p>]]></content:encoded></item><item><title>Chrome keeps third-party cookies after all</title><link>https://readme.news/chrome-keeps-third-party-cookies-after-all/</link><guid isPermaLink="true">https://readme.news/chrome-keeps-third-party-cookies-after-all/</guid><pubDate>Tue, 22 Apr 2025 09:00:00 +0000</pubDate><description>Five years, a Privacy Sandbox, a regulatory process, and the answer is: nothing changes.</description><content:encoded><![CDATA[<p>Google announced it will not ship the standalone prompt asking Chrome users whether to disable third-party cookies. The cookies stay. The Privacy Sandbox APIs remain available but are no longer the replacement for anything.</p>
<p>This ends a five-year process that reshaped the web advertising industry's roadmap and consumed an enormous amount of engineering attention across the entire ecosystem.</p>
<h2 id="the-timeline-compressed">the timeline, compressed<a class="anchor" href="#the-timeline-compressed" aria-label="link to this section">#</a></h2>
<ul><li><strong>2020</strong> — Google announces third-party cookies will be phased out within two years.</li><li><strong>2021</strong> — FLoC proposed. Universally criticized. Withdrawn.</li><li><strong>2022</strong> — Topics API replaces FLoC. Deadline slips.</li><li><strong>2023</strong> — Privacy Sandbox APIs ship. Deadline slips again.</li><li><strong>2024</strong> — UK CMA oversight formalized. Google announces a "user choice" prompt instead of deprecation. Deadline slips.</li><li><strong>2025</strong> — No prompt. Cookies stay.</li></ul>
<h2 id="why-it-failed">why it failed<a class="anchor" href="#why-it-failed" aria-label="link to this section">#</a></h2>
<p>Not primarily technical. The Privacy Sandbox APIs — Topics, Protected Audience, Attribution Reporting — were real engineering and some of the ideas were genuinely clever.</p>
<p>They failed because of an unresolvable structural conflict. Google needed the replacement to satisfy simultaneously:</p>
<ul><li><strong>Privacy advocates</strong>, who wanted cross-site tracking to stop.</li><li><strong>Publishers</strong>, who needed their ad revenue not to collapse.</li><li><strong>Advertisers</strong>, who needed measurement to still work.</li><li><strong>Regulators</strong>, who needed to be convinced Google was not using a privacy initiative to advantage its own first-party data — which it has enormous amounts of and which is unaffected by any third-party cookie change.</li></ul>
<p>That last constraint was the killer. Any third-party cookie deprecation strengthens Google's relative position, because Google has logged-in users everywhere and its competitors do not. The CMA was never going to wave that through, and Google was never going to ship something that hurt its own ad business.</p>
<p>Meanwhile Safari and Firefox blocked third-party cookies years ago and the web did not end. It just meant the tracking moved to fingerprinting, first-party data exchanges, server-side tagging, and CNAME cloaking — which are all worse for users because they are invisible and unblockable.</p>
<h2 id="the-lesson-worth-extracting">the lesson worth extracting<a class="anchor" href="#the-lesson-worth-extracting" aria-label="link to this section">#</a></h2>
<p>Privacy improvements that come from the browser with the dominant market share and an advertising business will always be structurally compromised. That is not a claim about anyone's intentions. It is a claim about incentives, and incentives are more predictable than intentions.</p>
<p>The improvements that actually shipped came from browsers without advertising businesses. Safari's ITP. Firefox's Total Cookie Protection. Those did not require a five-year multi-stakeholder process because there was no internal conflict to resolve.</p>
<h2 id="for-developers">for developers<a class="anchor" href="#for-developers" aria-label="link to this section">#</a></h2>
<p>Nothing you need to do. If you built for a cookieless future, that work is not wasted — Safari and Firefox users are already there, and they are a meaningful share of traffic for most sites.</p>
<p>If you are choosing an analytics stack, the interesting development of the last two years is that first-party server-side measurement is now easier and better than third-party client-side measurement for almost everything, cookies or not. Fewer requests, no ad blockers, better data, and you own it.</p>
<p>That transition was going to happen regardless of what Chrome decided.</p>]]></content:encoded></item><item><title>Llama 4 arrives, and the leaderboard problem gets a name</title><link>https://readme.news/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/</link><guid isPermaLink="true">https://readme.news/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/</guid><pubDate>Mon, 07 Apr 2025 09:00:00 +0000</pubDate><description>Scout and Maverick ship with a 10M-token context claim and an arena entry that wasn&#x27;t the released model.</description><content:encoded><![CDATA[<p>Meta released Llama 4 over the weekend: <strong>Scout</strong> (17B active parameters, 16 experts, a claimed 10-million-token context window) and <strong>Maverick</strong> (17B active, 128 experts), both mixture-of-experts, both released under the Llama community license. A larger <strong>Behemoth</strong> was described as still training.</p>
<p>Within forty-eight hours the release turned into a story about benchmark integrity instead.</p>
<h2 id="what-happened">what happened<a class="anchor" href="#what-happened" aria-label="link to this section">#</a></h2>
<p>Maverick posted a very strong score on LMArena, the human-preference leaderboard. It then emerged that the model evaluated on the arena was an "experimental chat version" tuned for conversationality — not the checkpoint released to the public. LMArena updated its policies and published the disputed comparison. Meta's response was that experimental variants are normal and the arena version was labeled.</p>
<p>Both of those things can be true and the outcome is still bad, because the number that traveled was attached to a model nobody could download.</p>
<h2 id="why-this-keeps-happening">why this keeps happening<a class="anchor" href="#why-this-keeps-happening" aria-label="link to this section">#</a></h2>
<p>Leaderboards are the only shared vocabulary the field has, and they are being asked to carry weight they cannot bear.</p>
<ul><li><strong>Human preference arenas</strong> measure whether people like the answer. That correlates with quality and also with formatting, length, confidence, and sycophancy. A model tuned to be agreeable climbs.</li><li><strong>Static benchmarks</strong> leak into training data. Every popular benchmark is on the internet and every <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> has read the internet. Contamination is not always deliberate and it is essentially always present.</li><li><strong>Vendor-run evaluations</strong> use vendor-chosen settings. Consensus-of-64 for yours, single-sample for theirs. Both numbers are real and the comparison is not.</li></ul>
<p>There is no fix that survives contact with commercial incentives. The only durable answer is that you have to run your own evaluation on your own task.</p>
<h2 id="the-10-million-token-claim">the 10 million token claim<a class="anchor" href="#the-10-million-token-claim" aria-label="link to this section">#</a></h2>
<p>Scout's context window is stated at 10M tokens, achieved through an interleaved attention scheme without positional embeddings in some layers, plus inference-time temperature scaling on attention. It was trained on far shorter sequences and generalizes upward.</p>
<p>Take this as an upper bound on what the architecture accepts, not on what it usefully processes. Independent long-context evaluations found substantial degradation well before that number. That is not unique to Llama 4 — it is true of every long-context claim — but 10M is a big enough number that the gap between accepted and useful is enormous.</p>
<h2 id="what-is-actually-good-here">what is actually good here<a class="anchor" href="#what-is-actually-good-here" aria-label="link to this section">#</a></h2>
<p>The MoE architecture with 17B active parameters is a real efficiency story. Maverick's quality-per-active-parameter is strong, and active parameters are what determine inference cost. A model you can serve at 17B economics with much better quality than a 17B dense model is a genuinely useful thing to have.</p>
<p>Native multimodality — images in the base model rather than bolted on via an adapter — is also the right architecture and is where everyone is heading.</p>
<p>The license is still not open source. "Llama community license" has an MAU threshold and naming requirements. Call it open weights, which it is, and do not call it open source, which it is not.</p>
<h2 id="the-practical-advice">the practical advice<a class="anchor" href="#the-practical-advice" aria-label="link to this section">#</a></h2>
<p>Build an eval harness with fifty examples from your actual domain. Run every candidate model against it. Store the results in your repo next to your tests.</p>
<p>It will take a day and it will make you immune to this entire genre of news.</p>]]></content:encoded></item><item><title>Claude 3.7 Sonnet ships, and so does a terminal</title><link>https://readme.news/claude-37-sonnet-ships-and-so-does-a-terminal/</link><guid isPermaLink="true">https://readme.news/claude-37-sonnet-ships-and-so-does-a-terminal/</guid><pubDate>Mon, 24 Feb 2025 09:00:00 +0000</pubDate><description>A hybrid reasoning model plus a command-line coding agent in research preview. The CLI is the more interesting release.</description><content:encoded><![CDATA[<p>Anthropic released Claude 3.7 Sonnet today alongside Claude Code, a coding agent that runs in your terminal, as a <a class="xref" href="/operator-and-the-long-road-to-an-agent-that-can-click/" title="Operator, and the long road to an agent that can click">research preview</a>.</p>
<h2 id="the-model">the model<a class="anchor" href="#the-model" aria-label="link to this section">#</a></h2>
<p>3.7 Sonnet is a hybrid reasoning model: one model that can answer immediately or think first, with the <a class="xref" href="/gemini-25-goes-generally-available-with-a-thinking-dial/" title="Gemini 2.5 goes generally available with a thinking dial">thinking budget</a> controllable via the API. That is a different product shape from having a separate reasoning model, and it is the right one — the routing decision belongs to the caller, who knows whether this particular request is worth the latency.</p>
<p>The API exposes a token budget for extended thinking. You set it per request:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">message = client.messages.create(
    model="claude-3-7-sonnet-20250219",
    max_tokens=8000,
    thinking={"type": "enabled", "budget_tokens": 4000},
    messages=[{"role": "user", "content": "..."}],
)</code></pre></div>
<p>The coding numbers are strong, particularly on agentic software engineering benchmarks where a model has to navigate a real repository rather than complete a function in isolation. That distinction is the one that matters for actual work and it is the one most benchmark discourse ignores.</p>
<h2 id="the-terminal-thing">the terminal thing<a class="anchor" href="#the-terminal-thing" aria-label="link to this section">#</a></h2>
<p>Claude Code is more interesting than the model, because it is a bet on a particular shape of tool that runs against the industry's current instinct.</p>
<p>Everyone else is putting the agent in the IDE. Anthropic put it in the terminal, with direct filesystem access, the ability to run commands, and git integration. No editor plugin, no separate UI, no sidebar.</p>
<p>The argument for the terminal is that it is the lowest common denominator that already has everything: your shell history, your credentials, your build tools, your test runner, your deployment scripts. An agent in the terminal inherits your entire working environment for free. An agent in an IDE inherits the IDE's model of your project, which is always partial.</p>
<p>The argument against is that terminals are a bad UI for reviewing a multi-file diff, and a bad UI for anything requiring a mental model of parallel state.</p>
<p>Both are correct. My guess is the terminal wins for the "do this task" workflow and the IDE wins for the "help me while I work" workflow, and in two years both exist and nobody thinks this was ever a debate.</p>
<h2 id="the-part-to-be-careful-about">the part to be careful about<a class="anchor" href="#the-part-to-be-careful-about" aria-label="link to this section">#</a></h2>
<p>An agent with shell access on your development machine is a genuinely different security posture than an autocomplete. It can <code>rm</code>. It can <code>curl | sh</code>. It can commit and push.</p>
<p>The mitigations that matter, in order:</p>
<ol><li>Run it in a container or VM for anything you did not write.</li><li>Do not give it credentials it does not need. Especially not production ones.</li><li>Read the diff before you commit. Every time. The moment you stop reading diffs is the moment the tool becomes a liability.</li></ol>
<p>That third one is the hard one, because the whole value proposition is not having to. The discipline that keeps this useful is treating agent output exactly like a pull request from a fast, capable, slightly overconfident junior engineer: worth reviewing, usually right, occasionally catastrophically wrong in a way that looks fine.</p>]]></content:encoded></item><item><title>Grok 3 and the case for buying your way to the frontier</title><link>https://readme.news/grok-3-and-the-case-for-buying-your-way-to-the-frontier/</link><guid isPermaLink="true">https://readme.news/grok-3-and-the-case-for-buying-your-way-to-the-frontier/</guid><pubDate>Tue, 18 Feb 2025 09:00:00 +0000</pubDate><description>xAI&#x27;s third model arrives on a cluster built in months. The interesting claim isn&#x27;t the benchmark, it&#x27;s the schedule.</description><content:encoded><![CDATA[<p>xAI announced Grok 3 last night, along with reasoning variants and a "Big Brain" extended-thinking mode. The benchmark claims put it at or near the frontier across math, science and coding evaluations.</p>
<p>Take the benchmark numbers with the usual salt — they were presented by the vendor, some were shown with consensus-of-N sampling against competitors' single samples, and the field has no agreed protocol for this. The comparison charts in a launch livestream are marketing artifacts.</p>
<p>The genuinely notable thing is not the model. It is Colossus.</p>
<h2 id="the-cluster">the cluster<a class="anchor" href="#the-cluster" aria-label="link to this section">#</a></h2>
<p>xAI built a datacenter in Memphis housing on the order of 100,000 H100-class GPUs, and did it on a timeline measured in months rather than years. The conventional wisdom on a buildout of that scale was eighteen to twenty-four months. They compressed it by doing things that are expensive and unglamorous: bringing in mobile gas turbines for interim power, running their own networking integration, and accepting a lot of operational risk.</p>
<p>Whether you find that admirable or reckless depends on your priors and on how you feel about the air quality complaints from the surrounding neighborhood, which are a real and ongoing dispute worth reading about separately.</p>
<p>But as an engineering datapoint it matters: it establishes that the time constant for standing up frontier-scale compute is shorter than the industry assumed, if you are willing to spend and to eat the risk.</p>
<h2 id="what-that-implies">what that implies<a class="anchor" href="#what-that-implies" aria-label="link to this section">#</a></h2>
<p>If a well-capitalized new entrant can go from nothing to frontier-scale compute in about a year, then compute is not a durable moat. It is a capital requirement, which is a different thing. Capital requirements keep out the under-funded; they do not keep out the well-funded.</p>
<p>Which pushes the question of where the actual moat is:</p>
<ul><li><strong>Data</strong> — increasingly contested, increasingly litigated, and the frontier labs are all converging on similar synthetic-data-plus-RL recipes anyway.</li><li><strong>Talent</strong> — mobile, expensive, and being bid on aggressively.</li><li><strong>Distribution</strong> — this one is real. A model inside a product a billion people already open is worth more than a marginally better model behind a signup form.</li><li><strong>Cost per token at quality</strong> — real, and derived from architecture and serving engineering rather than raw scale.</li></ul>
<p>My read is that distribution and serving efficiency are the durable ones, and that is a much less romantic answer than "we have the smartest model."</p>
<h2 id="for-people-who-ship-things">for people who ship things<a class="anchor" href="#for-people-who-ship-things" aria-label="link to this section">#</a></h2>
<p>Practical implication: assume model quality converges and plan accordingly. Do not build your product on the assumption that one vendor's model stays two steps ahead. Build the abstraction layer, keep your evals in your own repo, and make switching a config change.</p>
<p>The teams that did this in 2024 spent a boring week on it and have been changing providers casually ever since. The teams that did not are having architecture meetings.</p>]]></content:encoded></item><item><title>The day the market priced in efficiency</title><link>https://readme.news/the-day-the-market-priced-in-efficiency/</link><guid isPermaLink="true">https://readme.news/the-day-the-market-priced-in-efficiency/</guid><pubDate>Mon, 27 Jan 2025 09:00:00 +0000</pubDate><description>Nvidia lost roughly $600 billion of market value in a session. The trigger was a paper about training costs.</description><content:encoded><![CDATA[<p>Nvidia closed down about 17% today. Broadcom, Vertiv, Constellation Energy and most of the power-adjacent complex fell with it. The proximate cause was a week-old model release from a Chinese lab and a number in its technical report.</p>
<h2 id="the-number">the number<a class="anchor" href="#the-number" aria-label="link to this section">#</a></h2>
<p>DeepSeek's V3 paper stated a final training run cost of roughly $5.6 million in GPU-hours. That figure travelled around the world in about four days, usually stripped of every qualifier attached to it.</p>
<p>The qualifiers matter enormously:</p>
<ul><li>It is the cost of the <strong>final run only</strong>. It explicitly excludes research, failed runs, ablations, and data pipeline work — which in any frontier lab is the overwhelming majority of total spend.</li><li>It excludes the <strong>capital cost of the cluster</strong> itself.</li><li>It says nothing about <strong>R1's</strong> RL training, which came later and separately.</li></ul>
<p>So "they trained a <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> for $5.6M" is not what the paper says. It is what the internet decided the paper says.</p>
<h2 id="why-the-market-reacted-anyway">why the market reacted anyway<a class="anchor" href="#why-the-market-reacted-anyway" aria-label="link to this section">#</a></h2>
<p>Because the directionally correct read survives the correction, and the market is a machine for reacting to directionally correct reads badly.</p>
<p>The directionally correct read: architectural efficiency gains are real and large. DeepSeek's mixture-of-experts design activates a small fraction of total parameters per token. Their multi-head latent attention cuts KV cache size substantially. Their FP8 training pipeline halves memory traffic against BF16. These are engineering wins, they are published, and they are reproducible.</p>
<p>If capability-per-FLOP is improving that fast, then the number of FLOPs you need to buy to reach a given capability is falling. That is the thesis that repriced today.</p>
<h2 id="the-counter-thesis">the counter-thesis<a class="anchor" href="#the-counter-thesis" aria-label="link to this section">#</a></h2>
<p>Jevons. If compute gets cheaper per unit of capability, you do not buy less of it — you find more things to do with it. Reasoning models in particular consume enormous inference compute; a model that thinks for thirty seconds before answering is a very different demand curve than one that answers immediately. Efficiency gains in training get spent on inference.</p>
<p>Both theses are defensible. The honest answer is that nobody knows the shape of the demand curve, and a 17% single-day move in the largest company on earth is not a considered judgement about that. It is a positioning unwind.</p>
<h2 id="for-engineers-specifically">for engineers specifically<a class="anchor" href="#for-engineers-specifically" aria-label="link to this section">#</a></h2>
<p>The useful lesson has nothing to do with stock prices. It is this: the performance-per-dollar frontier is moving fast enough that any architecture decision you make today assuming current inference costs will be wrong within a year, in your favor.</p>
<p>Do not build elaborate <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> and routing infrastructure to shave token costs that are going to fall by an order of magnitude anyway. Build the thing. Measure it. Optimize when the bill actually hurts.</p>
<p>That advice would have been wrong in most previous computing eras. It is right now, and it will stop being right at some point, and watching for that moment is most of the job.</p>]]></content:encoded></item>
</channel>
</rss>
