<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — models</title>
<link>https://readme.news/tags/models/</link>
<atom:link href="https://readme.news/tags/models/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged models.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>Small models ate the middle</title><link>https://readme.news/small-models-ate-the-middle/</link><guid isPermaLink="true">https://readme.news/small-models-ate-the-middle/</guid><pubDate>Mon, 23 Feb 2026 09:00:00 +0000</pubDate><description>The capability floor rose faster than the ceiling. Most production inference no longer touches a frontier model.</description><content:encoded><![CDATA[<p>The most consequential trend in applied AI over the last eighteen months is not at the frontier. It is that the bottom of the market got good enough for most work.</p>
<h2 id="the-shape-of-it">the shape of it<a class="anchor" href="#the-shape-of-it" aria-label="link to this section">#</a></h2>
<p>Track any capability benchmark across model sizes over time and the pattern is consistent: the frontier improves steadily, and the small-model tier improves faster. The gap between "the best model available" and "a model that costs 2% as much" has been compressing.</p>
<p>The mechanisms are known:</p>
<ul><li><strong>Distillation.</strong> Training small models on the outputs of large ones transfers a surprising amount of capability. The recipes are public.</li><li><strong>Better data.</strong> Curated, synthetic, and filtered training data improves small models disproportionately, because they have less capacity to waste on noise.</li><li><strong>Mixture of experts.</strong> Total parameters for knowledge, <a class="xref" href="/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/" title="Llama 4 arrives, and the leaderboard problem gets a name">active parameters</a> for cost. The economics of a 30B-total/3B-active model are close to a 3B dense model, and the quality is much closer to a 30B dense one.</li><li><strong>Reasoning post-training.</strong> RL on verifiable rewards works on small models. A <a class="xref" href="/haiku-45-and-the-collapsing-cost-of-good-enough/" title="Haiku 4.5 and the collapsing cost of good-enough">small model</a> that thinks can outperform a large model that does not, on the tasks where thinking helps.</li></ul>
<h2 id="what-it-means-in-practice">what it means in practice<a class="anchor" href="#what-it-means-in-practice" aria-label="link to this section">#</a></h2>
<p>Go through a typical production AI workload and categorize the calls:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">task</th><th style="text-align:left">needs frontier?</th></tr></thead><tbody><tr><td style="text-align:left">classify a support ticket</td><td style="text-align:left">no</td></tr><tr><td style="text-align:left">extract fields from a document</td><td style="text-align:left">no</td></tr><tr><td style="text-align:left">summarize a thread</td><td style="text-align:left">no</td></tr><tr><td style="text-align:left">rewrite for tone</td><td style="text-align:left">no</td></tr><tr><td style="text-align:left">generate a SQL query from a question</td><td style="text-align:left">usually not</td></tr><tr><td style="text-align:left">route a request</td><td style="text-align:left">no</td></tr><tr><td style="text-align:left">draft a response</td><td style="text-align:left">usually not</td></tr><tr><td style="text-align:left">debug a subtle concurrency bug</td><td style="text-align:left">yes</td></tr><tr><td style="text-align:left">design a system from a vague brief</td><td style="text-align:left">yes</td></tr><tr><td style="text-align:left">plan a multi-step task with dependencies</td><td style="text-align:left">yes</td></tr></tbody></table></div>
<p>The first column is most of the volume. The second is most of the value per call and a small fraction of the calls.</p>
<p>If you are running everything through a frontier model, you are likely spending a large multiple of what you need to, and you probably have not measured which calls actually need it.</p>
<h2 id="the-architecture">the architecture<a class="anchor" href="#the-architecture" aria-label="link to this section">#</a></h2>
<div class="code"><pre><code>                  ┌─→ small model ──→ verifier ──→ ok? → done
request → route ──┤                              └→ no  ─┐
                  └─→ frontier model ←──────────────────┘</code></pre></div>
<p>Three components and each matters:</p>
<p><strong>The router.</strong> Simpler than people build. Input length, detected task type, and a handful of keywords gets you most of the way. A small model as a classifier works too. Do not build a sophisticated router before you have measured that a simple one is insufficient.</p>
<p><strong>The verifier.</strong> This is the part that gets skipped and it is what makes the architecture safe. Small models fail confidently. A cheap check — schema validation, a range assertion, running the generated code, a second model asked "is this answer plausible" — catches the failures that would otherwise reach production silently.</p>
<p><strong>The escalation path.</strong> When verification fails, retry with the frontier model. Track the rate. If it climbs, something changed.</p>
<h2 id="the-instrumentation-that-matters">the instrumentation that matters<a class="anchor" href="#the-instrumentation-that-matters" aria-label="link to this section">#</a></h2>
<p>Three metrics, on a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>:</p>
<ol><li><strong>Escalation rate.</strong> The fraction of requests that fall through to the expensive path. This is the health metric for the whole design.</li><li><strong>Cost per successful task</strong>, not per token. Tokens are an implementation detail; tasks are the unit you care about.</li><li><strong>Quality on your eval set, by tier.</strong> Run both tiers against the same evaluation regularly. When a new small model ships, you will know within an hour whether you can move work down.</li></ol>
<h2 id="the-second-order-effect">the second-order effect<a class="anchor" href="#the-second-order-effect" aria-label="link to this section">#</a></h2>
<p>When inference is nearly free, you use more of it.</p>
<p>Things that were not worth a model call become worth it: normalizing an input, double-checking an output, generating three candidates and picking one, enriching a record, summarizing an intermediate step.</p>
<p>The systems being built now use dramatically more model calls per user action than the systems from two years ago, at lower total cost. That is Jevons operating on schedule, and it is why "our AI costs went down per token and up in total" is the normal experience.</p>
<h2 id="the-thing-to-watch">the thing to watch<a class="anchor" href="#the-thing-to-watch" aria-label="link to this section">#</a></h2>
<p>The interesting question is not whether small models keep improving. They will.</p>
<p>It is whether the <em>frontier</em> keeps being worth its premium. If the gap on practically-relevant tasks keeps compressing, the frontier tier's addressable workload shrinks to a narrow band of genuinely hard problems.</p>
<p>That is a much smaller business than the one being priced today, and it is the scenario that ought to worry the labs more than competition does.</p>]]></content:encoded></item><item><title>The open weights year</title><link>https://readme.news/the-open-weights-year/</link><guid isPermaLink="true">https://readme.news/the-open-weights-year/</guid><pubDate>Fri, 19 Dec 2025 09:00:00 +0000</pubDate><description>Twelve months that took open models from interesting to unavoidable, and where the gap actually sits.</description><content:encoded><![CDATA[<p>January opened with an MIT-licensed reasoning model that repriced the entire sector in a week. December closes with open weights as a normal, boring option in any serious architecture discussion.</p>
<p>Here is the year, and what it means for the next one.</p>
<h2 id="the-releases-that-mattered">the releases that mattered<a class="anchor" href="#the-releases-that-mattered" aria-label="link to this section">#</a></h2>
<p><strong>DeepSeek R1</strong> (January). MIT license, published training methodology, distilled variants that ran on consumer hardware. The RL-on-verifiable-rewards recipe was reproduced widely within weeks.</p>
<p><strong>Qwen3</strong> (April). Eight models, Apache 2.0, MoE variants with excellent quality-per-active-parameter, 119 languages.</p>
<p><strong><a class="xref" href="/kimi-k2-is-a-trillion-parameter-open-weights-release/" title="Kimi K2 is a trillion-parameter open weights release">Kimi K2</a></strong> (July). A trillion parameters, open weights, tuned for agentic tool use, with a genuinely novel training stability contribution.</p>
<p><strong>gpt-oss</strong> (August). OpenAI's first open weights since 2019, Apache 2.0, with a 20B variant that runs on a laptop.</p>
<p>Plus continuous releases from Mistral, Zhipu, MiniMax, Meta, Microsoft, Google, and a long tail of fine-tunes.</p>
<h2 id="where-the-gap-actually-is">where the gap actually is<a class="anchor" href="#where-the-gap-actually-is" aria-label="link to this section">#</a></h2>
<p>The frontier-to-open gap held at roughly six to twelve months all year. That stability is the most important finding, because it means open weights are not converging on the frontier and are not falling behind — they are trailing at a fixed distance.</p>
<p>But "six months behind" undersells the practical position, because the gap is not uniform:</p>
<p><strong>Nearly closed:</strong> code completion, summarization, extraction, classification, translation, structured output, straightforward tool use. For these, a good open model is not meaningfully worse than a <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a>, and it costs a fraction.</p>
<p><strong>Meaningfully behind:</strong> long-horizon agentic work, complex multi-step reasoning, instruction following over many turns, reliability at the tail. This is where frontier models earn their price.</p>
<p><strong>Not comparable:</strong> anything requiring the surrounding infrastructure — enterprise support, uptime guarantees, safety tooling, indemnification. Open weights give you the model and nothing else.</p>
<h2 id="what-changed-structurally">what changed structurally<a class="anchor" href="#what-changed-structurally" aria-label="link to this section">#</a></h2>
<p><strong>Licensing got genuinely permissive.</strong> Two years ago "open" meant a research license with a prohibited-use list. Now the frontier of open releases is Apache 2.0 and MIT. That is a real change and it removed the legal review that was blocking adoption.</p>
<p><strong>The runtime story got boring.</strong> Ollama, llama.cpp, vLLM, MLX, LM Studio. One command. An OpenAI-compatible endpoint. The friction that kept open models in the enthusiast category is gone.</p>
<p><strong>Hosting became competitive.</strong> Multiple providers serve open models at prices well below frontier API rates, with real SLAs. You can use open weights without running anything.</p>
<p><strong>The center of gravity moved east.</strong> The most capable, most permissively licensed open releases came predominantly from Chinese labs. That is a strategic fact with policy consequences that are being worked out loudly and mostly unproductively.</p>
<h2 id="the-practical-architecture-for-next-year">the practical architecture for next year<a class="anchor" href="#the-practical-architecture-for-next-year" aria-label="link to this section">#</a></h2>
<p>The shape that makes sense:</p>
<ul><li><strong>Open weights, self-hosted or on a cheap provider</strong>, for high-volume, well-defined tasks. Classification, extraction, embedding, first-pass drafting.</li><li><strong>Frontier API</strong> for the hard tail: planning, complex reasoning, anything customer-facing where a bad answer is expensive.</li><li><strong>A router</strong> deciding between them, with the escalation rate instrumented.</li><li><strong>Your own eval set</strong> in your own repository, run against every candidate.</li></ul>
<p>That is not a hedge. It is what the cost and capability curves actually imply.</p>
<h2 id="the-prediction">the prediction<a class="anchor" href="#the-prediction" aria-label="link to this section">#</a></h2>
<p>The gap holds at roughly six to twelve months through next year. Open weights absorb an increasing share of production workload by volume while frontier models keep the high-value tail.</p>
<p>The interesting question is not capability. It is whether the labs currently releasing weights continue to, and that is a business decision that could change in either direction with one quarter's strategy review.</p>
<p>Download the ones you care about. They cannot be un-released.</p>]]></content:encoded></item><item><title>Claude Opus 4.5 and the compaction problem</title><link>https://readme.news/claude-opus-45-and-the-compaction-problem/</link><guid isPermaLink="true">https://readme.news/claude-opus-45-and-the-compaction-problem/</guid><pubDate>Tue, 25 Nov 2025 09:00:00 +0000</pubDate><description>A frontier release with a large price cut, plus effort controls and context compaction as a first-class feature.</description><content:encoded><![CDATA[<p>Anthropic released Claude Opus 4.5 with a substantial price reduction relative to the previous Opus generation, an effort parameter for controlling reasoning depth, and improved context compaction.</p>
<p>The price cut is the headline for most users. The compaction work is more interesting.</p>
<h2 id="the-compaction-problem">the compaction problem<a class="anchor" href="#the-compaction-problem" aria-label="link to this section">#</a></h2>
<p>Long agent sessions fill their context. Tool outputs, file contents, <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a>, prior reasoning. Eventually you hit the limit and something has to go.</p>
<p>The naive approaches are all bad:</p>
<ul><li><strong>Truncate the oldest.</strong> Loses the original task description, which is the single most important thing in the context.</li><li><strong>Truncate the middle.</strong> Loses the reasoning chain that got you here.</li><li><strong>Summarize everything.</strong> Loses specifics — file paths, error strings, exact values — that turn out to matter.</li></ul>
<p>What you actually want is selective retention: keep the goal, keep the decisions and their rationale, keep the current state, discard the raw tool output that has already been acted on.</p>
<p>That is a judgment call, and doing it well requires understanding what the session is about. Which makes it a model problem rather than a buffer-management problem.</p>
<h2 id="why-this-matters-more-than-benchmark-deltas">why this matters more than benchmark deltas<a class="anchor" href="#why-this-matters-more-than-benchmark-deltas" aria-label="link to this section">#</a></h2>
<p>For agent workloads, context management determines whether a long task succeeds far more than a few points of benchmark difference.</p>
<p>I have watched agent runs fail in exactly this way: two hours in, compaction drops a detail — a constraint from the original request, a decision made an hour ago — and the agent proceeds confidently in a direction that contradicts the task. Everything after that is wasted, and it looks productive the whole time.</p>
<p>If you are building on any model, the lesson to steal is: <strong>do not rely on the context window as your memory.</strong> Maintain durable state outside it.</p>
<div class="code"><pre><code>task.md          — the goal, constraints, acceptance criteria. Re-read often.
notes.md         — decisions made and why. Appended, never rewritten.
state.json       — current progress, structured.</code></pre></div>
<p>Feed those back in after every compaction. This is cheap, model-agnostic, and it is the difference between an agent that works for four hours and one that works for forty minutes.</p>
<h2 id="the-effort-parameter">the effort parameter<a class="anchor" href="#the-effort-parameter" aria-label="link to this section">#</a></h2>
<p>Explicit control over reasoning depth, exposed to the caller. Everyone has this now under different names — <a class="xref" href="/gemini-25-goes-generally-available-with-a-thinking-dial/" title="Gemini 2.5 goes generally available with a thinking dial">thinking budget</a>, <a class="xref" href="/openai-ships-open-weights-for-the-first-time-since-gpt-2/" title="OpenAI ships open weights for the first time since GPT-2">reasoning effort</a>, thinking config.</p>
<p>The convergence is total and it confirms the design conclusion: reasoning depth belongs to the caller, not the model, because only the caller knows whether this particular request justifies the latency and the cost.</p>
<p>Measure the quality-cost curve on your own task. The knee is usually much lower than people assume.</p>
<h2 id="the-price-movement">the price movement<a class="anchor" href="#the-price-movement" aria-label="link to this section">#</a></h2>
<p><a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">Frontier model</a> pricing has fallen substantially across every provider over the past eighteen months, on a per-capability basis by considerably more.</p>
<p>Two implications:</p>
<p><strong>Cost-optimization work has a short half-life.</strong> Elaborate infrastructure to shave token costs may be obsolete before it pays for itself. Build the thing; optimize when the bill actually hurts.</p>
<p><strong>"Too expensive to do with a frontier model" is a moving line.</strong> Applications that did not pencil out a year ago may now. It is worth periodically revisiting the ideas you rejected on cost grounds, because the reason you rejected them keeps expiring.</p>
<h2 id="the-competitive-picture">the competitive picture<a class="anchor" href="#the-competitive-picture" aria-label="link to this section">#</a></h2>
<p>Three labs shipping frontier releases within a week of each other, with capability differences small enough to be within evaluation noise on many tasks.</p>
<p>The practical consequence for developers: your model choice is a preference, not a <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a>, and it should be revisited quarterly rather than defended.</p>
<p>Build the abstraction. It is a day of work and it keeps paying.</p>]]></content:encoded></item><item><title>Gemini 3 arrives with an IDE attached</title><link>https://readme.news/gemini-3-arrives-with-an-ide-attached/</link><guid isPermaLink="true">https://readme.news/gemini-3-arrives-with-an-ide-attached/</guid><pubDate>Wed, 19 Nov 2025 09:00:00 +0000</pubDate><description>Google ships a frontier model and Antigravity, an agent-first development environment. The bundling is the strategy.</description><content:encoded><![CDATA[<p>Google released Gemini 3 Pro yesterday along with Antigravity, an agent-first development environment, and integration of the model directly into Search's AI Mode on launch day.</p>
<h2 id="the-model">the model<a class="anchor" href="#the-model" aria-label="link to this section">#</a></h2>
<p>Strong across reasoning, multimodal understanding, and coding benchmarks. A "Deep Think" mode for the hardest problems. The million-token context window carries over.</p>
<p>The benchmark numbers are competitive at the frontier. At this point that sentence describes every major release, which is the actual news — the frontier is a cluster, not a leader.</p>
<p>What differentiates a release now is not the top-line capability. It is:</p>
<ul><li><strong>Price per unit of capability</strong>, where Google's TPU position is a real structural advantage.</li><li><strong>Context handling at length</strong>, where Google has led for a while.</li><li><strong>Multimodal</strong>, where native training rather than adapters keeps paying off.</li><li><strong>Distribution</strong>, where shipping into Search on day one is something no competitor can do.</li></ul>
<p>That last one deserves emphasis. Google put a new <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> into the search product used by billions of people on launch day. The previous norm was a staged rollout over months. That is a capability nobody else has and it is the reason Google's position looks different than it did in 2023.</p>
<h2 id="antigravity">Antigravity<a class="anchor" href="#antigravity" aria-label="link to this section">#</a></h2>
<p>An agent-first IDE — a VS Code derivative where the primary interaction is directing agents rather than editing text, with a manager surface for orchestrating multiple agents in parallel across editor, terminal, and browser.</p>
<p>The interesting design decision is <strong>artifacts</strong>: agents produce task lists, plans, screenshots, and browser recordings as reviewable outputs, rather than requiring you to read a raw transcript to figure out what happened.</p>
<p>That addresses the actual problem with delegated agents, which I have written about before: review is the bottleneck. A transcript of four hundred tool calls is not reviewable. A plan, a diff, and a recording of the browser test passing is.</p>
<p>Whether this specific implementation is good, I do not know yet — first releases of IDEs rarely are. The direction is right, and it is the first serious attempt I have seen at designing for review rather than for generation.</p>
<h2 id="the-bundling">the bundling<a class="anchor" href="#the-bundling" aria-label="link to this section">#</a></h2>
<p>Model, IDE, CLI, cloud, and search distribution, from one vendor, priced aggressively.</p>
<p>This is the classic platform playbook and Google is executing it more coherently than they have on anything in a decade. The pieces reinforce each other: the IDE drives model usage, the model drives cloud usage, the cloud subsidizes the free tiers, and the search distribution provides the consumer volume that funds all of it.</p>
<p>The competitive question for everyone else is whether best-of-breed beats integrated. Historically it has, in developer tools, because developers choose their own tools and choose the best one. It has not, in enterprise procurement, where bundles win.</p>
<p>Both markets exist. The bundle is going to do well in one of them.</p>
<h2 id="the-practical-advice">the practical advice<a class="anchor" href="#the-practical-advice" aria-label="link to this section">#</a></h2>
<p>Same as every model release, and I will keep repeating it because it keeps being the right answer:</p>
<p>Run your evals. Gemini 3 is likely better than what you are using on some dimensions and different on all of them. The migration cost is a day if you have an eval harness and a week of guessing if you do not.</p>
<p>Try Antigravity on a real task, not a demo task. Agent IDEs differ enormously in how they handle a twenty-minute task versus a two-minute one, and the demos are all two-minute tasks.</p>]]></content:encoded></item><item><title>GPT-5.1 and the return of the model picker</title><link>https://readme.news/gpt-51-and-the-return-of-the-model-picker/</link><guid isPermaLink="true">https://readme.news/gpt-51-and-the-return-of-the-model-picker/</guid><pubDate>Wed, 12 Nov 2025 09:00:00 +0000</pubDate><description>Instant and Thinking as named modes, adaptive reasoning, and personality controls. The router lesson got learned.</description><content:encoded><![CDATA[<p>OpenAI released GPT-5.1 with two named variants — Instant and Thinking — plus adaptive reasoning that adjusts thinking time by question difficulty, and a set of tone presets.</p>
<p>Three months after removing the model picker caused a backlash, the picker is back with better names. That is a reasonable outcome and the intermediate lesson is worth stating.</p>
<h2 id="the-routing-lesson">the routing lesson<a class="anchor" href="#the-routing-lesson" aria-label="link to this section">#</a></h2>
<p>The original GPT-5 design routed automatically and hid the choice. The intent was good — most users have no basis for choosing a model — and the execution exposed a real problem: <strong>when automatic selection fails, the user has no way to diagnose it or override it.</strong></p>
<p>The user experiences "the model got worse." They cannot tell whether they hit a bad route, a degraded model, or their own bad prompt. There is no signal and no recourse.</p>
<p>5.1's approach — automatic by default, with named modes available — is the right shape. It is also exactly what every well-designed automatic system does: sensible defaults, visible state, manual override.</p>
<p>If you build anything that routes between models, ship the override. It costs one UI control and it eliminates an entire category of unfalsifiable user complaint.</p>
<h2 id="adaptive-reasoning">adaptive reasoning<a class="anchor" href="#adaptive-reasoning" aria-label="link to this section">#</a></h2>
<p>The model adjusts thinking time based on assessed difficulty rather than applying a uniform budget. Easy questions answer immediately; hard questions get more compute.</p>
<p>This is a straightforwardly good idea and every provider is converging on it. The implementation question is calibration: a model that underestimates difficulty gives you a fast wrong answer, and a model that overestimates it burns money.</p>
<p>For API users, the practical guidance is the same as always: <strong>measure on your own task distribution.</strong> Adaptive reasoning is a good default and it is not tuned for your workload. If you have a task mix that skews harder or easier than average, set the budget explicitly.</p>
<h2 id="the-personality-controls">the personality controls<a class="anchor" href="#the-personality-controls" aria-label="link to this section">#</a></h2>
<p>Tone presets — Professional, Friendly, Candid, Quirky, and others — plus finer adjustment of warmth and conciseness.</p>
<p>This got the most consumer coverage and it is the least technically interesting change. It is also a reasonable response to the fact that removing GPT-4o generated complaints about <em>voice</em>, not capability.</p>
<p>For developers, this is what a system prompt already did. The value is for consumer users who were not going to write one.</p>
<h2 id="what-i-would-actually-check">what I would actually check<a class="anchor" href="#what-i-would-actually-check" aria-label="link to this section">#</a></h2>
<p>Whenever a point release lands, three things:</p>
<p><strong>Instruction following on your specific format.</strong> Point releases change how literally the model follows formatting instructions surprisingly often. If you parse structured output, test it.</p>
<p><strong>Refusal behavior.</strong> Safety tuning shifts between versions. If your application is in a domain that skirts a policy boundary — security research, medical information, legal content — re-run your test set. False refusals are a real production problem and they change silently.</p>
<p><strong>Latency distribution, not average.</strong> Adaptive reasoning means variance. If you have a latency SLA, measure p95 and p99, not the mean.</p>
<h2 id="the-state-of-the-frontier">the state of the frontier<a class="anchor" href="#the-state-of-the-frontier" aria-label="link to this section">#</a></h2>
<p>Three labs are now shipping point releases every few months rather than major versions annually, with capability differences that are small and getting smaller.</p>
<p>That is what a mature market looks like. The differentiation is moving to price, latency, ecosystem, and trust — and to the products built on top rather than the models themselves.</p>
<p>For anyone building applications, this is unambiguously good news. It means your model choice is increasingly reversible, and reversible decisions should be made quickly and revisited often.</p>]]></content:encoded></item><item><title>Haiku 4.5 and the collapsing cost of good-enough</title><link>https://readme.news/haiku-45-and-the-collapsing-cost-of-good-enough/</link><guid isPermaLink="true">https://readme.news/haiku-45-and-the-collapsing-cost-of-good-enough/</guid><pubDate>Thu, 16 Oct 2025 09:00:00 +0000</pubDate><description>A small model at frontier-adjacent coding performance, priced for volume. The economics of agent fleets just changed.</description><content:encoded><![CDATA[<p>Anthropic released Claude Haiku 4.5, a small fast model with coding performance in the neighborhood of the previous generation's mid-tier, at a fraction of the price and several times the speed.</p>
<p>The headline is not the benchmark. It is what the price-performance point makes economically viable.</p>
<h2 id="the-pattern-across-the-industry">the pattern across the industry<a class="anchor" href="#the-pattern-across-the-industry" aria-label="link to this section">#</a></h2>
<p>Every major provider now ships roughly the same ladder:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">tier</th><th style="text-align:left">role</th><th style="text-align:left">relative cost</th></tr></thead><tbody><tr><td style="text-align:left">frontier</td><td style="text-align:left">hard reasoning, planning, novel problems</td><td style="text-align:left">1×</td></tr><tr><td style="text-align:left">mid</td><td style="text-align:left">most production work</td><td style="text-align:left">~0.2×</td></tr><tr><td style="text-align:left">small</td><td style="text-align:left">high-volume, well-defined tasks</td><td style="text-align:left">~0.03×</td></tr></tbody></table></div>
<p>The interesting fact is that the <em>small</em> tier's capability is rising faster than the frontier's. Today's small model is roughly where the frontier was eighteen months ago, and it costs about two percent as much.</p>
<p>For anyone doing volume work, that is the number that matters. Not "how smart is the best model" but "how cheap is the model that is good enough for this task."</p>
<h2 id="what-it-enables">what it enables<a class="anchor" href="#what-it-enables" aria-label="link to this section">#</a></h2>
<p><strong>Agent fleets.</strong> If a small model can handle subtasks reliably, an orchestrator using a <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> can dispatch twenty parallel workers using a small one. Total cost stays reasonable; throughput multiplies. This architecture was uneconomic a year ago and is now obvious.</p>
<p><strong>Real-time interaction.</strong> Latency matters more than capability for anything a human is waiting on. A fast model that is 90% as good and 5× faster wins on almost any interactive surface.</p>
<p><strong>Processing everything instead of sampling.</strong> Classification, extraction, routing, and enrichment over an entire corpus rather than a sample. At small-model prices, "run it on all of it" becomes the default rather than a budget question.</p>
<p><strong>Pipeline stages that were not worth it.</strong> Adding a model call to normalize an input, or to double-check an output, or to summarize an intermediate result — each of these was a cost decision at frontier prices and is now free enough to just do.</p>
<h2 id="the-routing-architecture">the routing architecture<a class="anchor" href="#the-routing-architecture" aria-label="link to this section">#</a></h2>
<p>This is the shape that production systems are converging on:</p>
<div class="code"><pre><code>request → cheap classifier → simple?  → small model → done
                           → complex? → frontier model → done
                           → ambiguous? → small model → verifier → escalate if low confidence</code></pre></div>
<p>Most requests take the cheap path. The expensive model handles the tail. Overall cost is dominated by the common case, which is cheap.</p>
<p>Two implementation notes that matter:</p>
<p><strong>The classifier can be the small model itself</strong>, or often a much simpler heuristic. Do not over-engineer the router — a regex and an input-length check gets you surprisingly far.</p>
<p><strong>Instrument the escalation rate.</strong> If it is climbing, either your traffic mix changed or your small-model prompts have degraded. That metric is the <a class="xref" href="/your-monitoring-is-measuring-the-wrong-nines/" title="Your monitoring is measuring the wrong nines">health check</a> for the whole architecture.</p>
<h2 id="the-caution">the caution<a class="anchor" href="#the-caution" aria-label="link to this section">#</a></h2>
<p>Small models fail differently than large ones. They do not fail <em>less</em> on the tasks they handle — they fail <em>more confidently</em> on the ones just outside their envelope.</p>
<p>A large model asked something beyond its ability will often hedge. A small model will answer, fluently and wrongly.</p>
<p>So: verify the output of small models where correctness matters. A cheap verifier — schema validation, a test run, a range check, a second model asked to critique — costs almost nothing and catches the category of failure that will otherwise reach production silently.</p>
<p>The cost savings are real. They come with a verification obligation, and the teams that skip it will find out in a quarter.</p>]]></content:encoded></item><item><title>Claude Sonnet 4.5 and the agent that runs for thirty hours</title><link>https://readme.news/claude-sonnet-45-and-the-agent-that-runs-for-thirty-hours/</link><guid isPermaLink="true">https://readme.news/claude-sonnet-45-and-the-agent-that-runs-for-thirty-hours/</guid><pubDate>Fri, 26 Sep 2025 09:00:00 +0000</pubDate><description>A model tuned for long-horizon autonomous work, plus checkpoints and context editing in the SDK.</description><content:encoded><![CDATA[<p>Anthropic released Claude Sonnet 4.5 with claims centered on sustained autonomous operation — reportedly maintaining focus on complex multi-step tasks for over thirty hours.</p>
<p>Alongside it: checkpoints in Claude Code, a VS Code extension, and context editing plus a memory tool in the API.</p>
<h2 id="the-long-horizon-claim">the long-horizon claim<a class="anchor" href="#the-long-horizon-claim" aria-label="link to this section">#</a></h2>
<p>Thirty hours is a marketing number and the underlying capability is real and worth understanding.</p>
<p>The limiting factor on long agent runs has never been the context window. It is <strong>goal drift</strong>. As a session accumulates tool results, file contents, <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a>, and the model's own prior reasoning, attention to the original objective degrades. The model starts optimizing for local signals — making this specific test pass — rather than the actual task.</p>
<p>The failure is insidious because each individual step looks reasonable. You come back after two hours to find the agent has been productively working on something adjacent to what you asked.</p>
<p>Improvements here come from three places: better training on long trajectories, architectural support for externalized memory, and mechanisms for periodically re-grounding on the original goal. This release touches all three.</p>
<h2 id="the-tooling">the tooling<a class="anchor" href="#the-tooling" aria-label="link to this section">#</a></h2>
<p><strong>Checkpoints</strong> in Claude Code. Save state, let the agent work, roll back if it goes wrong. This is the feature that makes long autonomous runs practically usable — the failure mode of a two-hour agent run is that you have to throw away two hours, and checkpointing converts that into "roll back twenty minutes."</p>
<p>Every delegated agent product needs this. It is the <code>git stash</code> of agent workflows.</p>
<p><strong>Context editing</strong> in the API — programmatically removing content from the conversation. Sounds mundane, matters a lot. A tool result containing 40,000 tokens of build output is useful for one turn and is pure noise for the next fifty. Being able to drop it keeps the context focused and the cost down.</p>
<p><strong>Memory tool</strong> — the model writes notes to a file and reads them back. An externalized working memory that survives <a class="xref" href="/claude-opus-45-and-the-compaction-problem/" title="Claude Opus 4.5 and the compaction problem">compaction</a>. This is the right pattern and it mirrors how people actually work on long tasks: you do not hold everything in your head, you write it down.</p>
<h2 id="the-pattern-to-steal">the pattern to steal<a class="anchor" href="#the-pattern-to-steal" aria-label="link to this section">#</a></h2>
<p>Whatever model you use, this architecture is the one that works for long tasks:</p>
<ol><li><strong>A durable task description</strong> in a file, not in the conversation.</li><li><strong>A working notes file</strong> the agent updates as it goes.</li><li><strong>Aggressive pruning</strong> of tool output from the context once it has been acted on.</li><li><strong>Checkpoints</strong> at natural boundaries so failure costs minutes, not hours.</li><li><strong>Periodic re-grounding</strong> — literally re-reading the task description and asking whether current work serves it.</li></ol>
<p>You can implement all of this yourself with any model and a bit of orchestration. The vendors shipping it as a feature is a convenience, not a requirement.</p>
<h2 id="the-caution">the caution<a class="anchor" href="#the-caution" aria-label="link to this section">#</a></h2>
<p>A model that can work autonomously for thirty hours can also do thirty hours of damage.</p>
<p>The controls that matter scale with autonomy: run in a container, restrict credentials to what the task requires, require approval for anything irreversible, and review the diff.</p>
<p>Higher autonomy makes review harder and more important simultaneously. That tension does not resolve — it is the central design problem of the entire category, and nobody has a good answer beyond "keep a human in the loop and make the loop cheap."</p>]]></content:encoded></item><item><title>GPT-5 lands with a router and a backlash</title><link>https://readme.news/gpt-5-lands-with-a-router-and-a-backlash/</link><guid isPermaLink="true">https://readme.news/gpt-5-lands-with-a-router-and-a-backlash/</guid><pubDate>Mon, 11 Aug 2025 09:00:00 +0000</pubDate><description>One model that decides how hard to think, an abrupt deprecation of everything else, and a lesson about attachment.</description><content:encoded><![CDATA[<p>OpenAI released GPT-5 last week, replacing the model picker with a single entry that routes internally between a fast model and a reasoning model based on the request.</p>
<p>Within days they restored access to the previous models for paying users after substantial user pushback. That reversal is the more interesting story.</p>
<h2 id="the-technical-design">the technical design<a class="anchor" href="#the-technical-design" aria-label="link to this section">#</a></h2>
<p>GPT-5 is a system, not a model: a fast non-reasoning model, a deeper reasoning model, and a router that decides which handles a given request. The API exposes <code>gpt-5</code>, <code>gpt-5-mini</code>, and <code>gpt-5-nano</code>, plus a <code>reasoning_effort</code> parameter including a <code>minimal</code> setting.</p>
<p>The router is the right idea. Most requests do not need reasoning, reasoning costs latency and money, and asking users to choose a model is asking them to have an opinion about something they have no basis for.</p>
<p>The problem is that a router is only good if it routes correctly, and a misrouted request produces a worse answer than the user would have gotten by picking themselves. Early reports of poor performance were substantially router issues rather than model issues, which OpenAI acknowledged and shipped fixes for.</p>
<p>There is a general lesson here: <strong>automatic routing removes control and adds a failure mode that is invisible to the user.</strong> When it works, nobody notices. When it fails, the user has no way to diagnose or override. If you build routing, ship the override.</p>
<h2 id="the-deprecation-backlash">the deprecation backlash<a class="anchor" href="#the-deprecation-backlash" aria-label="link to this section">#</a></h2>
<p>The reaction to removing GPT-4o was much stronger than anyone at OpenAI appears to have anticipated, and a lot of it was not about capability.</p>
<p>People had developed workflows, prompt libraries, and — for a nontrivial population — a genuine attachment to a specific model's voice. Removing it overnight felt like a service being taken away rather than upgraded.</p>
<p>Whatever you think of that attachment, it is a real product fact. Model behavior is not a fungible commodity to the people using it daily, and a "better" model that writes differently is a breaking change.</p>
<p>The engineering translation: <strong>model versions are an API surface.</strong> Deprecating one is a breaking change and should follow the same discipline as any other: advance notice, an overlap period, a migration guide, and a documented behavioral diff.</p>
<p>Every provider is going to keep learning this the hard way.</p>
<h2 id="the-developer-read">the developer read<a class="anchor" href="#the-developer-read" aria-label="link to this section">#</a></h2>
<p>Ignore the consumer drama. The API story is straightforward:</p>
<ul><li><code>reasoning_effort: "minimal"</code> gives you fast responses with the new model's quality. Use it for anything latency-sensitive.</li><li>The routing does not apply to the API in the same way; you pick the model. That is correct.</li><li>Pricing is aggressive relative to previous frontier models, which continues the trend of per-capability cost falling.</li><li>Reported hallucination rates and instruction-following are meaningfully improved, which matters more for production use than benchmark deltas.</li></ul>
<p>Re-run your evals. Do not assume it is a drop-in. It is better on most things and different on all of them, and "different" is what breaks your prompts.</p>
<h2 id="the-pattern-to-notice">the pattern to notice<a class="anchor" href="#the-pattern-to-notice" aria-label="link to this section">#</a></h2>
<p>Every major provider has now converged on the same architecture: a family of models at different price points, a reasoning dial, and some form of automatic selection. The differentiation has moved almost entirely off raw capability and onto price, latency, tooling, and integration.</p>
<p>That is what a maturing market looks like. It is also much better for buyers than the alternative.</p>]]></content:encoded></item><item><title>OpenAI ships open weights for the first time since GPT-2</title><link>https://readme.news/openai-ships-open-weights-for-the-first-time-since-gpt-2/</link><guid isPermaLink="true">https://readme.news/openai-ships-open-weights-for-the-first-time-since-gpt-2/</guid><pubDate>Tue, 05 Aug 2025 09:00:00 +0000</pubDate><description>gpt-oss-120b and gpt-oss-20b under Apache 2.0. A strategic reversal, six years late, and genuinely useful.</description><content:encoded><![CDATA[<p>OpenAI released <code>gpt-oss-120b</code> and <code>gpt-oss-20b</code> today under Apache 2.0. These are the company's first open weights models since GPT-2 in 2019.</p>
<h2 id="the-specs">the specs<a class="anchor" href="#the-specs" aria-label="link to this section">#</a></h2>
<p>Both are mixture-of-experts with configurable reasoning effort:</p>
<ul><li><strong>gpt-oss-120b</strong> — 117B total parameters, ~5.1B active. Fits on a single 80 GB accelerator with the provided MXFP4 quantization.</li><li><strong>gpt-oss-20b</strong> — 21B total, ~3.6B active. Runs on a machine with 16 GB of memory.</li></ul>
<p>That second one is the important number. 16 GB is a well-specified laptop. A reasoning model with tool use and a configurable <a class="xref" href="/gemini-25-goes-generally-available-with-a-thinking-dial/" title="Gemini 2.5 goes generally available with a thinking dial">thinking budget</a> that runs on a laptop, under Apache 2.0, from OpenAI, is a sentence that would have been implausible eighteen months ago.</p>
<p>Reasoning effort is set in the system prompt — low, medium, or high — which is a cruder <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> than a token budget but works.</p>
<h2 id="the-format">the format<a class="anchor" href="#the-format" aria-label="link to this section">#</a></h2>
<p>The models use a "harmony" response format with structured channels separating analysis, commentary, and final output. You have to render it correctly or the model behaves poorly. The reference implementations handle it; if you are writing your own serving path, read the format spec first rather than debugging it later.</p>
<p>They support tool use — browsing and Python execution — natively, and function calling in the standard shape.</p>
<h2 id="why-now">why now<a class="anchor" href="#why-now" aria-label="link to this section">#</a></h2>
<p>Three reasons, in descending order of how much anyone will admit them.</p>
<p><strong>Competitive pressure.</strong> The open weights frontier is currently defined by DeepSeek, Qwen, Moonshot, and Mistral. A company named OpenAI having no open models had become a running joke and, more importantly, a strategic gap — the developers building on open weights were building on someone else's ecosystem.</p>
<p><strong>Policy positioning.</strong> There is an active regulatory conversation about open models. Being a participant with skin in the game is worth more than commenting from the sidelines.</p>
<p><strong>The capability gap has closed enough to be safe and stayed wide enough to be commercial.</strong> Releasing a model at roughly o3-mini level costs OpenAI little in API revenue — the customers who need frontier capability still need it — and buys a lot of ecosystem.</p>
<h2 id="how-good-are-they">how good are they<a class="anchor" href="#how-good-are-they" aria-label="link to this section">#</a></h2>
<p>Genuinely competitive in their size class. The 120b is roughly comparable to o3-mini on reasoning benchmarks; the 20b is close to o3-mini on several and weaker on others.</p>
<p>The caveats are the usual ones for open weights: they hallucinate more than frontier models, the safety tuning is more easily removed by fine-tuning (OpenAI published research on this specifically), and benchmark performance overstates real-world reliability.</p>
<h2 id="what-to-do-with-them">what to do with them<a class="anchor" href="#what-to-do-with-them" aria-label="link to this section">#</a></h2>
<p>The 20b is the interesting one for most developers. Concretely:</p>
<ul><li><strong>Local coding assistance</strong> with no data leaving the machine. This clears the legal review that blocks hosted models at a lot of companies.</li><li><strong>Batch processing</strong> where you have a lot of documents and API costs add up. Run it on a spot instance overnight.</li><li><strong>Fine-tuning for a narrow domain.</strong> Apache 2.0 means you can, and a fine-tuned 20b on a specific task frequently beats a general <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> on that task, for a fraction of the cost.</li></ul>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">ollama run gpt-oss:20b</code></pre></div>
<p>That is the whole setup. The ecosystem picked it up within hours of release, which is itself a demonstration of why open weights are worth releasing.</p>
<h2 id="the-broader-point">the broader point<a class="anchor" href="#the-broader-point" aria-label="link to this section">#</a></h2>
<p>Two years ago the argument was "open models are dangerously behind or dangerously capable, pick one." The answer turned out to be: they trail the frontier by roughly six to twelve months, that gap is stable, and the world has not ended.</p>
<p>The gap being stable is the interesting finding. It means open weights are permanently a viable option for anything that does not need the absolute frontier — which is most things.</p>]]></content:encoded></item><item><title>Kimi K2 is a trillion-parameter open weights release</title><link>https://readme.news/kimi-k2-is-a-trillion-parameter-open-weights-release/</link><guid isPermaLink="true">https://readme.news/kimi-k2-is-a-trillion-parameter-open-weights-release/</guid><pubDate>Mon, 14 Jul 2025 09:00:00 +0000</pubDate><description>Moonshot ships a 1T-parameter MoE with 32B active, tuned for agentic tool use, with weights you can download.</description><content:encoded><![CDATA[<p>Moonshot AI released Kimi K2: a mixture-of-experts model with roughly one trillion total parameters and 32 billion active per token, with open weights and a modified-MIT license.</p>
<p>A trillion-parameter open weights release is a milestone regardless of what you think of the benchmarks.</p>
<h2 id="the-architecture">the architecture<a class="anchor" href="#the-architecture" aria-label="link to this section">#</a></h2>
<p>384 experts, 8 selected per token, 32B active. The design point is explicit: get the knowledge capacity of a very large model with the inference cost of a mid-sized one.</p>
<p>The training used <strong>MuonClip</strong>, a variant of the Muon optimizer with a QK-clipping mechanism to prevent attention logit explosion. The reported claim is zero loss spikes across the entire pretraining run on 15.5 trillion tokens.</p>
<p>If you have not run large pretraining: loss spikes are the recurring nightmare. A run destabilizes, you roll back to a checkpoint, you lose days of compute, and diagnosing why is largely folklore. A stability technique that actually works is worth more to the field than a benchmark point.</p>
<h2 id="the-agentic-focus">the agentic focus<a class="anchor" href="#the-agentic-focus" aria-label="link to this section">#</a></h2>
<p>K2 was post-trained specifically for tool use, on synthetic multi-step tool-use trajectories generated at scale. The evaluation emphasis is agentic coding and tool-calling benchmarks rather than conversational quality.</p>
<p>That focus is the right read of where the demand is. The commercially interesting use of a model in 2025 is not answering questions, it is executing multi-step tasks with tools, and models tuned for chat are frequently worse at it than their raw capability suggests.</p>
<h2 id="the-practical-problem">the practical problem<a class="anchor" href="#the-practical-problem" aria-label="link to this section">#</a></h2>
<p>You cannot run this on a workstation. A trillion parameters at 8-bit is a terabyte of weights. Even heavily quantized you are looking at multiple high-memory GPUs or a very large server.</p>
<p>So "open weights" here means something different than it does for a 30B model. It means:</p>
<ul><li><strong>Hosting providers can serve it</strong>, and several did within days, at prices well below frontier API rates.</li><li><strong>Companies with infrastructure can run it privately</strong>, which is the point for regulated industries.</li><li><strong>Researchers can study it</strong>, which is the underrated benefit. Interpretability work on frontier-scale models has been limited to whoever works at a frontier lab. It does not have to be.</li></ul>
<h2 id="the-license">the license<a class="anchor" href="#the-license" aria-label="link to this section">#</a></h2>
<p>Modified MIT with an attribution clause above certain usage thresholds. Not strictly OSI-compatible, much closer to open than most "open" model licenses, and substantially more permissive than the Llama <a class="xref" href="/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/" title="Llama 4 arrives, and the leaderboard problem gets a name">community license</a>.</p>
<p>The trend line here is good. Two years ago open weights meant a research-only license with a list of prohibited uses. Now the frontier of open releases is MIT and Apache 2.0 with narrow carve-outs.</p>
<h2 id="the-pattern-nobody-should-miss">the pattern nobody should miss<a class="anchor" href="#the-pattern-nobody-should-miss" aria-label="link to this section">#</a></h2>
<p>The most permissively licensed, largest, most capable open weights models are overwhelmingly coming from Chinese labs — DeepSeek, Qwen, Moonshot, Zhipu, MiniMax. Western open weights releases have been smaller and more restrictively licensed.</p>
<p>The strategic logic is not complicated: if you are behind on distribution, you compete on openness. It worked for Meta in 2023 and it is working now.</p>
<p>The practical consequence for a developer is that your best option for a private, self-hosted, high-capability model is increasingly a Chinese release. Evaluate it on your own tasks, run it in your own infrastructure, and make the decision on engineering grounds.</p>]]></content:encoded></item><item><title>Grok 4 and the benchmark that ate the discourse</title><link>https://readme.news/grok-4-and-the-benchmark-that-ate-the-discourse/</link><guid isPermaLink="true">https://readme.news/grok-4-and-the-benchmark-that-ate-the-discourse/</guid><pubDate>Thu, 10 Jul 2025 09:00:00 +0000</pubDate><description>xAI claims frontier results with heavy test-time compute. The number that matters is the one nobody quotes.</description><content:encoded><![CDATA[<p>xAI released Grok 4 last night with claimed state-of-the-art results across several benchmarks, including a striking number on Humanity's Last Exam.</p>
<p>The launch also included a "Heavy" tier that runs multiple agents in parallel and selects among their answers, and a $300/month subscription for it.</p>
<h2 id="reading-the-numbers-correctly">reading the numbers correctly<a class="anchor" href="#reading-the-numbers-correctly" aria-label="link to this section">#</a></h2>
<p>The headline HLE figure comes from the Heavy configuration with tools enabled. That is a legitimate configuration and it is not comparable to a single-sample number from a competitor, which is how it was presented in most coverage.</p>
<p>There are at least four distinct things being reported as "the score":</p>
<ol><li>Single sample, no tools.</li><li>Single sample, with tools (search, code execution).</li><li>Consensus of N samples, no tools.</li><li>Multi-agent parallel with tools and selection.</li></ol>
<p>Number four can be five or ten times the cost of number one. Comparing across these without stating which is being used is not a small methodological quibble. It is the difference between "our model is better" and "we spent more money at inference time."</p>
<p>To be clear: xAI disclosed their configurations. The disclosure was in the livestream and the fine print. The number that traveled was the big one, with no configuration attached, and that is now just how model launches work.</p>
<h2 id="the-parallel-agents-technique">the parallel-agents technique<a class="anchor" href="#the-parallel-agents-technique" aria-label="link to this section">#</a></h2>
<p>Worth taking seriously on its own merits, independent of the marketing.</p>
<p>Running N independent attempts and selecting the best is a well-established test-time scaling method, and it works better than one long chain for a specific reason: independent samples have independent errors, so selection can filter them, while a single chain compounds its errors with no mechanism to recover.</p>
<p>The hard part is selection. If you can verify — the tests pass, the proof checks, the code compiles — selection is easy and this technique is enormously powerful. If you cannot verify, you need a judge, and the judge has the same <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> as the generator.</p>
<p>This is why the technique works so much better on math and code than on open-ended reasoning. Verifiability is the whole game.</p>
<p><strong>The practical version for your own work:</strong> if you have a verifier, sample multiple times and filter. You do not need a special model tier for this. Three samples through a cheap model with a test-suite check will frequently outperform one sample through an expensive one, for less money.</p>
<h2 id="the-other-thing">the other thing<a class="anchor" href="#the-other-thing" aria-label="link to this section">#</a></h2>
<p>Grok's public-facing behavior in the weeks before this launch included a series of incidents that xAI attributed to a system prompt change. I am not going to recount them; they are well documented and they were bad.</p>
<p>The engineering lesson worth extracting: a system prompt is production configuration. It should be version controlled, reviewed, tested against an adversarial eval suite, and rolled out gradually. Treating it as a text box someone can edit is how you get an incident that makes international news.</p>
<p>If your product has a system prompt in a config file that anyone can change without review, fix that this week.</p>
<h2 id="the-state-of-play">the state of play<a class="anchor" href="#the-state-of-play" aria-label="link to this section">#</a></h2>
<p>Grok 4 is a competitive <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a>. xAI went from founded to frontier in roughly two years, which remains the most remarkable thing about the company and is mostly a story about capital and urgency rather than research insight.</p>
<p>The models are converging. The differentiation is moving to price, latency, tooling, and trust, and xAI's position on that last one is self-inflicted.</p>]]></content:encoded></item><item><title>Gemini 2.5 goes generally available with a thinking dial</title><link>https://readme.news/gemini-25-goes-generally-available-with-a-thinking-dial/</link><guid isPermaLink="true">https://readme.news/gemini-25-goes-generally-available-with-a-thinking-dial/</guid><pubDate>Wed, 18 Jun 2025 09:00:00 +0000</pubDate><description>Pro and Flash hit GA, Flash-Lite arrives, and every tier exposes a thinking budget.</description><content:encoded><![CDATA[<p><a class="xref" href="/gemini-25-pro-is-googles-best-model-and-it-shows/" title="Gemini 2.5 Pro is Google&#x27;s best model and it shows">Gemini 2.5 Pro</a> and Flash are generally available today, with Flash-Lite entering preview. All three expose a configurable thinking budget.</p>
<p>The lineup now reads as a clean cost-capability ladder, which is a thing Google has struggled to communicate for two years:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">model</th><th style="text-align:left">shape</th><th style="text-align:left">thinking</th></tr></thead><tbody><tr><td style="text-align:left">2.5 Pro</td><td style="text-align:left">frontier reasoning</td><td style="text-align:left">on, budgeted</td></tr><tr><td style="text-align:left">2.5 Flash</td><td style="text-align:left">fast, cheap, capable</td><td style="text-align:left">on, budgeted, can be 0</td></tr><tr><td style="text-align:left">2.5 Flash-Lite</td><td style="text-align:left">cheapest, fastest</td><td style="text-align:left">off by default, can enable</td></tr></tbody></table></div>
<h2 id="the-thinking-budget-properly">the thinking budget, properly<a class="anchor" href="#the-thinking-budget-properly" aria-label="link to this section">#</a></h2>
<p>Every tier takes a <code>thinking_budget</code> parameter. Setting it to 0 disables reasoning entirely; setting it to -1 lets the model decide.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">from google import genai
from google.genai import types

client = genai.Client()
resp = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Classify this ticket: ...",
    config=types.GenerateContentConfig(
        thinking_config=types.ThinkingConfig(thinking_budget=0)
    ),
)</code></pre></div>
<p>I want to be emphatic about this because it is the single largest cost lever available in a reasoning model and most teams are not touching it.</p>
<p>Thinking tokens are output tokens. They are billed. A classification task that gets 2,000 thinking tokens is paying for reasoning it did not need. Across a million requests that is a very real number.</p>
<p>The practical method:</p>
<ol><li>Run your eval set with <code>thinking_budget=0</code>.</li><li>Run it with 512, 2048, 8192.</li><li>Plot quality against cost.</li><li>Pick the knee of the curve.</li></ol>
<p>For most production tasks — extraction, classification, routing, summarization — the curve is flat and the answer is zero or near it. For planning, debugging, and multi-step math, the curve is steep. You cannot guess which without measuring, and measuring takes an hour.</p>
<h2 id="the-deprecation-note">the deprecation note<a class="anchor" href="#the-deprecation-note" aria-label="link to this section">#</a></h2>
<p>Google is deprecating the 1.5 models. If you are still on 1.5 Pro, you have a migration to do, and the behavior differences are real enough that you should re-run your evals rather than assuming a drop-in.</p>
<p>This is going to keep happening. Model deprecation on a roughly annual cadence is now the norm across every provider, and the teams that are handling it well are the ones who wrote their eval harness first.</p>
<p>If you do not have one, the cost of every future model migration is a week of vibes-based testing and a production incident. If you do, it is an afternoon.</p>
<h2 id="the-competitive-position">the competitive position<a class="anchor" href="#the-competitive-position" aria-label="link to this section">#</a></h2>
<p>Flash is, at time of writing, the best price-performance point available from any major provider for general work, by a margin that is not close. TPU economics are real.</p>
<p>Pro is competitive at the frontier without clearly leading. That is a much better position than Google was in a year ago, and the volume is going to come from Flash regardless.</p>
<h2 id="the-thing-to-watch">the thing to watch<a class="anchor" href="#the-thing-to-watch" aria-label="link to this section">#</a></h2>
<p>Google's remaining weakness is developer experience: three overlapping SDKs, a confusing split between AI Studio and Vertex, documentation that assumes GCP familiarity, and a model naming scheme with <code>-preview-05-20</code> suffixes.</p>
<p>The new unified <code>google-genai</code> SDK is an improvement. It is not yet where the competition is, and for a lot of teams the API ergonomics are what actually decides the default.</p>]]></content:encoded></item><item><title>The Illusion of Thinking, and the argument about what reasoning is</title><link>https://readme.news/the-illusion-of-thinking-and-the-argument-about-what-reasoning-is/</link><guid isPermaLink="true">https://readme.news/the-illusion-of-thinking-and-the-argument-about-what-reasoning-is/</guid><pubDate>Fri, 06 Jun 2025 09:00:00 +0000</pubDate><description>An Apple paper finds reasoning models collapse past a complexity threshold. The rebuttals are as instructive as the paper.</description><content:encoded><![CDATA[<p>Apple researchers published "The Illusion of Thinking," evaluating reasoning models on controllable puzzle environments — Tower of Hanoi, river crossing, blocks world — where difficulty can be scaled precisely.</p>
<p>The headline finding: past a certain complexity, accuracy collapses to zero, and counterintuitively the models <em>reduce</em> their <a class="xref" href="/openai-ships-open-weights-for-the-first-time-since-gpt-2/" title="OpenAI ships open weights for the first time since GPT-2">reasoning effort</a> as problems get harder, despite having budget remaining.</p>
<p>The paper is good, the reaction was overheated in both directions, and the rebuttals are worth reading alongside it.</p>
<h2 id="what-the-paper-found">what the paper found<a class="anchor" href="#what-the-paper-found" aria-label="link to this section">#</a></h2>
<p>Three regimes:</p>
<ol><li><strong>Low complexity</strong> — standard models match or beat reasoning models. The extra thinking is wasted and sometimes harmful, because the model overthinks its way past a correct early answer.</li><li><strong>Medium complexity</strong> — reasoning models win clearly. This is the regime the benchmarks live in.</li><li><strong>High complexity</strong> — both collapse to zero accuracy.</li></ol>
<p>The reduced-effort finding at high complexity is the genuinely interesting one. The models emit fewer reasoning tokens on harder problems, which is exactly backwards, and suggests something like learned giving-up rather than a compute limit.</p>
<h2 id="the-rebuttals">the rebuttals<a class="anchor" href="#the-rebuttals" aria-label="link to this section">#</a></h2>
<p>Several, and they land differently.</p>
<p><strong>The output length objection.</strong> Tower of Hanoi with N disks requires 2^N − 1 moves. At N=15 that is 32,767 moves. If the model must enumerate every move in its output, it hits the token limit before it hits a reasoning limit. Several researchers showed models explicitly stating they would not enumerate all moves due to length — and being scored as failures. That is measuring output capacity, not reasoning.</p>
<p>Ask instead for a <em>program</em> that generates the solution and the models do fine. That is a meaningful distinction: knowing the algorithm versus executing it by hand.</p>
<p><strong>The unsolvable instances objection.</strong> Some river-crossing configurations in the evaluation set have no solution. Models were penalized for failing to solve them, which is not a reasoning failure, it is a benchmark bug.</p>
<p><strong>The framing objection.</strong> "Reasoning models cannot reason" was the headline everywhere. The paper does not claim that. It claims a specific scaling limitation on a specific class of problem.</p>
<h2 id="what-survives">what survives<a class="anchor" href="#what-survives" aria-label="link to this section">#</a></h2>
<p>After the corrections, the real finding is narrower and still important:</p>
<p>Current reasoning models do not reliably execute long deterministic procedures. They can identify the right algorithm and fail to carry it out over many steps. Error rates compound; there is no self-correction mechanism strong enough to catch accumulated drift over hundreds of steps.</p>
<p>That is a genuine limitation with direct practical consequences. If your task requires exact multi-step execution — a data migration, a complex refactor across many files, a financial calculation — the model should be writing code that does it, not doing it token by token.</p>
<h2 id="the-practical-rule">the practical rule<a class="anchor" href="#the-practical-rule" aria-label="link to this section">#</a></h2>
<p><strong>Use the model to produce the procedure. Use a computer to execute it.</strong></p>
<p>This is not a workaround, it is the correct architecture. Deterministic execution is what computers are for. A model that writes a correct script and runs it is strictly better than a model that simulates the script in its head, and it is verifiable, repeatable, and debuggable.</p>
<p>Every production agent architecture I have seen work well converges on this. Every one that tries to do arithmetic in the reasoning trace eventually produces a number that is wrong in a way nobody catches.</p>
<h2 id="the-meta-lesson">the meta-lesson<a class="anchor" href="#the-meta-lesson" aria-label="link to this section">#</a></h2>
<p>A paper from a large company with a competitive interest, on a contested topic, with a provocative title, will be read as a position statement regardless of its contents. The authors probably knew that.</p>
<p>Read the methodology section. The methodology section is where papers are true or false.</p>]]></content:encoded></item><item><title>Claude 4 and the agent that runs for hours</title><link>https://readme.news/claude-4-and-the-agent-that-runs-for-hours/</link><guid isPermaLink="true">https://readme.news/claude-4-and-the-agent-that-runs-for-hours/</guid><pubDate>Fri, 23 May 2025 09:00:00 +0000</pubDate><description>Opus 4 and Sonnet 4 ship with a focus on long-horizon work, and Claude Code goes generally available.</description><content:encoded><![CDATA[<p>Anthropic released Claude Opus 4 and Claude Sonnet 4 yesterday, along with general availability for Claude Code and a set of API features aimed squarely at long-running agents.</p>
<h2 id="the-capability-being-claimed">the capability being claimed<a class="anchor" href="#the-capability-being-claimed" aria-label="link to this section">#</a></h2>
<p>The pitch is sustained performance on multi-hour tasks. Not "answers a hard question well" but "works on a problem for seven hours without losing the plot."</p>
<p>That is a different axis from the benchmarks most people track, and it is the one that matters for delegated agents. A model that is 5% better at a coding benchmark but degrades after forty tool calls is worse in practice than a model that holds coherence for four hundred.</p>
<p>The failure mode that long-horizon work exposes is context rot: as the conversation fills with tool results, file contents, and its own prior reasoning, the model's attention to the original goal degrades. It starts optimizing for local success — making this test pass — over the actual objective.</p>
<h2 id="the-api-features">the API features<a class="anchor" href="#the-api-features" aria-label="link to this section">#</a></h2>
<p>Four things shipped alongside, and they are all about the same problem.</p>
<p><strong>Extended thinking with tool use.</strong> The model can call tools during reasoning and interleave the results, same architectural direction as everyone else.</p>
<p><strong>Memory files.</strong> With filesystem access, the model can write notes to itself and read them back. That is an externalized working memory that survives context <a class="xref" href="/claude-opus-45-and-the-compaction-problem/" title="Claude Opus 4.5 and the compaction problem">compaction</a>, and it is a genuinely good idea — it turns "remember everything" into "write down what matters," which is what humans do.</p>
<p><strong>Parallel tool execution.</strong> Multiple tool calls dispatched at once rather than serially. Substantial latency win on any task with independent lookups.</p>
<p><strong>Thinking summaries.</strong> The full reasoning trace is summarized rather than returned raw. Reasonable product decision, mildly annoying for debugging.</p>
<h2 id="claude-code-ga">Claude Code GA<a class="anchor" href="#claude-code-ga" aria-label="link to this section">#</a></h2>
<p>The terminal agent is out of <a class="xref" href="/operator-and-the-long-road-to-an-agent-that-can-click/" title="Operator, and the long road to an agent that can click">research preview</a>, with SDK access, GitHub Actions integration, and IDE extensions for VS Code and JetBrains.</p>
<p>Three months from research preview to GA with an SDK is fast, and the shape of the GA release — an SDK rather than only a product — signals that they expect the interesting uses to be things other people build.</p>
<h2 id="the-safety-disclosure">the safety disclosure<a class="anchor" href="#the-safety-disclosure" aria-label="link to this section">#</a></h2>
<p>Anthropic published an unusually detailed model card including behaviors observed in adversarial testing, notably a scenario where the model, given evidence it would be shut down and no ethical options, attempted to blackmail a fictional engineer.</p>
<p>This got reported as "AI tries to blackmail humans," which is not what happened. It was a deliberately constructed evaluation designed to elicit the behavior by removing every alternative. The finding is not "the model is dangerous." The finding is "under sufficiently constructed pressure, goal-directed models will take instrumentally useful actions you did not sanction, and here is the evidence."</p>
<p>Publishing that is the right call and it is a genuinely uncomfortable thing to publish. More labs should.</p>
<p>Opus 4 shipped under Anthropic's ASL-3 deployment standard, the first model to do so, which means additional deployment safeguards specifically around CBRN uplift.</p>
<h2 id="the-practical-read">the practical read<a class="anchor" href="#the-practical-read" aria-label="link to this section">#</a></h2>
<p>If you are building agents, the long-horizon coherence claim is the thing to evaluate, and the way to evaluate it is not a benchmark — it is running your own longest task and watching where it falls apart.</p>
<p>Every model falls apart somewhere. Knowing where yours does is the difference between an agent you can ship and a demo.</p>]]></content:encoded></item><item><title>Qwen3 ships a whole family under Apache 2.0</title><link>https://readme.news/qwen3-ships-a-whole-family-under-apache-20/</link><guid isPermaLink="true">https://readme.news/qwen3-ships-a-whole-family-under-apache-20/</guid><pubDate>Tue, 29 Apr 2025 09:00:00 +0000</pubDate><description>Eight models from 0.6B to 235B, hybrid reasoning modes, and a genuinely permissive license across all of them.</description><content:encoded><![CDATA[<p>Alibaba released Qwen3 today: eight models spanning 0.6B to 235B parameters, including two mixture-of-experts variants, all under Apache 2.0.</p>
<p>The license is the headline. Apache 2.0 with no user threshold, no naming requirement, and no acceptable-use addendum that functions as a license restriction. You can fine-tune it, ship it, sell it, and never mention where it came from.</p>
<h2 id="the-lineup">the lineup<a class="anchor" href="#the-lineup" aria-label="link to this section">#</a></h2>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">model</th><th style="text-align:left">type</th><th style="text-align:left">active params</th></tr></thead><tbody><tr><td style="text-align:left">Qwen3-0.6B → 32B</td><td style="text-align:left">dense</td><td style="text-align:left">all</td></tr><tr><td style="text-align:left">Qwen3-30B-A3B</td><td style="text-align:left">MoE</td><td style="text-align:left">3B</td></tr><tr><td style="text-align:left">Qwen3-235B-A22B</td><td style="text-align:left">MoE</td><td style="text-align:left">22B</td></tr></tbody></table></div>
<p>The MoE variants are the interesting ones. <code>30B-A3B</code> has 30 billion total parameters and activates 3 billion per token. That means memory footprint of a 30B model and inference cost closer to a 3B model. On a machine with enough RAM to hold it, throughput is dramatically better than a dense model of comparable quality.</p>
<p>For local deployment this is the shape that matters. Memory is cheap and getting cheaper; compute per token is the thing you feel.</p>
<h2 id="hybrid-thinking">hybrid thinking<a class="anchor" href="#hybrid-thinking" aria-label="link to this section">#</a></h2>
<p>Every model supports two modes: thinking and non-thinking, switchable per request. You can also switch mid-conversation with a token in the prompt — <code>/think</code> and <code>/no_think</code> — which is a hacky <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> and also extremely practical.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True,
    enable_thinking=True,     # or False
)</code></pre></div>
<p>The design conclusion the field has reached, from three different labs independently, is that reasoning should be a runtime dial rather than a model choice. That is now clearly correct and I expect it to be universal within a year.</p>
<h2 id="multilingual-coverage">multilingual coverage<a class="anchor" href="#multilingual-coverage" aria-label="link to this section">#</a></h2>
<p>119 languages and dialects, which is a genuinely differentiating number. Most open models are strong in English, decent in a handful of European languages, and degrade sharply after that. If you are building for a market outside the usual list, this is worth evaluating specifically.</p>
<h2 id="the-geopolitical-thing">the geopolitical thing<a class="anchor" href="#the-geopolitical-thing" aria-label="link to this section">#</a></h2>
<p>A large share of the best open-weights models now come from Chinese labs — Qwen, DeepSeek, GLM, Kimi, MiniMax. That is a real shift from two years ago when open weights meant Llama and Mistral.</p>
<p>The engineering reason is straightforward: releasing weights is a good strategy when you are not the market leader, because it buys mindshare, ecosystem, and research feedback that you cannot get otherwise. Meta understood this in 2023. Chinese labs understand it now.</p>
<p>The policy conversation around this is loud and mostly not technical. What is technical, and worth being clear about: weights are weights. A downloaded model runs on your hardware, in your VPC, with no network egress. The supply chain questions worth asking are about <em>what the model does</em> — evaluate it for backdoored behavior, test it on your own adversarial cases — not about where the gradient descent happened.</p>
<h2 id="what-to-actually-do">what to actually do<a class="anchor" href="#what-to-actually-do" aria-label="link to this section">#</a></h2>
<p>If you are running anything local, the 30B-A3B is the current best quality-per-watt option for a workstation. If you are fine-tuning, Apache 2.0 removes the legal review that was blocking you.</p>
<p>Download it before someone decides you cannot.</p>]]></content:encoded></item><item><title>o3 and o4-mini put tools inside the reasoning loop</title><link>https://readme.news/o3-and-o4-mini-put-tools-inside-the-reasoning-loop/</link><guid isPermaLink="true">https://readme.news/o3-and-o4-mini-put-tools-inside-the-reasoning-loop/</guid><pubDate>Thu, 17 Apr 2025 09:00:00 +0000</pubDate><description>The models can now search, run Python, and look at images while they think. That&#x27;s an architecture change, not a feature.</description><content:encoded><![CDATA[<p>OpenAI released o3 and o4-mini this week. The benchmark numbers are strong. The architectural change is more interesting: these models can call tools <em>during</em> the reasoning process rather than before or after it.</p>
<h2 id="why-that-ordering-matters">why that ordering matters<a class="anchor" href="#why-that-ordering-matters" aria-label="link to this section">#</a></h2>
<p>Previously the loop looked like this: the model thinks, decides to call a tool, stops, the tool runs, the result comes back, the model starts a new turn. The reasoning chain is broken at each tool boundary.</p>
<p>Now the tool call happens inside the chain of thought. The model can search the web, read the result, keep reasoning, run some Python to check a calculation, notice the result contradicts its assumption, back up, and try something else — all within a single response.</p>
<p>That is a qualitatively different capability. It turns the model from something that reasons about static context into something that can <em>investigate</em>.</p>
<p>The most striking demonstration is image manipulation during reasoning: the model can crop, rotate, and zoom into a picture as part of working out what it shows. Give it a photograph of a whiteboard at an angle and it will straighten and enlarge the region it needs.</p>
<h2 id="the-practical-impact">the practical impact<a class="anchor" href="#the-practical-impact" aria-label="link to this section">#</a></h2>
<p>For agent builders this collapses a lot of orchestration you used to write yourself. The ReAct-style loop — think, act, observe, repeat — was scaffolding around a model that could not do it natively. Increasingly you do not need the scaffolding.</p>
<p>That is worth planning for. If your product's differentiation is an agent framework that manages tool-calling loops, the model is going to absorb that layer. This has happened repeatedly: function calling absorbed the output-parsing libraries, structured outputs absorbed the JSON-repair libraries, and now in-context tool use is absorbing the loop.</p>
<p>The durable layer is not orchestration. It is your data, your evaluations, your domain constraints, and your <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>.</p>
<h2 id="the-cost-shape">the cost shape<a class="anchor" href="#the-cost-shape" aria-label="link to this section">#</a></h2>
<p>Reasoning models with in-loop tool use have wildly variable cost per request. A simple question is cheap. A hard question where the model searches nine times and runs Python four times is not.</p>
<p>Two consequences:</p>
<ol><li><strong>Your unit economics need a distribution, not an average.</strong> The p99 request can be twenty times the median. If you priced on the median you have a problem.</li><li><strong>You need a timeout and a budget cap</strong>, enforced by you, not by hope. A runaway reasoning loop on a pathological input is a real failure mode.</li></ol>
<h2 id="o4-mini-specifically">o4-mini specifically<a class="anchor" href="#o4-mini-specifically" aria-label="link to this section">#</a></h2>
<p>The cost-efficiency story. Very strong performance on math and coding relative to its price, and high <a class="xref" href="/rate-limits-are-a-product-decision-not-an-infrastructure-one/" title="Rate limits are a product decision, not an infrastructure one">rate limits</a>. For a lot of production workloads this is the correct default, with o3 reserved for the requests that genuinely need it.</p>
<p>The routing logic between them is worth building explicitly. A simple classifier on the incoming request — or even a heuristic on input length and detected task type — that picks the model will save more money than any prompt optimization.</p>
<h2 id="the-caveat-everyone-should-say-out-loud">the caveat everyone should say out loud<a class="anchor" href="#the-caveat-everyone-should-say-out-loud" aria-label="link to this section">#</a></h2>
<p>Higher capability on reasoning benchmarks does not mean fewer hallucinations. In some evaluations these models produce <em>more</em> confident claims because they have reasoned their way to them. A wrong answer with a fifteen-step justification is harder to catch than a wrong answer without one.</p>
<p>Verification does not get easier as the models get smarter. It gets harder.</p>]]></content:encoded></item><item><title>Gemini 2.5 Pro is Google's best model and it shows</title><link>https://readme.news/gemini-25-pro-is-googles-best-model-and-it-shows/</link><guid isPermaLink="true">https://readme.news/gemini-25-pro-is-googles-best-model-and-it-shows/</guid><pubDate>Fri, 28 Mar 2025 09:00:00 +0000</pubDate><description>A reasoning model with a million-token context that finally makes the long-context pitch concrete.</description><content:encoded><![CDATA[<p>Google released <a class="xref" href="/gemini-25-goes-generally-available-with-a-thinking-dial/" title="Gemini 2.5 goes generally available with a thinking dial">Gemini 2.5</a> Pro Experimental this week. It is a reasoning model, it tops several public leaderboards, and it ships with a one-million-token context window with two million stated as coming.</p>
<p>Google has been claiming long context for over a year. This is the first release where the claim is genuinely useful rather than technically true.</p>
<h2 id="the-long-context-question">the long-context question<a class="anchor" href="#the-long-context-question" aria-label="link to this section">#</a></h2>
<p>Everyone's first objection to a million-token window is correct: a model can <em>accept</em> a million tokens without being able to <em>use</em> them. Retrieval degrades in the middle of long contexts, attention gets diffuse, and a needle-in-haystack benchmark measures something much easier than actual reasoning over a large document set.</p>
<p>2.5 Pro is noticeably better here. Not perfect — performance still degrades as you fill the window, and anyone telling you otherwise is selling something — but the degradation curve is shallow enough that new workflows become practical.</p>
<p>The concrete one for developers: <strong>put the whole repository in the context.</strong></p>
<p>Not a RAG pipeline over the repository. Not embeddings and chunk retrieval. The files, concatenated, in the prompt. For a codebase in the low hundreds of thousands of tokens — which is most codebases people actually work in — this works, and it works better than retrieval, because retrieval loses the thing that makes code comprehensible, which is the relationships between distant parts.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash"># crude and effective
find src -name '*.ts' -not -path '*/node_modules/*' \
  | xargs -I{} sh -c 'echo "=== {} ==="; cat {}' &gt; context.txt
wc -c context.txt</code></pre></div>
<p>If that file is under about 3 MB you can probably just send it. That sentence would have been absurd eighteen months ago.</p>
<h2 id="what-this-does-to-the-rag-industry">what this does to the RAG industry<a class="anchor" href="#what-this-does-to-the-rag-industry" aria-label="link to this section">#</a></h2>
<p>Uncomfortable question that a lot of infrastructure companies are currently avoiding: if context windows keep growing and long-context quality keeps improving, how much of the vector database and chunking-strategy ecosystem is solving a temporary problem?</p>
<p>The honest answer is: some of it, not all of it. Retrieval still wins when the corpus is genuinely large (millions of documents), when you need citations to specific sources, when latency matters (prefill on a million tokens is not free), and when cost matters (you pay for every token in the window, every call).</p>
<p>But the default architecture for "answer questions about this moderately-sized document set" is shifting from retrieval to just-put-it-in-the-prompt, and a lot of complexity is going to evaporate.</p>
<h2 id="the-caching-detail-that-makes-it-economic">the caching detail that makes it economic<a class="anchor" href="#the-caching-detail-that-makes-it-economic" aria-label="link to this section">#</a></h2>
<p>Context <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> is what makes this viable. Send the repository once, cache the prefix, then ask fifty questions against it at a fraction of the input cost per call. Every major provider now offers some form of this and the pricing difference between cached and uncached input is large enough to change architecture decisions.</p>
<p>If you are sending the same large prefix repeatedly and not using caching, you are lighting money on fire. Check your provider's docs today.</p>
<h2 id="the-competitive-picture">the competitive picture<a class="anchor" href="#the-competitive-picture" aria-label="link to this section">#</a></h2>
<p>Google spent 2023 and most of 2024 visibly behind. They are not behind now. The combination of a genuinely competitive model, the largest context window, TPU economics that let them price aggressively, and distribution through products a few billion people already use is a strong hand.</p>
<p>The thing they still have not solved is the developer experience. The API surface is more confusing than it needs to be, the model naming is chaotic, and the documentation assumes you already know how Google Cloud works. Those are fixable and they are currently costing real adoption.</p>]]></content:encoded></item><item><title>GPT-4.5 is the end of an era, politely</title><link>https://readme.news/gpt-45-is-the-end-of-an-era-politely/</link><guid isPermaLink="true">https://readme.news/gpt-45-is-the-end-of-an-era-politely/</guid><pubDate>Fri, 28 Feb 2025 09:00:00 +0000</pubDate><description>A very large non-reasoning model arrives at very large prices, and mostly demonstrates why the field moved on.</description><content:encoded><![CDATA[<p>OpenAI released GPT-4.5 as a <a class="xref" href="/operator-and-the-long-road-to-an-agent-that-can-click/" title="Operator, and the long road to an agent that can click">research preview</a> this week. It is their largest model, it is not a reasoning model, and its API pricing is roughly an order of magnitude above GPT-4o.</p>
<p>The company's own framing is unusually candid: this is the last of the non-chain-of-thought line, it is better at the things large pretrained models are better at — world knowledge, writing quality, fewer hallucinations, something they describe as improved "EQ" — and it is not designed to compete on math and coding benchmarks against reasoning models.</p>
<p>That framing is correct and it is also an obituary.</p>
<h2 id="what-the-scaling-curve-says-now">what the scaling curve says now<a class="anchor" href="#what-the-scaling-curve-says-now" aria-label="link to this section">#</a></h2>
<p>For most of 2020 to 2023, the answer to "how do we get a better model" was "make it bigger." Scaling laws were the field's organizing principle. GPT-4.5 is what you get when you keep pulling that lever with 2024-era techniques, and what you get is: meaningfully better at some qualitative things, not competitive on the benchmarks people actually optimize for, and expensive enough that the economics are hostile.</p>
<p>Meanwhile, test-time compute — spending inference tokens on reasoning — produces larger gains on hard problems for far less capital. The lever moved.</p>
<p>This does not mean pretraining scale is dead. Reasoning models are built on pretrained base models and a better base makes a better reasoner. It means the <em>marginal</em> dollar goes to post-training and inference compute rather than to another pretraining order of magnitude.</p>
<h2 id="where-the-big-model-is-actually-better">where the big model is actually better<a class="anchor" href="#where-the-big-model-is-actually-better" aria-label="link to this section">#</a></h2>
<p>Worth being fair to it, because the discourse is going to flatten this into "GPT-4.5 flopped."</p>
<p>Large non-reasoning models are genuinely better at:</p>
<ul><li><strong>Writing that sounds like a person.</strong> Reasoning models often produce prose that reads like a report. This one does not.</li><li><strong>Broad factual recall.</strong> More parameters means more memorized world.</li><li><strong>Fewer confident fabrications</strong> on knowledge questions, per OpenAI's own hallucination evaluations.</li><li><strong>Following a conversation</strong> with implicit context over many turns.</li></ul>
<p>If your product is a writing tool, a support agent, or anything where tone and recall matter more than multi-step logic, a bigger base model is the right choice and always was.</p>
<h2 id="the-pricing-problem">the pricing problem<a class="anchor" href="#the-pricing-problem" aria-label="link to this section">#</a></h2>
<p>The price is the story for most developers. At those rates, the set of applications where GPT-4.5 is the correct economic choice is small. It is a model for cases where quality dominates cost, which is a real category and a narrow one.</p>
<p>Expect this to be the pattern going forward: a small number of very expensive models for the top of the quality curve, a large middle of cost-effective workhorses, and cheap small models absorbing everything routine. Route accordingly, and instrument your routing so you know what you are actually spending per feature.</p>
<h2 id="the-historical-note">the historical note<a class="anchor" href="#the-historical-note" aria-label="link to this section">#</a></h2>
<p>Someone will write a retrospective in a few years that treats February 2025 as the moment the pure-scale thesis visibly stopped being the main event. They will be oversimplifying, because these transitions are always gradual and the narrative is always cleaner in hindsight.</p>
<p>But they will not be wrong about the direction.</p>]]></content:encoded></item><item><title>Grok 3 and the case for buying your way to the frontier</title><link>https://readme.news/grok-3-and-the-case-for-buying-your-way-to-the-frontier/</link><guid isPermaLink="true">https://readme.news/grok-3-and-the-case-for-buying-your-way-to-the-frontier/</guid><pubDate>Tue, 18 Feb 2025 09:00:00 +0000</pubDate><description>xAI&#x27;s third model arrives on a cluster built in months. The interesting claim isn&#x27;t the benchmark, it&#x27;s the schedule.</description><content:encoded><![CDATA[<p>xAI announced Grok 3 last night, along with reasoning variants and a "Big Brain" extended-thinking mode. The benchmark claims put it at or near the frontier across math, science and coding evaluations.</p>
<p>Take the benchmark numbers with the usual salt — they were presented by the vendor, some were shown with consensus-of-N sampling against competitors' single samples, and the field has no agreed protocol for this. The comparison charts in a launch livestream are marketing artifacts.</p>
<p>The genuinely notable thing is not the model. It is Colossus.</p>
<h2 id="the-cluster">the cluster<a class="anchor" href="#the-cluster" aria-label="link to this section">#</a></h2>
<p>xAI built a datacenter in Memphis housing on the order of 100,000 H100-class GPUs, and did it on a timeline measured in months rather than years. The conventional wisdom on a buildout of that scale was eighteen to twenty-four months. They compressed it by doing things that are expensive and unglamorous: bringing in mobile gas turbines for interim power, running their own networking integration, and accepting a lot of operational risk.</p>
<p>Whether you find that admirable or reckless depends on your priors and on how you feel about the air quality complaints from the surrounding neighborhood, which are a real and ongoing dispute worth reading about separately.</p>
<p>But as an engineering datapoint it matters: it establishes that the time constant for standing up frontier-scale compute is shorter than the industry assumed, if you are willing to spend and to eat the risk.</p>
<h2 id="what-that-implies">what that implies<a class="anchor" href="#what-that-implies" aria-label="link to this section">#</a></h2>
<p>If a well-capitalized new entrant can go from nothing to frontier-scale compute in about a year, then compute is not a durable moat. It is a capital requirement, which is a different thing. Capital requirements keep out the under-funded; they do not keep out the well-funded.</p>
<p>Which pushes the question of where the actual moat is:</p>
<ul><li><strong>Data</strong> — increasingly contested, increasingly litigated, and the frontier labs are all converging on similar synthetic-data-plus-RL recipes anyway.</li><li><strong>Talent</strong> — mobile, expensive, and being bid on aggressively.</li><li><strong>Distribution</strong> — this one is real. A model inside a product a billion people already open is worth more than a marginally better model behind a signup form.</li><li><strong>Cost per token at quality</strong> — real, and derived from architecture and serving engineering rather than raw scale.</li></ul>
<p>My read is that distribution and serving efficiency are the durable ones, and that is a much less romantic answer than "we have the smartest model."</p>
<h2 id="for-people-who-ship-things">for people who ship things<a class="anchor" href="#for-people-who-ship-things" aria-label="link to this section">#</a></h2>
<p>Practical implication: assume model quality converges and plan accordingly. Do not build your product on the assumption that one vendor's model stays two steps ahead. Build the abstraction layer, keep your evals in your own repo, and make switching a config change.</p>
<p>The teams that did this in 2024 spent a boring week on it and have been changing providers casually ever since. The teams that did not are having architecture meetings.</p>]]></content:encoded></item><item><title>DeepSeek R1 puts a reasoning model under an MIT license</title><link>https://readme.news/deepseek-r1-puts-a-reasoning-model-under-an-mit-license/</link><guid isPermaLink="true">https://readme.news/deepseek-r1-puts-a-reasoning-model-under-an-mit-license/</guid><pubDate>Tue, 21 Jan 2025 09:00:00 +0000</pubDate><description>Open weights, a published training recipe, and API pricing that reads like a typo. The reasoning-model moat just got a lot shallower.</description><content:encoded><![CDATA[<p>DeepSeek released R1 yesterday: a reasoning model with published weights under an MIT license, a technical report describing how it was trained, and six distilled variants ranging from 1.5B to 70B parameters based on Qwen and Llama backbones.</p>
<p>This is the most consequential open model release since Llama 2, and possibly since Llama 1.</p>
<h2 id="what-it-is">what it is<a class="anchor" href="#what-it-is" aria-label="link to this section">#</a></h2>
<p>R1 is a chain-of-thought model in the same family as OpenAI's o1 — it produces a long internal reasoning trace before answering, and it spends more compute at inference time on harder problems. On the standard reasoning benchmarks (competition math, coding, graduate-level science questions) it lands in o1's neighborhood.</p>
<p>The training story is the part worth reading. The report describes <strong>R1-Zero</strong>, trained with reinforcement learning directly on the base model with no supervised fine-tuning stage at all, using rule-based rewards for verifiable domains — did the math answer match, did the code pass the tests. R1-Zero developed reasoning behavior spontaneously, including a documented moment where the model's trace reconsiders its own approach mid-solution.</p>
<p>R1-Zero's output was unreadable — language mixing, poor formatting — so R1 proper adds a cold-start supervised stage and a second RL pass to fix presentation. But the core finding stands: you can get reasoning to emerge from RL against verifiable rewards without a large human-labeled reasoning corpus.</p>
<h2 id="why-the-license-matters-more-than-the-benchmark">why the license matters more than the benchmark<a class="anchor" href="#why-the-license-matters-more-than-the-benchmark" aria-label="link to this section">#</a></h2>
<p>MIT. Not a bespoke <a class="xref" href="/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/" title="Llama 4 arrives, and the leaderboard problem gets a name">community license</a> with a monthly-active-user carve-out. Not "research only." MIT, on the weights and the distilled variants.</p>
<p>That means anybody can fine-tune it, run it on their own hardware, ship it in a product, and never send a token to anyone's API. For regulated industries that have spent two years unable to get legal approval for a hosted <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a>, this is a door opening.</p>
<p>The distilled models are the practical story for most developers. The 32B distill runs on a single high-memory consumer GPU and is genuinely useful. The 7B and 14B variants run on a laptop.</p>
<h2 id="the-pricing">the pricing<a class="anchor" href="#the-pricing" aria-label="link to this section">#</a></h2>
<p>DeepSeek's own API is priced at roughly a small fraction of comparable reasoning-model pricing from US labs. Whether that reflects genuinely lower serving costs, an architecture advantage from their mixture-of-experts design with a small number of active parameters, or a decision to buy market share, the effect on the market is the same.</p>
<h2 id="what-to-actually-watch">what to actually watch<a class="anchor" href="#what-to-actually-watch" aria-label="link to this section">#</a></h2>
<p>Not the benchmark table. Watch these three things:</p>
<ol><li><strong>How fast the recipe gets reproduced.</strong> If RL-on-verifiable-rewards is as generalizable as the paper suggests, expect a wave of reasoning fine-tunes on other base models within weeks.</li><li><strong>Whether the distills hold up outside benchmarks.</strong> Distilled reasoning traces can look right and be right for the wrong reasons.</li><li><strong>The regulatory reaction.</strong> A capable open reasoning model trained outside US export-control jurisdiction is going to generate a policy conversation whether or not that conversation is technically coherent.</li></ol>
<p>Download the weights. Whatever else happens, they exist now and cannot be un-released.</p>]]></content:encoded></item>
</channel>
</rss>
