<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — hardware</title>
<link>https://readme.news/tags/hardware/</link>
<atom:link href="https://readme.news/tags/hardware/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged hardware.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>GTC 2026 and the rack that draws a megawatt</title><link>https://readme.news/gtc-2026-and-the-rack-that-draws-a-megawatt/</link><guid isPermaLink="true">https://readme.news/gtc-2026-and-the-rack-that-draws-a-megawatt/</guid><pubDate>Mon, 09 Mar 2026 09:00:00 +0000</pubDate><description>New silicon, new interconnect, and an industry where the unit of purchase is a room rather than a card.</description><content:encoded><![CDATA[<p>Nvidia's annual conference happened this week and the shape of the announcements confirms what has been obvious for two years: the product is no longer a chip.</p>
<h2 id="the-unit-of-sale">the unit of sale<a class="anchor" href="#the-unit-of-sale" aria-label="link to this section">#</a></h2>
<p>The thing being sold is a <a class="xref" href="/gtc-2025-a-roadmap-to-2027-and-a-warning-about-power/" title="GTC 2025: a roadmap to 2027 and a warning about power">rack-scale</a> system — dozens of accelerators, a co-designed interconnect, integrated liquid cooling, and a power distribution architecture — that arrives as a unit and is installed by people who specialize in installing it.</p>
<p>That is a fundamentally different business from selling PCIe cards, and it is the moat. Competing on a single accelerator's FLOPs per watt is a fight several companies can have. Competing on an integrated rack with a proprietary high-bandwidth interconnect between every device in it is a fight almost nobody can have, because the interconnect is where a decade of engineering lives.</p>
<h2 id="the-power-numbers">the power numbers<a class="anchor" href="#the-power-numbers" aria-label="link to this section">#</a></h2>
<p>The per-rack power figures continue climbing. The industry has moved from "how many kilowatts per rack" to "how many racks per megawatt," which is a different mental model.</p>
<p>The consequences, which are the actual story:</p>
<p><strong>Existing datacenter shells are mostly unusable.</strong> A facility designed for 10 kW racks cannot host these regardless of floor space. The buildout is new construction, near new substations, on a permitting timeline.</p>
<p><strong>Liquid cooling is not optional and not new.</strong> Direct-to-chip liquid cooling is now the baseline, and the interesting engineering has moved to facility-level heat rejection — where does the heat go, and can it be sold to someone.</p>
<p><strong>Higher-voltage DC distribution</strong> is being adopted to reduce conversion losses and copper mass at these current levels. This is a genuine architectural change to how datacenters distribute power and it is being driven entirely by this workload.</p>
<h2 id="the-software-announcements">the software announcements<a class="anchor" href="#the-software-announcements" aria-label="link to this section">#</a></h2>
<p>The inference serving stack continues to be where the practical developer value is. Disaggregated serving — running prefill and decode on separate hardware pools because they have completely different compute and memory characteristics — is now mature enough to be the default recommendation rather than an optimization.</p>
<p>If you operate inference at any scale and have not looked at disaggregation, that remains the largest single efficiency win available. Prefill is compute-bound and parallelizable. Decode is memory-bandwidth-bound and sequential. Running them on one homogeneous pool wastes a substantial fraction of your silicon on whichever phase is not currently limiting.</p>
<h2 id="the-competitive-picture">the competitive picture<a class="anchor" href="#the-competitive-picture" aria-label="link to this section">#</a></h2>
<p>Every hyperscaler now has its own accelerator in volume production. The software gap versus CUDA remains the deciding factor, and it is narrowing slowly rather than quickly, because CUDA's advantage is fifteen years of accumulated libraries, kernels, tooling, and institutional knowledge rather than any single technical feature.</p>
<p>The realistic near-term outcome is not displacement. It is a market where hyperscalers run their own silicon for their own high-volume internal workloads and buy Nvidia for everything else, which is enough to constrain pricing without threatening the position.</p>
<h2 id="what-a-developer-should-take-from-this">what a developer should take from this<a class="anchor" href="#what-a-developer-should-take-from-this" aria-label="link to this section">#</a></h2>
<p><strong>Almost none of it directly.</strong> You will not buy one of these. You will rent time on one.</p>
<p>What matters to you:</p>
<ul><li><strong>Capacity comes in steps</strong>, tied to construction schedules. Plan around availability, not just price.</li><li><strong>Efficiency work compounds.</strong> Every optimization that reduces tokens, batches better, or caches more is worth more in an environment where the underlying resource is physically constrained.</li><li><strong>The abstraction layer is your friend.</strong> The hardware underneath your API calls will change several times. If your code cares, that is a design problem you created.</li></ul>
<h2 id="the-sentence-that-summarizes-the-era">the sentence that summarizes the era<a class="anchor" href="#the-sentence-that-summarizes-the-era" aria-label="link to this section">#</a></h2>
<p>The bottleneck on artificial intelligence is currently electrical engineering and civil construction, and it has been for about two years.</p>
<p>Every discussion about model capability that ignores this is discussing a hypothetical.</p>]]></content:encoded></item><item><title>CES 2026: the power supply is the product</title><link>https://readme.news/ces-2026-the-power-supply-is-the-product/</link><guid isPermaLink="true">https://readme.news/ces-2026-the-power-supply-is-the-product/</guid><pubDate>Tue, 06 Jan 2026 09:00:00 +0000</pubDate><description>Another year of AI in appliances, and one genuine trend hiding underneath: everything is now thermally constrained.</description><content:encoded><![CDATA[<p>The annual Las Vegas exercise in stapling language models to inventory has concluded. As always, the interesting signal is in the parts nobody put on a keynote slide.</p>
<h2 id="the-actual-trend">the actual trend<a class="anchor" href="#the-actual-trend" aria-label="link to this section">#</a></h2>
<p>Every category at this show is now constrained by thermals and power delivery rather than by compute.</p>
<p><strong>Laptops</strong> with neural accelerators that cannot sustain their rated throughput for more than a few minutes in a thin chassis. The TOPS number on the sticker is a peak, and the peak lasts as long as the thermal mass does.</p>
<p><strong>Handhelds</strong> where battery life is the entire product decision and every added capability is a subtraction from it.</p>
<p><strong>Home devices</strong> that want to run models locally and cannot, because the power envelope of something that sits on a shelf is a few watts.</p>
<p><strong>Desktops</strong> where the power supply recommendations for a current-generation GPU have crept past what a lot of household circuits comfortably deliver alongside everything else in the room.</p>
<p>This is what a computing era looks like when it hits a physical wall. The interesting engineering for the next several years is efficiency, not capability — which is historically when the best engineering happens.</p>
<h2 id="what-to-actually-take-from-the-show">what to actually take from the show<a class="anchor" href="#what-to-actually-take-from-the-show" aria-label="link to this section">#</a></h2>
<p><strong><a class="xref" href="/pixel-10-and-the-on-device-model-as-a-platform-feature/" title="Pixel 10 and the on-device model as a platform feature">On-device</a> inference is a spec sheet lie in the thin-and-light category.</strong> If you are shipping software that assumes a local NPU, test on a machine that has been running for twenty minutes, not on a cold one. The difference is large and nobody benchmarks it.</p>
<p><strong>Unified memory keeps winning.</strong> Every serious local-AI machine announced has shared CPU/GPU memory at high capacity. Discrete VRAM is a bad fit for a workload where capacity matters more than peak bandwidth, and the industry has figured that out.</p>
<p><strong>The robot demos are still teleoperated.</strong> They have been for four years. Watch for the operator's hands, or the suspiciously smooth trajectory, or the fact that the demo never deviates from the script. There is real progress in robotics and it is happening in warehouses, not on stages.</p>
<h2 id="the-accessory-economy">the accessory economy<a class="anchor" href="#the-accessory-economy" aria-label="link to this section">#</a></h2>
<p>A large fraction of the floor was accessories for devices that do not exist yet: cases, docks, and mounts for AI wearables. That is a leading indicator of nothing except that the accessory industry moves fast and takes risks cheaply.</p>
<h2 id="the-one-thing-i-would-actually-buy">the one thing I would actually buy<a class="anchor" href="#the-one-thing-i-would-actually-buy" aria-label="link to this section">#</a></h2>
<p>The mundane one: displays got better and cheaper. High-refresh, high-resolution, accurate-color monitors are at prices that were flagship prices three years ago.</p>
<p>If you have been running the same monitor since 2020 and you spend eight hours a day looking at text on it, that is the highest-value hardware purchase available to you, and it will not be mentioned in a single keynote.</p>
<h2 id="the-meta-observation">the meta-observation<a class="anchor" href="#the-meta-observation" aria-label="link to this section">#</a></h2>
<p>CES has been a bad predictor of what matters for about a decade. The genuinely important hardware of the last ten years — the M1, the H100, TPUs, the shift to ARM in servers — was announced at industry events or in blog posts, to audiences who understood it, without a stage show.</p>
<p>The consumer electronics show is now mostly a trade event for retail buyers, covered as if it were a technology forecast. Read the specs. Skip the narrative.</p>]]></content:encoded></item><item><title>Nvidia's quarter and the question nobody can answer</title><link>https://readme.news/nvidias-quarter-and-the-question-nobody-can-answer/</link><guid isPermaLink="true">https://readme.news/nvidias-quarter-and-the-question-nobody-can-answer/</guid><pubDate>Fri, 21 Nov 2025 09:00:00 +0000</pubDate><description>Another enormous beat, another set of concerns about circular financing. Both facts are real.</description><content:encoded><![CDATA[<p>Nvidia reported another quarter far above expectations, with data center revenue continuing to grow at a rate that would be implausible in any other context.</p>
<p>The stock's reaction was muted relative to the beat, which tells you the debate has moved from "is demand real" to "is demand <em>sustainable</em>, and how much of it is funded by Nvidia."</p>
<h2 id="the-bull-case">the bull case<a class="anchor" href="#the-bull-case" aria-label="link to this section">#</a></h2>
<p>Straightforward and well supported:</p>
<ul><li>Every hyperscaler raised capital expenditure guidance again.</li><li>Inference demand is growing faster than training demand, and inference is the recurring workload rather than the one-time one.</li><li>Reasoning models consume dramatically more inference compute per request than their predecessors, and adoption is rising.</li><li>Supply remains the constraint. Lead times are long. Customers are queueing.</li><li>The <a class="xref" href="/gtc-2025-a-roadmap-to-2027-and-a-warning-about-power/" title="GTC 2025: a roadmap to 2027 and a warning about power">rack-scale</a> systems business has a moat that individual chip competition does not touch — the interconnect is the product.</li></ul>
<h2 id="the-bear-case">the bear case<a class="anchor" href="#the-bear-case" aria-label="link to this section">#</a></h2>
<p>Also straightforward:</p>
<ul><li>A meaningful share of revenue traces to customers Nvidia has invested in or financed, which makes the demand signal less independent.</li><li>Depreciation schedules on AI hardware are assumed at five to six years. If the useful life is closer to three — which some operators argue, given the pace of generational improvement — reported profitability across the sector is overstated.</li><li>Hyperscalers are building their own silicon. Google's TPUs are mature, Amazon's Trainium is shipping in volume, and every one of those deployments is a substituted Nvidia sale.</li><li>Model efficiency improvements keep arriving. If capability-per-FLOP keeps improving as fast as it has, required FLOPs for a given capability fall.</li><li>The financing environment for the buildout depends on continued access to cheap debt.</li></ul>
<h2 id="what-nobody-knows">what nobody knows<a class="anchor" href="#what-nobody-knows" aria-label="link to this section">#</a></h2>
<p>Whether AI application revenue will eventually justify the infrastructure spend.</p>
<p>Current annualized revenue across the AI application layer is a fraction of annual AI capital expenditure. That gap can close — infrastructure is built ahead of demand in every capital cycle, and railroads, fiber, and cloud all looked insane at the equivalent stage.</p>
<p>It can also not close. Fiber overbuild in 2000 was followed by a decade of dark fiber and a lot of bankruptcies, and the eventual users of that fiber were not the companies that laid it.</p>
<p>Both patterns are real. The people confidently predicting which one applies here are pattern-matching, not analyzing, and that includes the ones I agree with.</p>
<h2 id="why-an-engineer-should-care">why an engineer should care<a class="anchor" href="#why-an-engineer-should-care" aria-label="link to this section">#</a></h2>
<p>Not for investing advice. For planning.</p>
<p><strong>Compute pricing is not going to fall smoothly.</strong> If the capex cycle continues, capacity comes online in steps and prices drift down. If it contracts, capacity tightens and prices firm. Do not build a business model that requires a specific trajectory.</p>
<p><strong>Efficiency work has enduring value.</strong> Whatever happens to the capex cycle, using less compute for the same result is good. Prompt <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a>, model routing, smaller models for routine tasks, batch processing where latency permits — all of it pays regardless of the macro environment.</p>
<p><strong>Multi-provider capability is cheap insurance.</strong> The cost of abstracting your model calls is a day. The cost of being locked to a provider whose pricing or availability changes is much larger.</p>
<h2 id="the-sentence-i-keep-coming-back-to">the sentence I keep coming back to<a class="anchor" href="#the-sentence-i-keep-coming-back-to" aria-label="link to this section">#</a></h2>
<p>The infrastructure being built is real, the demand today is real, and whether they match at the scale being assumed is genuinely unknown to everyone including the people spending the money.</p>
<p>That is an uncomfortable place to be, and pretending otherwise — in either direction — is the main thing to avoid.</p>]]></content:encoded></item><item><title>OpenAI and Nvidia sign a circular deal</title><link>https://readme.news/openai-and-nvidia-sign-a-circular-deal/</link><guid isPermaLink="true">https://readme.news/openai-and-nvidia-sign-a-circular-deal/</guid><pubDate>Tue, 23 Sep 2025 09:00:00 +0000</pubDate><description>Up to $100 billion of investment tied to gigawatts of deployment. The financing structures are getting interesting.</description><content:encoded><![CDATA[<p>Nvidia and OpenAI announced a letter of intent under which Nvidia would invest up to $100 billion in OpenAI, staged against the deployment of at least 10 gigawatts of Nvidia systems.</p>
<p>Read that structure carefully, because it is the interesting part.</p>
<h2 id="the-circularity">the circularity<a class="anchor" href="#the-circularity" aria-label="link to this section">#</a></h2>
<p>Nvidia invests in OpenAI. OpenAI uses the money to buy Nvidia systems. The purchase is recognized as Nvidia revenue. The investment is staged against deployment milestones.</p>
<p>This is not fraud and it is not unusual in capital-intensive industries — vendor financing has been standard in telecom, aviation, and semiconductor equipment for decades. A supplier finances a customer's purchase because the supplier has the balance sheet and wants the volume.</p>
<p>It does deserve scrutiny for a specific reason: it makes the demand signal less informative. When a supplier funds its customer's purchases, revenue growth no longer cleanly indicates independent market demand. Some portion of it is the supplier's own capital cycling through.</p>
<p>Analysts have been tracking a widening set of these arrangements across the AI sector — investments in customers, prepayments, equity stakes in companies that are also large purchasers. Individually each is defensible. Collectively they make the sector's growth figures harder to interpret.</p>
<h2 id="the-gigawatt-as-a-unit">the gigawatt as a unit<a class="anchor" href="#the-gigawatt-as-a-unit" aria-label="link to this section">#</a></h2>
<p>Note what is being measured. Not chips, not dollars, not FLOPs. <strong>Gigawatts.</strong></p>
<p>Ten gigawatts is on the order of the electricity consumption of a large metropolitan area. It is roughly ten large nuclear reactors' worth of continuous generation.</p>
<p>The industry has converged on power as the natural unit because power is the binding constraint. You can order chips. You cannot order a substation and have it next quarter.</p>
<p>The consequences flow outward: electricity prices in datacenter-heavy regions, grid interconnection <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a>, transmission buildout, and the political economy of who pays for it. Several US utility regulators are now handling rate cases that are effectively about whether residential customers subsidize datacenter connections.</p>
<p>That fight is going to define a lot of the next five years and it is being had in public utility commission hearings that nobody in tech reads.</p>
<h2 id="what-it-means-for-you">what it means for you<a class="anchor" href="#what-it-means-for-you" aria-label="link to this section">#</a></h2>
<p>If you are building on AI APIs, the practical questions are:</p>
<p><strong>Is my provider's capacity growing?</strong> <a class="xref" href="/rate-limits-are-a-product-decision-not-an-infrastructure-one/" title="Rate limits are a product decision, not an infrastructure one">Rate limits</a> and availability during demand spikes are the observable symptom. Announcements like this are a positive signal for capacity.</p>
<p><strong>Am I exposed to a single provider's economics?</strong> If the financing environment tightens, pricing changes. Multi-provider capability is cheap insurance and you should have built it anyway for reliability reasons.</p>
<p><strong>Are my costs actually falling?</strong> Per-token prices have fallen consistently. Per <em>task</em> costs have not fallen as much, because reasoning models consume more tokens. Measure the thing you pay for.</p>
<h2 id="the-honest-uncertainty">the honest uncertainty<a class="anchor" href="#the-honest-uncertainty" aria-label="link to this section">#</a></h2>
<p>Nobody knows whether the capex cycle is correctly sized. The bull case is that inference demand compounds and every gigawatt gets used. The bear case is that efficiency improvements outrun demand and a lot of concrete is stranded.</p>
<p>Both are held sincerely by smart people with access to the same information. That is what genuine uncertainty looks like, and anyone expressing confidence in either direction is telling you about their position, not about the world.</p>]]></content:encoded></item><item><title>iPhone 17, the Air, and a modem that finally isn't Qualcomm's</title><link>https://readme.news/iphone-17-the-air-and-a-modem-that-finally-isnt-qualcomms/</link><guid isPermaLink="true">https://readme.news/iphone-17-the-air-and-a-modem-that-finally-isnt-qualcomms/</guid><pubDate>Fri, 12 Sep 2025 09:00:00 +0000</pubDate><description>A19 Pro, a very thin phone as a technology demonstrator, and Apple&#x27;s first in-house cellular modem shipping at scale.</description><content:encoded><![CDATA[<p>Apple announced the iPhone 17 line this week. The interesting engineering is not in the flagship.</p>
<h2 id="the-c1-modem">the C1 modem<a class="anchor" href="#the-c1-modem" aria-label="link to this section">#</a></h2>
<p>Apple's in-house cellular modem ships in the iPhone Air. This project has been running for roughly a decade — including the acquisition of Intel's modem business in 2019 — and has slipped repeatedly.</p>
<p>Cellular modems are genuinely, notoriously hard. You are implementing a standard that is tens of thousands of pages, that has accumulated twenty years of backward compatibility, that must interoperate with carrier equipment from dozens of vendors deployed in every configuration imaginable, and that must be certified in every market on earth. The failure mode is not a crash; it is a dropped call in one city on one carrier's network under one specific handover condition.</p>
<p>Qualcomm's position has been the strongest moat in mobile for exactly this reason. Apple, with effectively unlimited resources and a decade, has produced a first modem that is competitive on power efficiency and behind on peak throughput. That is a reasonable first-generation result and it tells you how hard the problem is.</p>
<p>The power efficiency angle matters more than peak speed for a phone. Nobody saturates a 5G connection all day; everybody has a battery.</p>
<h2 id="the-air">the Air<a class="anchor" href="#the-air" aria-label="link to this section">#</a></h2>
<p>Extremely thin, single rear camera, smaller battery. As a product it is a trade-off most people will not want. As an engineering statement it is the platform for the C1, the new thermal design, and the packaging techniques that will show up in the rest of the line later.</p>
<p>Apple has done this before — the original MacBook Air was a compromised machine that established a direction the whole industry followed.</p>
<h2 id="a19-pro">A19 Pro<a class="anchor" href="#a19-pro" aria-label="link to this section">#</a></h2>
<p>Faster, more efficient, more neural accelerator throughput. The specific number that matters to developers is memory: <a class="xref" href="/pixel-10-and-the-on-device-model-as-a-platform-feature/" title="Pixel 10 and the on-device model as a platform feature">on-device model</a> work is memory-bound, and the amount of RAM in a phone determines what you can run.</p>
<p>Apple has been conservative here for years for power reasons. On-device AI is the first argument for more RAM that has real product weight behind it, and the trajectory is upward.</p>
<h2 id="the-developer-read">the developer read<a class="anchor" href="#the-developer-read" aria-label="link to this section">#</a></h2>
<p>Two practical items.</p>
<p><strong>On-device model capability is now a spec differentiator across the line.</strong> If you build a feature on Foundation Models, know which devices support it and design the fallback. "Newest phones only" is a real constraint for a consumer app.</p>
<p><strong>The modem transition means cellular behavior will vary by model</strong> in ways it has not for a decade. If your app does anything sensitive to network characteristics — real-time media, aggressive prefetching, connection reuse — test on both modem generations. Field behavior differences in the first year of a new modem are normal and you do not want to discover them through crash reports.</p>
<h2 id="the-pattern-worth-noting">the pattern worth noting<a class="anchor" href="#the-pattern-worth-noting" aria-label="link to this section">#</a></h2>
<p>Apple's vertical integration keeps extending: CPU, GPU, neural engine, now modem, with networking chips also in-house. Each one takes roughly a decade and each one removes a dependency and a margin stack.</p>
<p>The strategic logic has been consistent since 2008 and it keeps paying off. The cost is that when they get one wrong, there is nobody else to blame and no alternative supplier to switch to.</p>]]></content:encoded></item><item><title>Pixel 10 and the on-device model as a platform feature</title><link>https://readme.news/pixel-10-and-the-on-device-model-as-a-platform-feature/</link><guid isPermaLink="true">https://readme.news/pixel-10-and-the-on-device-model-as-a-platform-feature/</guid><pubDate>Thu, 21 Aug 2025 09:00:00 +0000</pubDate><description>Tensor G5 moves to TSMC, and Gemini Nano gets an API surface that third-party apps can actually use.</description><content:encoded><![CDATA[<p>Google announced the Pixel 10 line this week with Tensor G5, the first Tensor chip fabricated by TSMC rather than Samsung.</p>
<h2 id="the-fab-change">the fab change<a class="anchor" href="#the-fab-change" aria-label="link to this section">#</a></h2>
<p>Tensor's history has been a story of good ideas hampered by a manufacturing process that trailed the competition. Thermal throttling, modem power draw, and sustained-performance deficits against Snapdragon and Apple silicon were consistently traceable to process node rather than architecture.</p>
<p>Moving to TSMC's 3nm addresses that directly. Early efficiency numbers look substantially better, which for a phone matters more than peak performance — nobody runs a benchmark all day, everybody runs a battery.</p>
<p>The strategic significance: Google's silicon ambition was always about controlling the ML acceleration path for on-device inference. That only works if the chip is competitive on the fundamentals, and for four generations it was not.</p>
<h2 id="the-developer-surface">the developer surface<a class="anchor" href="#the-developer-surface" aria-label="link to this section">#</a></h2>
<p>The more relevant announcement is that Gemini Nano is exposed to third-party apps through ML Kit's GenAI APIs and the newer on-device inference paths.</p>
<div class="code"><span class="code-lang">kotlin</span><pre><code class="lang-kotlin">val summarizer = Summarization.getClient(
    SummarizerOptions.builder(context)
        .setInputType(InputType.ARTICLE)
        .setOutputType(OutputType.ONE_BULLET)
        .build()
)
val result = summarizer.runInference(text).await()</code></pre></div>
<p>The available primitives — summarization, proofreading, rewriting, image description — are deliberately narrow. That is a reasonable choice: a small on-device model is reliable within a bounded task and unreliable outside it, and shipping a constrained API prevents developers from discovering that the hard way.</p>
<p>The economics are the same as Apple's Foundation Models framework: free, private, offline, no <a class="xref" href="/rate-limits-are-a-product-decision-not-an-infrastructure-one/" title="Rate limits are a product decision, not an infrastructure one">rate limits</a>. For app features that were not worth a server bill, that changes the calculation entirely.</p>
<h2 id="the-convergence">the convergence<a class="anchor" href="#the-convergence" aria-label="link to this section">#</a></h2>
<p>Apple and Google have now independently arrived at the same architecture:</p>
<ul><li>A <a class="xref" href="/haiku-45-and-the-collapsing-cost-of-good-enough/" title="Haiku 4.5 and the collapsing cost of good-enough">small model</a> on device, exposed to third-party apps through a constrained system API.</li><li>A larger model in the cloud for anything the small one cannot handle.</li><li>A routing decision the OS makes, mostly invisible to the app.</li><li>Privacy positioning built on the on-device path.</li></ul>
<p>That is going to be the standard shape of mobile AI. Which means the useful developer skill is not "call an LLM API" but "decompose a feature so the on-device model handles the common case."</p>
<p>Concretely: design your feature so the 3B model handles 90% of inputs and the remaining 10% escalates. The 90% is free, instant, and offline. The 10% costs money and needs a network. Getting that split right is where the engineering is.</p>
<h2 id="the-caveat">the caveat<a class="anchor" href="#the-caveat" aria-label="link to this section">#</a></h2>
<p>On-device models are small and they will stay small, because the constraint is memory and thermal budget in a phone, and those improve slowly. A 3B model in 2027 will be better than a 3B model today, but it will still be a 3B model.</p>
<p>Do not design a feature that requires frontier reasoning and hope the device catches up. It will not. Design for the envelope you have, and escalate explicitly.</p>]]></content:encoded></item><item><title>OpenAI buys a hardware company that hasn't shipped anything</title><link>https://readme.news/openai-buys-a-hardware-company-that-hasnt-shipped-anything/</link><guid isPermaLink="true">https://readme.news/openai-buys-a-hardware-company-that-hasnt-shipped-anything/</guid><pubDate>Tue, 27 May 2025 09:00:00 +0000</pubDate><description>$6.5 billion for io, Jony Ive&#x27;s design studio. The bet is that the phone is the wrong shape for this.</description><content:encoded><![CDATA[<p>OpenAI announced an all-stock acquisition of io, the hardware startup founded by Jony Ive, valued at around $6.5 billion. No product exists. The first device is described as arriving in 2026.</p>
<p>Six and a half billion dollars for a design team and an idea is a lot of money, and it is worth taking the underlying thesis seriously even if you think the price is absurd.</p>
<h2 id="the-thesis">the thesis<a class="anchor" href="#the-thesis" aria-label="link to this section">#</a></h2>
<p>The smartphone is optimized for an interaction model that AI makes obsolete.</p>
<p>A phone is a rectangle of apps. You unlock it, find the app, navigate its hierarchy, and perform a task. That design solved the problem of "how do I access many different services on a small screen," and it solved it well enough that the form factor has been essentially static for fifteen years.</p>
<p>If the interaction model becomes "state your intent, the system figures out which services to use," then most of the phone's design — the grid, the navigation, the visual hierarchy — is scaffolding for a problem that no longer exists. What you need instead is something always available, ambient, primarily audio, with a screen only when a screen is genuinely the right modality.</p>
<p>That is a coherent thesis. It is also the thesis behind two products that failed badly and publicly in 2024: the Humane Ai Pin and the Rabbit R1.</p>
<h2 id="why-those-failed-and-whether-this-is-different">why those failed and whether this is different<a class="anchor" href="#why-those-failed-and-whether-this-is-different" aria-label="link to this section">#</a></h2>
<p>The Ai Pin and R1 failed for the same three reasons:</p>
<ol><li><strong>The models were not good enough.</strong> Both shipped in early 2024, before reliable tool use, before good latency, before models that could handle ambiguity gracefully. The demos were staged; the products were not ready.</li><li><strong>The hardware was bad.</strong> Thermal problems, terrible battery life, laser projection nobody could read in daylight.</li><li><strong>They competed with a phone that was in the same pocket.</strong> Any task the device could not do, the phone could. That is a brutal comparison to survive.</li></ol>
<p>Reason one has changed a lot in eighteen months and will change more by 2026. Reason two is precisely what you buy Jony Ive's team for. Reason three has not changed at all and is, I think, the actual problem.</p>
<h2 id="the-part-that-is-genuinely-hard">the part that is genuinely hard<a class="anchor" href="#the-part-that-is-genuinely-hard" aria-label="link to this section">#</a></h2>
<p>A dedicated AI device has to be better than a phone at <em>something</em> to justify existing. The candidates:</p>
<ul><li><strong>Always-on context.</strong> A device that sees and hears what you do all day has context a phone does not. That is also the most invasive product concept in consumer electronics history, and the social norms around it do not exist.</li><li><strong>Zero-friction capture.</strong> Say something, it is recorded, transcribed, actioned. Genuinely better than unlocking a phone. Also achievable by a watch or earbuds, which people already wear.</li><li><strong>Not being a phone.</strong> There is a real and growing market for devices that do not have a feed. That market is not obviously large enough for a $6.5B bet.</li></ul>
<h2 id="the-strategic-read">the strategic read<a class="anchor" href="#the-strategic-read" aria-label="link to this section">#</a></h2>
<p>OpenAI's dependence on Apple and Google for distribution is a structural risk. Every ChatGPT interaction on a phone happens on an operating system built by a competitor who can change the rules. Owning hardware is the only permanent solution to that, and it is worth a lot to not be a tenant.</p>
<p>Whether it is worth this much, executed by a team that has never shipped hardware at this company, on a two-year timeline, in a category with a hundred percent failure rate so far — that is a different question.</p>
<p>I would not bet against Ive on the object. I would bet against the category on the timeline.</p>]]></content:encoded></item><item><title>GTC 2025: a roadmap to 2027 and a warning about power</title><link>https://readme.news/gtc-2025-a-roadmap-to-2027-and-a-warning-about-power/</link><guid isPermaLink="true">https://readme.news/gtc-2025-a-roadmap-to-2027-and-a-warning-about-power/</guid><pubDate>Wed, 19 Mar 2025 09:00:00 +0000</pubDate><description>Blackwell Ultra this year, Rubin next, Feynman after. Nvidia is now publishing a schedule like a foundry.</description><content:encoded><![CDATA[<p>Jensen Huang did two hours at the San Jose arena and the substance came down to one slide: a named product cadence stretching to 2027.</p>
<ul><li><strong>Blackwell Ultra (GB300)</strong> — second half of 2025, more HBM per package.</li><li><strong>Vera Rubin</strong> — 2026, new CPU (Vera) and new GPU (Rubin), HBM4.</li><li><strong>Rubin Ultra</strong> — 2027, a much larger rack-scale configuration.</li><li><strong>Feynman</strong> — 2028, named, not detailed.</li></ul>
<p>Publishing a multi-year roadmap at this granularity is a foundry move, not a product-company move. The audience is not developers. It is the utilities, the construction firms, the HBM suppliers, and the CFOs who need to plan capital around it.</p>
<h2 id="the-number-that-should-worry-you">the number that should worry you<a class="anchor" href="#the-number-that-should-worry-you" aria-label="link to this section">#</a></h2>
<p>The power figures for the rack-scale systems are the part of this keynote that will still matter in five years. Current NVL72 racks draw on the order of 120 kW. The Rubin Ultra generation is being discussed in the hundreds of kilowatts per rack.</p>
<p>A traditional enterprise datacenter rack is provisioned for 5 to 15 kW. Air cooling tops out somewhere around 40 kW with heroic effort. Everything past that is liquid, and everything past about 150 kW is liquid plus a fundamentally different power distribution architecture — which is why the roadmap includes 800 VDC distribution.</p>
<p>Translation: existing datacenter shells are largely unusable for this. The buildout is not "install new servers," it is "build new buildings near new substations." That is a multi-year, capital-intensive, permit-bound process, and it is the actual rate limiter on AI capacity through the rest of the decade.</p>
<h2 id="the-software-announcements">the software announcements<a class="anchor" href="#the-software-announcements" aria-label="link to this section">#</a></h2>
<p><strong>Dynamo</strong>, an open-source inference serving framework, is the developer-relevant release. It handles disaggregated serving — running the prefill phase and the decode phase on different hardware pools, because they have completely different compute and memory characteristics. Prefill is compute-bound and parallel; decode is memory-bandwidth-bound and sequential. Running them on the same homogeneous pool wastes a lot of silicon.</p>
<p>If you operate inference at any scale, disaggregation is probably the largest single efficiency win available to you right now, and having a maintained open implementation lowers the bar considerably.</p>
<p><strong>NIM microservices</strong> continue to be Nvidia's attempt to own the deployment layer. Reasonable, containerized, and a <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a> vector you should evaluate with clear eyes.</p>
<h2 id="the-strategic-read">the strategic read<a class="anchor" href="#the-strategic-read" aria-label="link to this section">#</a></h2>
<p>Nvidia's actual product is no longer a chip. It is a rack, plus the networking between racks, plus the software stack on top. Selling a GPU means competing with AMD and with every hyperscaler's internal silicon team. Selling an integrated rack-scale system with a co-designed interconnect means competing with nobody, because nobody else has the interconnect.</p>
<p>That is the moat, it is being widened deliberately, and the roadmap slide is the announcement of that strategy rather than a product list.</p>]]></content:encoded></item><item><title>The Mac lineup now requires a spreadsheet</title><link>https://readme.news/the-mac-lineup-now-requires-a-spreadsheet/</link><guid isPermaLink="true">https://readme.news/the-mac-lineup-now-requires-a-spreadsheet/</guid><pubDate>Wed, 05 Mar 2025 09:00:00 +0000</pubDate><description>M4 MacBook Air, M3 Ultra Mac Studio, M4 Max Mac Studio. Apple&#x27;s chip naming has fully decoupled from chronology.</description><content:encoded><![CDATA[<p>Apple refreshed the MacBook Air with M4 and the Mac Studio with a choice of M4 Max or M3 Ultra. Read that last part again.</p>
<p>The Mac Studio's top configuration uses a previous-generation chip family because M4 Ultra does not exist — the M4 Max apparently lacks the interconnect that made the Ultra fusion possible, so the highest-end desktop part is an M3 generation die. This is defensible as engineering and indefensible as naming.</p>
<h2 id="what-the-hardware-actually-is">what the hardware actually is<a class="anchor" href="#what-the-hardware-actually-is" aria-label="link to this section">#</a></h2>
<p><strong>MacBook Air M4.</strong> Better base memory, a 12MP center-stage camera, support for two external displays with the lid open, and a starting price that came down. For a developer machine this is the sweet spot it has been for two years: silent, long battery life, and fast enough that the fanless design is not a compromise for anything short of sustained compilation.</p>
<p><strong>Mac Studio M3 Ultra.</strong> Up to 512 GB of unified memory. That number is the story. There is currently no other desktop workstation where you can put half a terabyte of memory directly addressable by a capable GPU at reasonable bandwidth.</p>
<p>For local model work this is a genuinely unique product. A 512 GB configuration holds a frontier-scale open-weights model in memory at reasonable quantization, running on a machine that sits on a desk and draws a few hundred watts. There is no comparable option. The nearest equivalents are multi-GPU servers that cost more, sound like a jet, and require a dedicated circuit.</p>
<p><strong>Mac Studio M4 Max.</strong> Faster single-thread, fewer memory options, better for most people who are not loading enormous models.</p>
<h2 id="the-memory-bandwidth-caveat">the memory bandwidth caveat<a class="anchor" href="#the-memory-bandwidth-caveat" aria-label="link to this section">#</a></h2>
<p>Capacity is not throughput. The Ultra's memory bandwidth is high by desktop standards and low compared to an H100's HBM. If you load a very large model into 512 GB, you will be able to run it, and it will generate tokens at a rate that is useful for batch work and frustrating for interactive chat.</p>
<p>This is the right trade for a lot of workloads — running evals overnight, processing a corpus, fine-tuning something small — and the wrong trade if you expected an interactive <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> on your desk.</p>
<h2 id="the-naming-problem-seriously">the naming problem, seriously<a class="anchor" href="#the-naming-problem-seriously" aria-label="link to this section">#</a></h2>
<p>Apple now ships M1 through M4, each in base / Pro / Max / Ultra variants, released on staggered schedules, sold simultaneously in different products. The current lineup contains M2, M3 and M4 generation silicon in actively-sold Macs.</p>
<p>The consequence for developers is real: "requires Apple Silicon" is no longer a useful compatibility statement, and neither is "M-series." Feature availability varies by generation — hardware ray tracing arrived in M3, the improved neural engine in M4 — so if you are targeting these capabilities you need to check the generation, not the family.</p>
<p><code>sysctl -n machdep.cpu.brand_string</code> will tell you what you are on. You are going to need it.</p>]]></content:encoded></item><item><title>Grok 3 and the case for buying your way to the frontier</title><link>https://readme.news/grok-3-and-the-case-for-buying-your-way-to-the-frontier/</link><guid isPermaLink="true">https://readme.news/grok-3-and-the-case-for-buying-your-way-to-the-frontier/</guid><pubDate>Tue, 18 Feb 2025 09:00:00 +0000</pubDate><description>xAI&#x27;s third model arrives on a cluster built in months. The interesting claim isn&#x27;t the benchmark, it&#x27;s the schedule.</description><content:encoded><![CDATA[<p>xAI announced Grok 3 last night, along with reasoning variants and a "Big Brain" extended-thinking mode. The benchmark claims put it at or near the frontier across math, science and coding evaluations.</p>
<p>Take the benchmark numbers with the usual salt — they were presented by the vendor, some were shown with consensus-of-N sampling against competitors' single samples, and the field has no agreed protocol for this. The comparison charts in a launch livestream are marketing artifacts.</p>
<p>The genuinely notable thing is not the model. It is Colossus.</p>
<h2 id="the-cluster">the cluster<a class="anchor" href="#the-cluster" aria-label="link to this section">#</a></h2>
<p>xAI built a datacenter in Memphis housing on the order of 100,000 H100-class GPUs, and did it on a timeline measured in months rather than years. The conventional wisdom on a buildout of that scale was eighteen to twenty-four months. They compressed it by doing things that are expensive and unglamorous: bringing in mobile gas turbines for interim power, running their own networking integration, and accepting a lot of operational risk.</p>
<p>Whether you find that admirable or reckless depends on your priors and on how you feel about the air quality complaints from the surrounding neighborhood, which are a real and ongoing dispute worth reading about separately.</p>
<p>But as an engineering datapoint it matters: it establishes that the time constant for standing up frontier-scale compute is shorter than the industry assumed, if you are willing to spend and to eat the risk.</p>
<h2 id="what-that-implies">what that implies<a class="anchor" href="#what-that-implies" aria-label="link to this section">#</a></h2>
<p>If a well-capitalized new entrant can go from nothing to frontier-scale compute in about a year, then compute is not a durable moat. It is a capital requirement, which is a different thing. Capital requirements keep out the under-funded; they do not keep out the well-funded.</p>
<p>Which pushes the question of where the actual moat is:</p>
<ul><li><strong>Data</strong> — increasingly contested, increasingly litigated, and the frontier labs are all converging on similar synthetic-data-plus-RL recipes anyway.</li><li><strong>Talent</strong> — mobile, expensive, and being bid on aggressively.</li><li><strong>Distribution</strong> — this one is real. A model inside a product a billion people already open is worth more than a marginally better model behind a signup form.</li><li><strong>Cost per token at quality</strong> — real, and derived from architecture and serving engineering rather than raw scale.</li></ul>
<p>My read is that distribution and serving efficiency are the durable ones, and that is a much less romantic answer than "we have the smartest model."</p>
<h2 id="for-people-who-ship-things">for people who ship things<a class="anchor" href="#for-people-who-ship-things" aria-label="link to this section">#</a></h2>
<p>Practical implication: assume model quality converges and plan accordingly. Do not build your product on the assumption that one vendor's model stays two steps ahead. Build the abstraction layer, keep your evals in your own repo, and make switching a config change.</p>
<p>The teams that did this in 2024 spent a boring week on it and have been changing providers casually ever since. The teams that did not are having architecture meetings.</p>]]></content:encoded></item><item><title>The day the market priced in efficiency</title><link>https://readme.news/the-day-the-market-priced-in-efficiency/</link><guid isPermaLink="true">https://readme.news/the-day-the-market-priced-in-efficiency/</guid><pubDate>Mon, 27 Jan 2025 09:00:00 +0000</pubDate><description>Nvidia lost roughly $600 billion of market value in a session. The trigger was a paper about training costs.</description><content:encoded><![CDATA[<p>Nvidia closed down about 17% today. Broadcom, Vertiv, Constellation Energy and most of the power-adjacent complex fell with it. The proximate cause was a week-old model release from a Chinese lab and a number in its technical report.</p>
<h2 id="the-number">the number<a class="anchor" href="#the-number" aria-label="link to this section">#</a></h2>
<p>DeepSeek's V3 paper stated a final training run cost of roughly $5.6 million in GPU-hours. That figure travelled around the world in about four days, usually stripped of every qualifier attached to it.</p>
<p>The qualifiers matter enormously:</p>
<ul><li>It is the cost of the <strong>final run only</strong>. It explicitly excludes research, failed runs, ablations, and data pipeline work — which in any frontier lab is the overwhelming majority of total spend.</li><li>It excludes the <strong>capital cost of the cluster</strong> itself.</li><li>It says nothing about <strong>R1's</strong> RL training, which came later and separately.</li></ul>
<p>So "they trained a <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> for $5.6M" is not what the paper says. It is what the internet decided the paper says.</p>
<h2 id="why-the-market-reacted-anyway">why the market reacted anyway<a class="anchor" href="#why-the-market-reacted-anyway" aria-label="link to this section">#</a></h2>
<p>Because the directionally correct read survives the correction, and the market is a machine for reacting to directionally correct reads badly.</p>
<p>The directionally correct read: architectural efficiency gains are real and large. DeepSeek's mixture-of-experts design activates a small fraction of total parameters per token. Their multi-head latent attention cuts KV cache size substantially. Their FP8 training pipeline halves memory traffic against BF16. These are engineering wins, they are published, and they are reproducible.</p>
<p>If capability-per-FLOP is improving that fast, then the number of FLOPs you need to buy to reach a given capability is falling. That is the thesis that repriced today.</p>
<h2 id="the-counter-thesis">the counter-thesis<a class="anchor" href="#the-counter-thesis" aria-label="link to this section">#</a></h2>
<p>Jevons. If compute gets cheaper per unit of capability, you do not buy less of it — you find more things to do with it. Reasoning models in particular consume enormous inference compute; a model that thinks for thirty seconds before answering is a very different demand curve than one that answers immediately. Efficiency gains in training get spent on inference.</p>
<p>Both theses are defensible. The honest answer is that nobody knows the shape of the demand curve, and a 17% single-day move in the largest company on earth is not a considered judgement about that. It is a positioning unwind.</p>
<h2 id="for-engineers-specifically">for engineers specifically<a class="anchor" href="#for-engineers-specifically" aria-label="link to this section">#</a></h2>
<p>The useful lesson has nothing to do with stock prices. It is this: the performance-per-dollar frontier is moving fast enough that any architecture decision you make today assuming current inference costs will be wrong within a year, in your favor.</p>
<p>Do not build elaborate <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> and routing infrastructure to shave token costs that are going to fall by an order of magnitude anyway. Build the thing. Measure it. Optimize when the bill actually hurts.</p>
<p>That advice would have been wrong in most previous computing eras. It is right now, and it will stop being right at some point, and watching for that moment is most of the job.</p>]]></content:encoded></item><item><title>CES 2025 was a GPU keynote in a robot costume</title><link>https://readme.news/ces-2025-was-a-gpu-keynote-in-a-robot-costume/</link><guid isPermaLink="true">https://readme.news/ces-2025-was-a-gpu-keynote-in-a-robot-costume/</guid><pubDate>Mon, 13 Jan 2025 09:00:00 +0000</pubDate><description>RTX 50-series, a desktop supercomputer named DIGITS, and an entire show floor that decided AI is a hardware category now.</description><content:encoded><![CDATA[<p>The convention center in Las Vegas hosted approximately nine thousand companies last week and roughly all of them said the word "AI" from a stage. The only part that will still matter in six months came from Nvidia.</p>
<h2 id="rtx-50-series">RTX 50-series<a class="anchor" href="#rtx-50-series" aria-label="link to this section">#</a></h2>
<p>Blackwell hits consumer cards. The headline specs are real — GDDR7, more memory bandwidth, a genuinely improved encoder — but the marketing number is not a rasterization number. It is a DLSS 4 number, and DLSS 4 introduces multi-frame generation, which produces up to three interpolated frames per rendered frame.</p>
<p>Comparing that against last generation's single generated frame and printing the ratio as a performance multiple is, let us be diplomatic, a choice. The underlying rendering improvement generation-over-generation is meaningful but nothing like the number on the slide. If you are buying a card to run a compiler and occasionally a game, look for pure raster benchmarks and mentally discard everything with "AI" in the chart title.</p>
<p>What is unambiguously good: the memory. For anyone running local models, VRAM is the binding constraint and always has been. More of it at higher bandwidth changes what fits.</p>
<h2 id="project-digits">Project DIGITS<a class="anchor" href="#project-digits" aria-label="link to this section">#</a></h2>
<p>The more interesting announcement, because it is a new shape of thing: a desktop-sized machine built on a GB10 Grace Blackwell part with 128 GB of unified memory, aimed squarely at developers who want to prototype against large models without renting an H100 by the hour.</p>
<p>The pitch is that you develop locally on the same CUDA stack you deploy to. That has always been Nvidia's actual moat — not the silicon, the software continuity — and this is that moat extended down to a box on your desk. The stated price is around $3,000, which sounds like a lot until you price four months of cloud GPU.</p>
<p>Two open questions: sustained thermal behavior in that form factor, and whether the memory bandwidth is enough to make the capacity useful rather than theoretical. Capacity without bandwidth means you can load the model and then watch it think very slowly.</p>
<h2 id="everything-else">everything else<a class="anchor" href="#everything-else" aria-label="link to this section">#</a></h2>
<p>The show floor's other theme was putting a language model into objects that did not ask for one. An AI-enabled grill. An AI-enabled crib. A smart ring for your dog. This is the phase of a hype cycle where the technology gets stapled to inventory, and it is worth remembering it happened with "smart," with "cloud," and with "blockchain."</p>
<p>The signal in the noise: robotics is quietly getting real. Not the humanoids on the keynote stage — those are demos with a lot of teleoperation behind them — but the boring warehouse and inspection stuff that has actual unit economics. That is where the useful engineering jobs will be.</p>
<p>Meanwhile, three separate companies announced a foldable laptop. Nobody asked for a foldable laptop.</p>]]></content:encoded></item><item><title>The year the model comes to the laptop</title><link>https://readme.news/the-year-the-model-comes-to-the-laptop/</link><guid isPermaLink="true">https://readme.news/the-year-the-model-comes-to-the-laptop/</guid><pubDate>Fri, 03 Jan 2025 09:00:00 +0000</pubDate><description>2025 opens with open weights good enough to matter and consumer hardware finally sized for them.</description><content:encoded><![CDATA[<p>Every January someone declares the Year of the Linux Desktop and everyone laughs. Let me try a less funny one: 2025 is the year a genuinely useful model runs on the machine already on your desk, and a meaningful number of developers stop paying per token for the boring half of their work.</p>
<p>The argument is not that local models will beat frontier models. They will not. The argument is that most of what a developer asks a model to do is not frontier work. It is renaming things. It is writing the test you already know the shape of. It is summarizing a stack trace. It is converting JSON to a struct. That band of work fell below the capability line of a 7B-to-30B open-weights model somewhere in the second half of 2024, and nobody sent a memo.</p>
<h2 id="what-actually-changed">what actually changed<a class="anchor" href="#what-actually-changed" aria-label="link to this section">#</a></h2>
<p>Three things converged, and none of them were a single dramatic release.</p>
<p><strong>Quantization stopped being lossy in ways you notice.</strong> Four-bit and five-bit weight formats went from "acceptable for a demo" to "I cannot tell in normal use" for most instruct-tuned models. The k-quant family and the newer importance matrix approaches put a 14B model in roughly 9 GB. That fits.</p>
<p><strong>Unified memory became normal.</strong> A machine with 32 GB of memory shared between CPU and GPU is now unremarkable. Apple Silicon got there first and made it a category. The PC side is following with soldered LPDDR5X in thin laptops, which developers will complain about right up until they load a 30B model on a plane.</p>
<p><strong>The runtimes got boring.</strong> <code>llama.cpp</code>, Ollama, LM Studio, MLX. Boring is the compliment. You do not compile anything. You do not fight CUDA versions. You run one command and there is an OpenAI-shaped endpoint on <code>localhost:11434</code>.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">ollama run qwen2.5-coder:14b
# or point your existing client at it and change nothing else
export OPENAI_BASE_URL=http://localhost:11434/v1
export OPENAI_API_KEY=whatever</code></pre></div>
<p>That last line is the whole story. The local ecosystem won by adopting somebody else's API shape instead of inventing a better one.</p>
<h2 id="what-this-does-not-solve">what this does not solve<a class="anchor" href="#what-this-does-not-solve" aria-label="link to this section">#</a></h2>
<p>Latency to first token on a cold model is bad. Context windows on local models are smaller and the quality degrades faster as you fill them. Long agentic loops with tool calls will still route to a hosted <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> for a while, because the difference in instruction-following over twenty turns is enormous and obvious.</p>
<p>And the honest failure mode: people benchmark a local model on the tasks it is good at, feel great, then use it for something requiring real reasoning and conclude the whole category is fake. Both halves of that are wrong.</p>
<h2 id="the-practical-setup">the practical setup<a class="anchor" href="#the-practical-setup" aria-label="link to this section">#</a></h2>
<p>Run a <a class="xref" href="/haiku-45-and-the-collapsing-cost-of-good-enough/" title="Haiku 4.5 and the collapsing cost of good-enough">small model</a> locally for completion, commit messages, and quick shell questions. Keep a frontier key for design work, gnarly debugging, and anything touching a large codebase. Route between them explicitly rather than hoping some product does it for you.</p>
<p>The interesting second-order effect is privacy. A local model means the code never leaves the machine, which means legal stops being the blocker for a whole category of shops that have spent two years saying no. That is going to move more adoption in 2025 than any benchmark.</p>
<p>Prediction, for the record, to be graded in December: by the end of the year the default developer setup is hybrid, and arguing about it will feel as dated as arguing about tabs.</p>]]></content:encoded></item>
</channel>
</rss>
