<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — open-weights</title>
<link>https://readme.news/tags/open-weights/</link>
<atom:link href="https://readme.news/tags/open-weights/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged open-weights.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>The open weights year</title><link>https://readme.news/the-open-weights-year/</link><guid isPermaLink="true">https://readme.news/the-open-weights-year/</guid><pubDate>Fri, 19 Dec 2025 09:00:00 +0000</pubDate><description>Twelve months that took open models from interesting to unavoidable, and where the gap actually sits.</description><content:encoded><![CDATA[<p>January opened with an MIT-licensed reasoning model that repriced the entire sector in a week. December closes with open weights as a normal, boring option in any serious architecture discussion.</p>
<p>Here is the year, and what it means for the next one.</p>
<h2 id="the-releases-that-mattered">the releases that mattered<a class="anchor" href="#the-releases-that-mattered" aria-label="link to this section">#</a></h2>
<p><strong>DeepSeek R1</strong> (January). MIT license, published training methodology, distilled variants that ran on consumer hardware. The RL-on-verifiable-rewards recipe was reproduced widely within weeks.</p>
<p><strong>Qwen3</strong> (April). Eight models, Apache 2.0, MoE variants with excellent quality-per-active-parameter, 119 languages.</p>
<p><strong><a class="xref" href="/kimi-k2-is-a-trillion-parameter-open-weights-release/" title="Kimi K2 is a trillion-parameter open weights release">Kimi K2</a></strong> (July). A trillion parameters, open weights, tuned for agentic tool use, with a genuinely novel training stability contribution.</p>
<p><strong>gpt-oss</strong> (August). OpenAI's first open weights since 2019, Apache 2.0, with a 20B variant that runs on a laptop.</p>
<p>Plus continuous releases from Mistral, Zhipu, MiniMax, Meta, Microsoft, Google, and a long tail of fine-tunes.</p>
<h2 id="where-the-gap-actually-is">where the gap actually is<a class="anchor" href="#where-the-gap-actually-is" aria-label="link to this section">#</a></h2>
<p>The frontier-to-open gap held at roughly six to twelve months all year. That stability is the most important finding, because it means open weights are not converging on the frontier and are not falling behind — they are trailing at a fixed distance.</p>
<p>But "six months behind" undersells the practical position, because the gap is not uniform:</p>
<p><strong>Nearly closed:</strong> code completion, summarization, extraction, classification, translation, structured output, straightforward tool use. For these, a good open model is not meaningfully worse than a <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a>, and it costs a fraction.</p>
<p><strong>Meaningfully behind:</strong> long-horizon agentic work, complex multi-step reasoning, instruction following over many turns, reliability at the tail. This is where frontier models earn their price.</p>
<p><strong>Not comparable:</strong> anything requiring the surrounding infrastructure — enterprise support, uptime guarantees, safety tooling, indemnification. Open weights give you the model and nothing else.</p>
<h2 id="what-changed-structurally">what changed structurally<a class="anchor" href="#what-changed-structurally" aria-label="link to this section">#</a></h2>
<p><strong>Licensing got genuinely permissive.</strong> Two years ago "open" meant a research license with a prohibited-use list. Now the frontier of open releases is Apache 2.0 and MIT. That is a real change and it removed the legal review that was blocking adoption.</p>
<p><strong>The runtime story got boring.</strong> Ollama, llama.cpp, vLLM, MLX, LM Studio. One command. An OpenAI-compatible endpoint. The friction that kept open models in the enthusiast category is gone.</p>
<p><strong>Hosting became competitive.</strong> Multiple providers serve open models at prices well below frontier API rates, with real SLAs. You can use open weights without running anything.</p>
<p><strong>The center of gravity moved east.</strong> The most capable, most permissively licensed open releases came predominantly from Chinese labs. That is a strategic fact with policy consequences that are being worked out loudly and mostly unproductively.</p>
<h2 id="the-practical-architecture-for-next-year">the practical architecture for next year<a class="anchor" href="#the-practical-architecture-for-next-year" aria-label="link to this section">#</a></h2>
<p>The shape that makes sense:</p>
<ul><li><strong>Open weights, self-hosted or on a cheap provider</strong>, for high-volume, well-defined tasks. Classification, extraction, embedding, first-pass drafting.</li><li><strong>Frontier API</strong> for the hard tail: planning, complex reasoning, anything customer-facing where a bad answer is expensive.</li><li><strong>A router</strong> deciding between them, with the escalation rate instrumented.</li><li><strong>Your own eval set</strong> in your own repository, run against every candidate.</li></ul>
<p>That is not a hedge. It is what the cost and capability curves actually imply.</p>
<h2 id="the-prediction">the prediction<a class="anchor" href="#the-prediction" aria-label="link to this section">#</a></h2>
<p>The gap holds at roughly six to twelve months through next year. Open weights absorb an increasing share of production workload by volume while frontier models keep the high-value tail.</p>
<p>The interesting question is not capability. It is whether the labs currently releasing weights continue to, and that is a business decision that could change in either direction with one quarter's strategy review.</p>
<p>Download the ones you care about. They cannot be un-released.</p>]]></content:encoded></item><item><title>OpenAI ships open weights for the first time since GPT-2</title><link>https://readme.news/openai-ships-open-weights-for-the-first-time-since-gpt-2/</link><guid isPermaLink="true">https://readme.news/openai-ships-open-weights-for-the-first-time-since-gpt-2/</guid><pubDate>Tue, 05 Aug 2025 09:00:00 +0000</pubDate><description>gpt-oss-120b and gpt-oss-20b under Apache 2.0. A strategic reversal, six years late, and genuinely useful.</description><content:encoded><![CDATA[<p>OpenAI released <code>gpt-oss-120b</code> and <code>gpt-oss-20b</code> today under Apache 2.0. These are the company's first open weights models since GPT-2 in 2019.</p>
<h2 id="the-specs">the specs<a class="anchor" href="#the-specs" aria-label="link to this section">#</a></h2>
<p>Both are mixture-of-experts with configurable reasoning effort:</p>
<ul><li><strong>gpt-oss-120b</strong> — 117B total parameters, ~5.1B active. Fits on a single 80 GB accelerator with the provided MXFP4 quantization.</li><li><strong>gpt-oss-20b</strong> — 21B total, ~3.6B active. Runs on a machine with 16 GB of memory.</li></ul>
<p>That second one is the important number. 16 GB is a well-specified laptop. A reasoning model with tool use and a configurable <a class="xref" href="/gemini-25-goes-generally-available-with-a-thinking-dial/" title="Gemini 2.5 goes generally available with a thinking dial">thinking budget</a> that runs on a laptop, under Apache 2.0, from OpenAI, is a sentence that would have been implausible eighteen months ago.</p>
<p>Reasoning effort is set in the system prompt — low, medium, or high — which is a cruder <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> than a token budget but works.</p>
<h2 id="the-format">the format<a class="anchor" href="#the-format" aria-label="link to this section">#</a></h2>
<p>The models use a "harmony" response format with structured channels separating analysis, commentary, and final output. You have to render it correctly or the model behaves poorly. The reference implementations handle it; if you are writing your own serving path, read the format spec first rather than debugging it later.</p>
<p>They support tool use — browsing and Python execution — natively, and function calling in the standard shape.</p>
<h2 id="why-now">why now<a class="anchor" href="#why-now" aria-label="link to this section">#</a></h2>
<p>Three reasons, in descending order of how much anyone will admit them.</p>
<p><strong>Competitive pressure.</strong> The open weights frontier is currently defined by DeepSeek, Qwen, Moonshot, and Mistral. A company named OpenAI having no open models had become a running joke and, more importantly, a strategic gap — the developers building on open weights were building on someone else's ecosystem.</p>
<p><strong>Policy positioning.</strong> There is an active regulatory conversation about open models. Being a participant with skin in the game is worth more than commenting from the sidelines.</p>
<p><strong>The capability gap has closed enough to be safe and stayed wide enough to be commercial.</strong> Releasing a model at roughly o3-mini level costs OpenAI little in API revenue — the customers who need frontier capability still need it — and buys a lot of ecosystem.</p>
<h2 id="how-good-are-they">how good are they<a class="anchor" href="#how-good-are-they" aria-label="link to this section">#</a></h2>
<p>Genuinely competitive in their size class. The 120b is roughly comparable to o3-mini on reasoning benchmarks; the 20b is close to o3-mini on several and weaker on others.</p>
<p>The caveats are the usual ones for open weights: they hallucinate more than frontier models, the safety tuning is more easily removed by fine-tuning (OpenAI published research on this specifically), and benchmark performance overstates real-world reliability.</p>
<h2 id="what-to-do-with-them">what to do with them<a class="anchor" href="#what-to-do-with-them" aria-label="link to this section">#</a></h2>
<p>The 20b is the interesting one for most developers. Concretely:</p>
<ul><li><strong>Local coding assistance</strong> with no data leaving the machine. This clears the legal review that blocks hosted models at a lot of companies.</li><li><strong>Batch processing</strong> where you have a lot of documents and API costs add up. Run it on a spot instance overnight.</li><li><strong>Fine-tuning for a narrow domain.</strong> Apache 2.0 means you can, and a fine-tuned 20b on a specific task frequently beats a general <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> on that task, for a fraction of the cost.</li></ul>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">ollama run gpt-oss:20b</code></pre></div>
<p>That is the whole setup. The ecosystem picked it up within hours of release, which is itself a demonstration of why open weights are worth releasing.</p>
<h2 id="the-broader-point">the broader point<a class="anchor" href="#the-broader-point" aria-label="link to this section">#</a></h2>
<p>Two years ago the argument was "open models are dangerously behind or dangerously capable, pick one." The answer turned out to be: they trail the frontier by roughly six to twelve months, that gap is stable, and the world has not ended.</p>
<p>The gap being stable is the interesting finding. It means open weights are permanently a viable option for anything that does not need the absolute frontier — which is most things.</p>]]></content:encoded></item><item><title>Kimi K2 is a trillion-parameter open weights release</title><link>https://readme.news/kimi-k2-is-a-trillion-parameter-open-weights-release/</link><guid isPermaLink="true">https://readme.news/kimi-k2-is-a-trillion-parameter-open-weights-release/</guid><pubDate>Mon, 14 Jul 2025 09:00:00 +0000</pubDate><description>Moonshot ships a 1T-parameter MoE with 32B active, tuned for agentic tool use, with weights you can download.</description><content:encoded><![CDATA[<p>Moonshot AI released Kimi K2: a mixture-of-experts model with roughly one trillion total parameters and 32 billion active per token, with open weights and a modified-MIT license.</p>
<p>A trillion-parameter open weights release is a milestone regardless of what you think of the benchmarks.</p>
<h2 id="the-architecture">the architecture<a class="anchor" href="#the-architecture" aria-label="link to this section">#</a></h2>
<p>384 experts, 8 selected per token, 32B active. The design point is explicit: get the knowledge capacity of a very large model with the inference cost of a mid-sized one.</p>
<p>The training used <strong>MuonClip</strong>, a variant of the Muon optimizer with a QK-clipping mechanism to prevent attention logit explosion. The reported claim is zero loss spikes across the entire pretraining run on 15.5 trillion tokens.</p>
<p>If you have not run large pretraining: loss spikes are the recurring nightmare. A run destabilizes, you roll back to a checkpoint, you lose days of compute, and diagnosing why is largely folklore. A stability technique that actually works is worth more to the field than a benchmark point.</p>
<h2 id="the-agentic-focus">the agentic focus<a class="anchor" href="#the-agentic-focus" aria-label="link to this section">#</a></h2>
<p>K2 was post-trained specifically for tool use, on synthetic multi-step tool-use trajectories generated at scale. The evaluation emphasis is agentic coding and tool-calling benchmarks rather than conversational quality.</p>
<p>That focus is the right read of where the demand is. The commercially interesting use of a model in 2025 is not answering questions, it is executing multi-step tasks with tools, and models tuned for chat are frequently worse at it than their raw capability suggests.</p>
<h2 id="the-practical-problem">the practical problem<a class="anchor" href="#the-practical-problem" aria-label="link to this section">#</a></h2>
<p>You cannot run this on a workstation. A trillion parameters at 8-bit is a terabyte of weights. Even heavily quantized you are looking at multiple high-memory GPUs or a very large server.</p>
<p>So "open weights" here means something different than it does for a 30B model. It means:</p>
<ul><li><strong>Hosting providers can serve it</strong>, and several did within days, at prices well below frontier API rates.</li><li><strong>Companies with infrastructure can run it privately</strong>, which is the point for regulated industries.</li><li><strong>Researchers can study it</strong>, which is the underrated benefit. Interpretability work on frontier-scale models has been limited to whoever works at a frontier lab. It does not have to be.</li></ul>
<h2 id="the-license">the license<a class="anchor" href="#the-license" aria-label="link to this section">#</a></h2>
<p>Modified MIT with an attribution clause above certain usage thresholds. Not strictly OSI-compatible, much closer to open than most "open" model licenses, and substantially more permissive than the Llama <a class="xref" href="/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/" title="Llama 4 arrives, and the leaderboard problem gets a name">community license</a>.</p>
<p>The trend line here is good. Two years ago open weights meant a research-only license with a list of prohibited uses. Now the frontier of open releases is MIT and Apache 2.0 with narrow carve-outs.</p>
<h2 id="the-pattern-nobody-should-miss">the pattern nobody should miss<a class="anchor" href="#the-pattern-nobody-should-miss" aria-label="link to this section">#</a></h2>
<p>The most permissively licensed, largest, most capable open weights models are overwhelmingly coming from Chinese labs — DeepSeek, Qwen, Moonshot, Zhipu, MiniMax. Western open weights releases have been smaller and more restrictively licensed.</p>
<p>The strategic logic is not complicated: if you are behind on distribution, you compete on openness. It worked for Meta in 2023 and it is working now.</p>
<p>The practical consequence for a developer is that your best option for a private, self-hosted, high-capability model is increasingly a Chinese release. Evaluate it on your own tasks, run it in your own infrastructure, and make the decision on engineering grounds.</p>]]></content:encoded></item><item><title>Qwen3 ships a whole family under Apache 2.0</title><link>https://readme.news/qwen3-ships-a-whole-family-under-apache-20/</link><guid isPermaLink="true">https://readme.news/qwen3-ships-a-whole-family-under-apache-20/</guid><pubDate>Tue, 29 Apr 2025 09:00:00 +0000</pubDate><description>Eight models from 0.6B to 235B, hybrid reasoning modes, and a genuinely permissive license across all of them.</description><content:encoded><![CDATA[<p>Alibaba released Qwen3 today: eight models spanning 0.6B to 235B parameters, including two mixture-of-experts variants, all under Apache 2.0.</p>
<p>The license is the headline. Apache 2.0 with no user threshold, no naming requirement, and no acceptable-use addendum that functions as a license restriction. You can fine-tune it, ship it, sell it, and never mention where it came from.</p>
<h2 id="the-lineup">the lineup<a class="anchor" href="#the-lineup" aria-label="link to this section">#</a></h2>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">model</th><th style="text-align:left">type</th><th style="text-align:left">active params</th></tr></thead><tbody><tr><td style="text-align:left">Qwen3-0.6B → 32B</td><td style="text-align:left">dense</td><td style="text-align:left">all</td></tr><tr><td style="text-align:left">Qwen3-30B-A3B</td><td style="text-align:left">MoE</td><td style="text-align:left">3B</td></tr><tr><td style="text-align:left">Qwen3-235B-A22B</td><td style="text-align:left">MoE</td><td style="text-align:left">22B</td></tr></tbody></table></div>
<p>The MoE variants are the interesting ones. <code>30B-A3B</code> has 30 billion total parameters and activates 3 billion per token. That means memory footprint of a 30B model and inference cost closer to a 3B model. On a machine with enough RAM to hold it, throughput is dramatically better than a dense model of comparable quality.</p>
<p>For local deployment this is the shape that matters. Memory is cheap and getting cheaper; compute per token is the thing you feel.</p>
<h2 id="hybrid-thinking">hybrid thinking<a class="anchor" href="#hybrid-thinking" aria-label="link to this section">#</a></h2>
<p>Every model supports two modes: thinking and non-thinking, switchable per request. You can also switch mid-conversation with a token in the prompt — <code>/think</code> and <code>/no_think</code> — which is a hacky <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> and also extremely practical.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True,
    enable_thinking=True,     # or False
)</code></pre></div>
<p>The design conclusion the field has reached, from three different labs independently, is that reasoning should be a runtime dial rather than a model choice. That is now clearly correct and I expect it to be universal within a year.</p>
<h2 id="multilingual-coverage">multilingual coverage<a class="anchor" href="#multilingual-coverage" aria-label="link to this section">#</a></h2>
<p>119 languages and dialects, which is a genuinely differentiating number. Most open models are strong in English, decent in a handful of European languages, and degrade sharply after that. If you are building for a market outside the usual list, this is worth evaluating specifically.</p>
<h2 id="the-geopolitical-thing">the geopolitical thing<a class="anchor" href="#the-geopolitical-thing" aria-label="link to this section">#</a></h2>
<p>A large share of the best open-weights models now come from Chinese labs — Qwen, DeepSeek, GLM, Kimi, MiniMax. That is a real shift from two years ago when open weights meant Llama and Mistral.</p>
<p>The engineering reason is straightforward: releasing weights is a good strategy when you are not the market leader, because it buys mindshare, ecosystem, and research feedback that you cannot get otherwise. Meta understood this in 2023. Chinese labs understand it now.</p>
<p>The policy conversation around this is loud and mostly not technical. What is technical, and worth being clear about: weights are weights. A downloaded model runs on your hardware, in your VPC, with no network egress. The supply chain questions worth asking are about <em>what the model does</em> — evaluate it for backdoored behavior, test it on your own adversarial cases — not about where the gradient descent happened.</p>
<h2 id="what-to-actually-do">what to actually do<a class="anchor" href="#what-to-actually-do" aria-label="link to this section">#</a></h2>
<p>If you are running anything local, the 30B-A3B is the current best quality-per-watt option for a workstation. If you are fine-tuning, Apache 2.0 removes the legal review that was blocking you.</p>
<p>Download it before someone decides you cannot.</p>]]></content:encoded></item><item><title>Llama 4 arrives, and the leaderboard problem gets a name</title><link>https://readme.news/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/</link><guid isPermaLink="true">https://readme.news/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/</guid><pubDate>Mon, 07 Apr 2025 09:00:00 +0000</pubDate><description>Scout and Maverick ship with a 10M-token context claim and an arena entry that wasn&#x27;t the released model.</description><content:encoded><![CDATA[<p>Meta released Llama 4 over the weekend: <strong>Scout</strong> (17B active parameters, 16 experts, a claimed 10-million-token context window) and <strong>Maverick</strong> (17B active, 128 experts), both mixture-of-experts, both released under the Llama community license. A larger <strong>Behemoth</strong> was described as still training.</p>
<p>Within forty-eight hours the release turned into a story about benchmark integrity instead.</p>
<h2 id="what-happened">what happened<a class="anchor" href="#what-happened" aria-label="link to this section">#</a></h2>
<p>Maverick posted a very strong score on LMArena, the human-preference leaderboard. It then emerged that the model evaluated on the arena was an "experimental chat version" tuned for conversationality — not the checkpoint released to the public. LMArena updated its policies and published the disputed comparison. Meta's response was that experimental variants are normal and the arena version was labeled.</p>
<p>Both of those things can be true and the outcome is still bad, because the number that traveled was attached to a model nobody could download.</p>
<h2 id="why-this-keeps-happening">why this keeps happening<a class="anchor" href="#why-this-keeps-happening" aria-label="link to this section">#</a></h2>
<p>Leaderboards are the only shared vocabulary the field has, and they are being asked to carry weight they cannot bear.</p>
<ul><li><strong>Human preference arenas</strong> measure whether people like the answer. That correlates with quality and also with formatting, length, confidence, and sycophancy. A model tuned to be agreeable climbs.</li><li><strong>Static benchmarks</strong> leak into training data. Every popular benchmark is on the internet and every <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> has read the internet. Contamination is not always deliberate and it is essentially always present.</li><li><strong>Vendor-run evaluations</strong> use vendor-chosen settings. Consensus-of-64 for yours, single-sample for theirs. Both numbers are real and the comparison is not.</li></ul>
<p>There is no fix that survives contact with commercial incentives. The only durable answer is that you have to run your own evaluation on your own task.</p>
<h2 id="the-10-million-token-claim">the 10 million token claim<a class="anchor" href="#the-10-million-token-claim" aria-label="link to this section">#</a></h2>
<p>Scout's context window is stated at 10M tokens, achieved through an interleaved attention scheme without positional embeddings in some layers, plus inference-time temperature scaling on attention. It was trained on far shorter sequences and generalizes upward.</p>
<p>Take this as an upper bound on what the architecture accepts, not on what it usefully processes. Independent long-context evaluations found substantial degradation well before that number. That is not unique to Llama 4 — it is true of every long-context claim — but 10M is a big enough number that the gap between accepted and useful is enormous.</p>
<h2 id="what-is-actually-good-here">what is actually good here<a class="anchor" href="#what-is-actually-good-here" aria-label="link to this section">#</a></h2>
<p>The MoE architecture with 17B active parameters is a real efficiency story. Maverick's quality-per-active-parameter is strong, and active parameters are what determine inference cost. A model you can serve at 17B economics with much better quality than a 17B dense model is a genuinely useful thing to have.</p>
<p>Native multimodality — images in the base model rather than bolted on via an adapter — is also the right architecture and is where everyone is heading.</p>
<p>The license is still not open source. "Llama community license" has an MAU threshold and naming requirements. Call it open weights, which it is, and do not call it open source, which it is not.</p>
<h2 id="the-practical-advice">the practical advice<a class="anchor" href="#the-practical-advice" aria-label="link to this section">#</a></h2>
<p>Build an eval harness with fifty examples from your actual domain. Run every candidate model against it. Store the results in your repo next to your tests.</p>
<p>It will take a day and it will make you immune to this entire genre of news.</p>]]></content:encoded></item><item><title>The day the market priced in efficiency</title><link>https://readme.news/the-day-the-market-priced-in-efficiency/</link><guid isPermaLink="true">https://readme.news/the-day-the-market-priced-in-efficiency/</guid><pubDate>Mon, 27 Jan 2025 09:00:00 +0000</pubDate><description>Nvidia lost roughly $600 billion of market value in a session. The trigger was a paper about training costs.</description><content:encoded><![CDATA[<p>Nvidia closed down about 17% today. Broadcom, Vertiv, Constellation Energy and most of the power-adjacent complex fell with it. The proximate cause was a week-old model release from a Chinese lab and a number in its technical report.</p>
<h2 id="the-number">the number<a class="anchor" href="#the-number" aria-label="link to this section">#</a></h2>
<p>DeepSeek's V3 paper stated a final training run cost of roughly $5.6 million in GPU-hours. That figure travelled around the world in about four days, usually stripped of every qualifier attached to it.</p>
<p>The qualifiers matter enormously:</p>
<ul><li>It is the cost of the <strong>final run only</strong>. It explicitly excludes research, failed runs, ablations, and data pipeline work — which in any frontier lab is the overwhelming majority of total spend.</li><li>It excludes the <strong>capital cost of the cluster</strong> itself.</li><li>It says nothing about <strong>R1's</strong> RL training, which came later and separately.</li></ul>
<p>So "they trained a <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> for $5.6M" is not what the paper says. It is what the internet decided the paper says.</p>
<h2 id="why-the-market-reacted-anyway">why the market reacted anyway<a class="anchor" href="#why-the-market-reacted-anyway" aria-label="link to this section">#</a></h2>
<p>Because the directionally correct read survives the correction, and the market is a machine for reacting to directionally correct reads badly.</p>
<p>The directionally correct read: architectural efficiency gains are real and large. DeepSeek's mixture-of-experts design activates a small fraction of total parameters per token. Their multi-head latent attention cuts KV cache size substantially. Their FP8 training pipeline halves memory traffic against BF16. These are engineering wins, they are published, and they are reproducible.</p>
<p>If capability-per-FLOP is improving that fast, then the number of FLOPs you need to buy to reach a given capability is falling. That is the thesis that repriced today.</p>
<h2 id="the-counter-thesis">the counter-thesis<a class="anchor" href="#the-counter-thesis" aria-label="link to this section">#</a></h2>
<p>Jevons. If compute gets cheaper per unit of capability, you do not buy less of it — you find more things to do with it. Reasoning models in particular consume enormous inference compute; a model that thinks for thirty seconds before answering is a very different demand curve than one that answers immediately. Efficiency gains in training get spent on inference.</p>
<p>Both theses are defensible. The honest answer is that nobody knows the shape of the demand curve, and a 17% single-day move in the largest company on earth is not a considered judgement about that. It is a positioning unwind.</p>
<h2 id="for-engineers-specifically">for engineers specifically<a class="anchor" href="#for-engineers-specifically" aria-label="link to this section">#</a></h2>
<p>The useful lesson has nothing to do with stock prices. It is this: the performance-per-dollar frontier is moving fast enough that any architecture decision you make today assuming current inference costs will be wrong within a year, in your favor.</p>
<p>Do not build elaborate <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a> and routing infrastructure to shave token costs that are going to fall by an order of magnitude anyway. Build the thing. Measure it. Optimize when the bill actually hurts.</p>
<p>That advice would have been wrong in most previous computing eras. It is right now, and it will stop being right at some point, and watching for that moment is most of the job.</p>]]></content:encoded></item><item><title>DeepSeek R1 puts a reasoning model under an MIT license</title><link>https://readme.news/deepseek-r1-puts-a-reasoning-model-under-an-mit-license/</link><guid isPermaLink="true">https://readme.news/deepseek-r1-puts-a-reasoning-model-under-an-mit-license/</guid><pubDate>Tue, 21 Jan 2025 09:00:00 +0000</pubDate><description>Open weights, a published training recipe, and API pricing that reads like a typo. The reasoning-model moat just got a lot shallower.</description><content:encoded><![CDATA[<p>DeepSeek released R1 yesterday: a reasoning model with published weights under an MIT license, a technical report describing how it was trained, and six distilled variants ranging from 1.5B to 70B parameters based on Qwen and Llama backbones.</p>
<p>This is the most consequential open model release since Llama 2, and possibly since Llama 1.</p>
<h2 id="what-it-is">what it is<a class="anchor" href="#what-it-is" aria-label="link to this section">#</a></h2>
<p>R1 is a chain-of-thought model in the same family as OpenAI's o1 — it produces a long internal reasoning trace before answering, and it spends more compute at inference time on harder problems. On the standard reasoning benchmarks (competition math, coding, graduate-level science questions) it lands in o1's neighborhood.</p>
<p>The training story is the part worth reading. The report describes <strong>R1-Zero</strong>, trained with reinforcement learning directly on the base model with no supervised fine-tuning stage at all, using rule-based rewards for verifiable domains — did the math answer match, did the code pass the tests. R1-Zero developed reasoning behavior spontaneously, including a documented moment where the model's trace reconsiders its own approach mid-solution.</p>
<p>R1-Zero's output was unreadable — language mixing, poor formatting — so R1 proper adds a cold-start supervised stage and a second RL pass to fix presentation. But the core finding stands: you can get reasoning to emerge from RL against verifiable rewards without a large human-labeled reasoning corpus.</p>
<h2 id="why-the-license-matters-more-than-the-benchmark">why the license matters more than the benchmark<a class="anchor" href="#why-the-license-matters-more-than-the-benchmark" aria-label="link to this section">#</a></h2>
<p>MIT. Not a bespoke <a class="xref" href="/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/" title="Llama 4 arrives, and the leaderboard problem gets a name">community license</a> with a monthly-active-user carve-out. Not "research only." MIT, on the weights and the distilled variants.</p>
<p>That means anybody can fine-tune it, run it on their own hardware, ship it in a product, and never send a token to anyone's API. For regulated industries that have spent two years unable to get legal approval for a hosted <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a>, this is a door opening.</p>
<p>The distilled models are the practical story for most developers. The 32B distill runs on a single high-memory consumer GPU and is genuinely useful. The 7B and 14B variants run on a laptop.</p>
<h2 id="the-pricing">the pricing<a class="anchor" href="#the-pricing" aria-label="link to this section">#</a></h2>
<p>DeepSeek's own API is priced at roughly a small fraction of comparable reasoning-model pricing from US labs. Whether that reflects genuinely lower serving costs, an architecture advantage from their mixture-of-experts design with a small number of active parameters, or a decision to buy market share, the effect on the market is the same.</p>
<h2 id="what-to-actually-watch">what to actually watch<a class="anchor" href="#what-to-actually-watch" aria-label="link to this section">#</a></h2>
<p>Not the benchmark table. Watch these three things:</p>
<ol><li><strong>How fast the recipe gets reproduced.</strong> If RL-on-verifiable-rewards is as generalizable as the paper suggests, expect a wave of reasoning fine-tunes on other base models within weeks.</li><li><strong>Whether the distills hold up outside benchmarks.</strong> Distilled reasoning traces can look right and be right for the wrong reasons.</li><li><strong>The regulatory reaction.</strong> A capable open reasoning model trained outside US export-control jurisdiction is going to generate a policy conversation whether or not that conversation is technically coherent.</li></ol>
<p>Download the weights. Whatever else happens, they exist now and cannot be un-released.</p>]]></content:encoded></item>
</channel>
</rss>
