<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — performance</title>
<link>https://readme.news/tags/performance/</link>
<atom:link href="https://readme.news/tags/performance/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged performance.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>Why your tests are slow</title><link>https://readme.news/why-your-tests-are-slow/</link><guid isPermaLink="true">https://readme.news/why-your-tests-are-slow/</guid><pubDate>Wed, 16 Sep 2026 09:00:00 +0000</pubDate><description>Six causes, ranked by how often they are the real problem, and the fix for each.</description><content:encoded><![CDATA[<p>A test suite that takes twenty minutes is not run locally. A suite that is not run locally is a suite that fails in CI, which means the feedback loop is now measured in pipeline runs.</p>
<p>Here is where the time actually goes, roughly in order of how often each is the dominant cause.</p>
<h2 id="1-everything-talks-to-a-real-database">1. everything talks to a real database<a class="anchor" href="#1-everything-talks-to-a-real-database" aria-label="link to this section">#</a></h2>
<p>The most common cause by a wide margin. Each test sets up a database, inserts fixtures, runs, and tears down.</p>
<p>Fixes, in order of return:</p>
<p><strong>Roll back instead of truncating.</strong> Wrap each test in a transaction and roll it back. Orders of magnitude faster than deleting rows, and it needs no cleanup code.</p>
<p><strong>Share the schema, not the data.</strong> Create the schema once per run, not per test.</p>
<p><strong>Use a tmpfs for the test database.</strong> The data does not need to survive; durable writes are pure cost. On Postgres, a data directory in memory plus <code>fsync=off</code> and <code>synchronous_commit=off</code> is dramatically faster and completely inappropriate for anything but tests.</p>
<p><strong>Move the logic out of the database's reach.</strong> The deepest fix: if business logic is a pure function of its inputs, its tests do not need a database at all. Tests that are hard to write without infrastructure are telling you about your design.</p>
<h2 id="2-the-suite-is-serial">2. the suite is serial<a class="anchor" href="#2-the-suite-is-serial" aria-label="link to this section">#</a></h2>
<p>Most runners parallelise and most projects have not turned it on, usually because it exposed shared state once and someone reverted it.</p>
<p>That shared state is a bug. Tests that pass only in a particular order have an ordering dependency, and ordering dependencies in tests usually mirror a real one in the code.</p>
<p>Turn on parallelism, fix what breaks, and you will find at least one genuine issue.</p>
<h2 id="3-sleeps">3. sleeps<a class="anchor" href="#3-sleeps" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">time.sleep(2)   # wait for the worker to pick it up</code></pre></div>
<p>Every one of these is pure latency, and they are always tuned to the slowest machine anyone has run on. Fifty of them is a hundred seconds of doing nothing.</p>
<p>Replace with polling on the actual condition:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">wait_until(lambda: job.reload().status == "done", timeout=5)</code></pre></div>
<p>Same worst case, typically a hundredth of the elapsed time, and it fails with a useful message instead of a mysterious assertion.</p>
<h2 id="4-the-fixtures-are-enormous">4. the fixtures are enormous<a class="anchor" href="#4-the-fixtures-are-enormous" aria-label="link to this section">#</a></h2>
<p>A test that needs one user builds an organisation, three teams, forty users and a year of history, because it reuses the "standard" fixture.</p>
<p>Build the minimum. Factories with sensible defaults, overridden per test, beat shared fixture files that grow to serve every case.</p>
<h2 id="5-it-is-doing-real-network-io">5. it is doing real network I/O<a class="anchor" href="#5-it-is-doing-real-network-io" aria-label="link to this section">#</a></h2>
<p>Tests hitting real HTTP endpoints — even internal ones, even mock servers over a socket — pay connection setup and scheduling per call.</p>
<p>Intercept at the client layer rather than the network layer. Most languages have a way to stub the HTTP client in-process, and it is both faster and more deterministic than a local server.</p>
<p>The exception is contract tests, which exist to catch exactly what stubs hide. Keep those, run them separately, and do not let them into the fast suite.</p>
<h2 id="6-everything-is-an-end-to-end-test">6. everything is an end-to-end test<a class="anchor" href="#6-everything-is-an-end-to-end-test" aria-label="link to this section">#</a></h2>
<p>Browser-driving tests are two to three orders of magnitude slower than unit tests. A suite made mostly of them is slow no matter what you do.</p>
<p>The ratio that works: a large number of fast unit tests, a moderate number of integration tests around real boundaries, and a small number of end-to-end tests covering the handful of journeys that must never break.</p>
<p>Inverting that ratio is the most expensive testing mistake a team can make, and it is usually the result of not trusting the lower layers rather than a deliberate choice.</p>
<h2 id="the-measurement-first">the measurement first<a class="anchor" href="#the-measurement-first" aria-label="link to this section">#</a></h2>
<p>Before fixing anything, get the distribution:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">pytest --durations=25
go test ./... -json | jq -r 'select(.Action=="pass") | "\(.Elapsed) \(.Test)"' | sort -rn | head -25</code></pre></div>
<p>Almost always, a small number of tests dominate. Fixing the slowest twenty is usually the whole job, and it is an afternoon rather than a project.</p>
<h2 id="the-target">the target<a class="anchor" href="#the-target" aria-label="link to this section">#</a></h2>
<p><strong>Under 10 seconds for the unit suite</strong>, run on every save. Under two minutes for everything that gates a merge.</p>
<p>The ten-second number is not arbitrary — it is roughly the threshold past which people stop running tests as they work and start batching. Once that happens you have lost the feedback loop, and the suite's speed stops being a convenience question and starts being a quality one.</p>]]></content:encoded></item><item><title>The queues you did not know you had</title><link>https://readme.news/the-queues-you-did-not-know-you-had/</link><guid isPermaLink="true">https://readme.news/the-queues-you-did-not-know-you-had/</guid><pubDate>Tue, 01 Sep 2026 09:00:00 +0000</pubDate><description>Every fixed-size resource is a queue. Most of them are unmonitored, and that is where latency hides.</description><content:encoded><![CDATA[<p>You know about the message queue, because you chose it and it has a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>. The queues that hurt are the ones nobody named.</p>
<p>Anywhere a fixed-size resource is shared by more requests than it has capacity, there is a queue. It has a depth, a wait time, and a failure mode, and almost none of them are instrumented.</p>
<h2 id="the-inventory">the inventory<a class="anchor" href="#the-inventory" aria-label="link to this section">#</a></h2>
<p><strong>The connection pool.</strong> Twenty connections, forty concurrent requests: twenty requests are waiting. Pool wait time is the single most under-measured latency component in web applications, and it is invisible in a database query timer because the clock starts after the connection is acquired.</p>
<p><strong>The thread pool or worker pool.</strong> Same shape. Requests queue for a worker, and the time spent waiting is not attributed to any handler.</p>
<p><strong>The TCP accept backlog.</strong> The kernel holds connections your process has not accepted yet. Overflow silently drops them, and the client sees a timeout with no server-side trace at all.</p>
<p><strong>The HTTP client's per-host connection limit.</strong> Most clients cap concurrent connections per destination. Exceed it and your requests queue in the client, before any network activity, invisible to server-side metrics on both ends.</p>
<p><strong>The DNS resolver.</strong> A limited number of in-flight lookups with a cache that can stampede on expiry.</p>
<p><strong>The disk queue.</strong> Storage devices have a queue depth. Exceed it and I/O waits.</p>
<p><strong>The garbage collector.</strong> Not a queue exactly, but the same behaviour: work that accumulates and is paid in a burst.</p>
<p><strong>Rate limiters.</strong> A limiter that delays rather than rejecting is a queue with an enforced service rate.</p>
<h2 id="why-this-matters-more-than-it-sounds">why this matters more than it sounds<a class="anchor" href="#why-this-matters-more-than-it-sounds" aria-label="link to this section">#</a></h2>
<p>Little's Law: <code>L = λW</code>. Items in the system equals arrival rate times time in system. Rearranged, <strong>wait time is queue length over service rate</strong>.</p>
<p>That means your queue depth choice <em>is</em> a latency choice, whether or not you made it deliberately. A connection pool with a 30-second acquisition timeout is a promise that some requests will wait 30 seconds.</p>
<p>And these queues compose. A request waits for a worker, then waits for a connection, then waits for the disk. Each is modest; the sum is the p99 nobody can explain, because each layer's own metrics look fine.</p>
<h2 id="how-to-find-them">how to find them<a class="anchor" href="#how-to-find-them" aria-label="link to this section">#</a></h2>
<p><strong>Measure acquisition, not just use.</strong> Time the wait for the resource separately from the work done with it:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">t0 = time.perf_counter()
with pool.acquire() as conn:
    t1 = time.perf_counter()
    result = conn.execute(query)
span.set_attribute("db.pool_wait_ms", (t1 - t0) * 1000)
span.set_attribute("db.query_ms", (time.perf_counter() - t1) * 1000)</code></pre></div>
<p>Two numbers instead of one. The first one is the one you did not have, and it is frequently the larger.</p>
<p><strong>Add up your span times.</strong> If the parent span is 400 ms and the children total 180 ms, the missing 220 ms is queueing somewhere. That gap is the most useful signal in a trace and almost nobody looks for it.</p>
<p><strong>Check <code>netstat -s</code> for accept queue overflows.</strong> A non-zero and growing "listen queue overflowed" counter means you are dropping connections before your application sees them.</p>
<p><strong>Load test past capacity deliberately.</strong> Push to 3× and watch which metric degrades first. That is your binding queue, and you will not find it at normal load.</p>
<h2 id="the-rule">the rule<a class="anchor" href="#the-rule" aria-label="link to this section">#</a></h2>
<p>Every queue needs three things, and most have none:</p>
<ol><li><strong>A bounded size.</strong> Unbounded is a memory leak with a friendly name.</li><li><strong>A wait metric</strong> — specifically the <em>age of the oldest waiter</em>, which tells you how far behind you are in time rather than in count.</li><li><strong>A decision about overflow.</strong> Reject, drop, or block. Pick one deliberately, because the default is usually "queue forever," and queueing forever means serving requests whose callers gave up ten seconds ago.</li></ol>
<p>Do that for the queues you chose. Then go find the six you did not.</p>]]></content:encoded></item><item><title>Reading a flame graph</title><link>https://readme.news/reading-a-flame-graph/</link><guid isPermaLink="true">https://readme.news/reading-a-flame-graph/</guid><pubDate>Tue, 11 Aug 2026 09:00:00 +0000</pubDate><description>The single most useful performance visualisation, and the four shapes worth recognising in one.</description><content:encoded><![CDATA[<p>A flame graph answers one question extremely well: <em>where is the time going?</em> Most people who look at one have never been told how to read it, so they squint at a colourful pile of rectangles and conclude that profiling is hard.</p>
<p>It is not. There are four rules and four shapes.</p>
<h2 id="the-four-rules">the four rules<a class="anchor" href="#the-four-rules" aria-label="link to this section">#</a></h2>
<p><strong>The x-axis is not time.</strong> This is the rule everyone gets wrong. Left-to-right is alphabetical or arbitrary — it is <em>not</em> chronological. A frame on the left did not happen before a frame on the right.</p>
<p><strong>Width is total time.</strong> A frame's width is the proportion of samples that included it. Wide means expensive. That is the entire message.</p>
<p><strong>Height is stack depth.</strong> A frame sitting on another means it was called by it. Tall is not bad; tall just means deep call stacks.</p>
<p><strong>Colour is usually meaningless.</strong> In most tools it is random, chosen to make adjacent frames distinguishable. Do not read anything into it unless the tool explicitly says otherwise — some use it for language or module.</p>
<p>That is it. Wide is expensive, stacked means called-by, colour is decoration.</p>
<h2 id="the-four-shapes">the four shapes<a class="anchor" href="#the-four-shapes" aria-label="link to this section">#</a></h2>
<p><strong>A wide plateau near the top.</strong> One function, doing real work, dominating. This is the good case: a clear, single hotspot. Optimise that function or call it less.</p>
<p><strong>A wide plateau near the bottom, narrowing above.</strong> The time is spread across many children. There is no single hotspot; the cost is the whole subtree. Look one level up — usually the answer is "call this subtree fewer times" rather than "make some leaf faster".</p>
<p><strong>A staircase.</strong> Deep, narrow, repeating structure. Often recursion, sometimes a framework's middleware chain. Each frame is cheap; the depth is the cost. Look for a way to shorten the chain rather than to optimise any frame in it.</p>
<p><strong>Many thin spikes with nothing dominant.</strong> Death by a thousand cuts, or your profile is too short. Check the sample count first — a profile of two hundred samples looks like this regardless of the workload. If the sample count is healthy and it still looks like this, the program is genuinely uniform and your wins are architectural rather than local.</p>
<h2 id="the-practical-workflow">the practical workflow<a class="anchor" href="#the-practical-workflow" aria-label="link to this section">#</a></h2>
<p><strong>Profile the thing you care about, under load that resembles production.</strong> A profile of a cold process running a synthetic benchmark measures start-up and your benchmark harness.</p>
<p><strong>Profile long enough.</strong> Seconds, not milliseconds. You are sampling; you need samples.</p>
<p><strong>Look at the widest frame you did not expect.</strong> Not the widest frame — the widest <em>surprising</em> one. Everyone's profile is dominated by something obvious and irreducible. The win is in the frame that has no business being 8% of your runtime.</p>
<p><strong>Check for time you cannot see.</strong> A flame graph of CPU samples shows CPU. If your program spends most of its wall clock waiting on a socket, an on-CPU profile will be nearly empty and entirely misleading. That is what off-CPU profiling and tracing are for, and confusing the two is the most common way people conclude a profiler is lying to them.</p>
<h2 id="differential-flame-graphs">differential flame graphs<a class="anchor" href="#differential-flame-graphs" aria-label="link to this section">#</a></h2>
<p>Underused and excellent. Profile before, profile after, render the difference — frames that got wider are coloured one way, narrower the other.</p>
<p>This turns "did my change help?" from an argument about noisy averages into a picture. It is also the fastest way to find a performance regression introduced between two releases: profile both, diff, look at what turned red.</p>
<h2 id="the-honest-caveat">the honest caveat<a class="anchor" href="#the-honest-caveat" aria-label="link to this section">#</a></h2>
<p>A flame graph tells you where time went. It does not tell you whether that time was necessary, and it will happily show you a perfectly optimised hot loop that should not have been called at all.</p>
<p>The largest performance wins are almost always "stop doing this work" rather than "do this work faster," and no profiler can suggest that. It can only show you where to point the question.</p>]]></content:encoded></item><item><title>The performance budget</title><link>https://readme.news/the-performance-budget/</link><guid isPermaLink="true">https://readme.news/the-performance-budget/</guid><pubDate>Mon, 27 Jul 2026 09:00:00 +0000</pubDate><description>A number, agreed in advance, that fails the build. The only mechanism that has ever kept a system fast.</description><content:encoded><![CDATA[<p>Every system starts fast and gets slower. Not because of one bad decision — because of a hundred small ones, each of which added two milliseconds, none of which anyone could reasonably object to.</p>
<p>The only mechanism I have seen reliably prevent this is a budget: a number, agreed in advance, that fails the build when exceeded.</p>
<h2 id="why-we-should-keep-it-fast-does-not-work">why "we should keep it fast" does not work<a class="anchor" href="#why-we-should-keep-it-fast-does-not-work" aria-label="link to this section">#</a></h2>
<p>Because "fast" has no threshold, so there is never a moment where a specific change is the problem.</p>
<p>Every individual addition is defensible. The library is 12 KB and saves a week. The query is 8 ms and enables a feature. The middleware is 3 ms and improves security.</p>
<p>Ten of those and your page is 400 ms slower, and no single change was wrong. There was no point at which anyone could say no, because the comparison was always "this change versus nothing" rather than "this change versus the budget."</p>
<p>A budget changes the comparison. Now the question is "what are you willing to remove to make room for this," which is a real conversation.</p>
<h2 id="setting-the-numbers">setting the numbers<a class="anchor" href="#setting-the-numbers" aria-label="link to this section">#</a></h2>
<p>Derive them from user-facing outcomes, not from what you currently have.</p>
<p><strong>For a web frontend:</strong></p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">metric</th><th style="text-align:left">budget</th></tr></thead><tbody><tr><td style="text-align:left">JavaScript, compressed</td><td style="text-align:left">170 KB</td></tr><tr><td style="text-align:left">CSS, compressed</td><td style="text-align:left">60 KB</td></tr><tr><td style="text-align:left">Largest Contentful Paint (p75, mobile)</td><td style="text-align:left">2.5 s</td></tr><tr><td style="text-align:left">Interaction to Next Paint (p75)</td><td style="text-align:left">200 ms</td></tr><tr><td style="text-align:left">Total requests, initial load</td><td style="text-align:left">40</td></tr></tbody></table></div>
<p><strong>For an API:</strong></p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">metric</th><th style="text-align:left">budget</th></tr></thead><tbody><tr><td style="text-align:left">p50 latency</td><td style="text-align:left">50 ms</td></tr><tr><td style="text-align:left">p99 latency</td><td style="text-align:left">500 ms</td></tr><tr><td style="text-align:left">database queries per request</td><td style="text-align:left">10</td></tr><tr><td style="text-align:left">memory per instance</td><td style="text-align:left">512 MB</td></tr></tbody></table></div>
<p><strong>Per-layer budgets are the underrated part.</strong> A single end-to-end number tells you something regressed. Per-layer budgets tell you <em>where</em>.</p>
<div class="code"><pre><code>total request budget: 200 ms
  auth:        10 ms
  validation:   5 ms
  database:    80 ms
  business:    40 ms
  serialize:   15 ms
  overhead:    50 ms</code></pre></div>
<p>When the database layer goes to 120 ms, that specific budget fails, and the person who changed the query knows immediately rather than in a month when someone investigates a general slowdown.</p>
<p>This is borrowed directly from game development, where per-subsystem frame budgets have been standard practice for decades.</p>
<h2 id="enforcement">enforcement<a class="anchor" href="#enforcement" aria-label="link to this section">#</a></h2>
<p>The budget must fail something, or it is a wish.</p>
<p><strong>In CI, on every pull request:</strong></p>
<div class="code"><span class="code-lang">yaml</span><pre><code class="lang-yaml">- name: bundle size
  run: npx size-limit          # fails if over the configured budget

- name: performance test
  run: k6 run --threshold 'http_req_duration{p(99)}&lt;500' load.js</code></pre></div>
<p><strong>Report the delta on the pull request.</strong> "This change adds 8 KB to the main bundle (142 KB → 150 KB, budget 170 KB)." Visible, in context, at the moment of decision.</p>
<p><strong>Allow overrides with a written reason.</strong> Not blocking forever — blocking until somebody says why. The record of overrides is itself useful; if you are overriding every week, the budget is wrong or the system is losing.</p>
<h2 id="the-budget-for-query-count">the budget for query count<a class="anchor" href="#the-budget-for-query-count" aria-label="link to this section">#</a></h2>
<p>The single most useful backend budget, and the least common.</p>
<p>Count database queries per request in tests. Assert a maximum.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">with assert_max_queries(10):
    client.get("/api/orders")</code></pre></div>
<p>This catches N+1 queries at the moment they are introduced, which is the only cheap time to catch them. An N+1 that reaches production is found weeks later by someone investigating a slow endpoint, and by then it is embedded in an ORM relationship that four other things depend on.</p>
<p>Almost every framework has a way to do this, and almost nobody does.</p>
<h2 id="what-happens-when-you-exceed-it">what happens when you exceed it<a class="anchor" href="#what-happens-when-you-exceed-it" aria-label="link to this section">#</a></h2>
<p>The conversation that a budget forces, in order:</p>
<ol><li><strong>Can we make the new thing cheaper?</strong> Lazy load it, defer it, make the query better.</li><li><strong>Can we remove something else?</strong> This is the valuable one. It surfaces the feature nobody uses and the library that is doing 5% of what it costs.</li><li><strong>Should the budget change?</strong> Sometimes yes, deliberately, with a reason recorded. A budget that never changes is a budget that will be ignored.</li><li><strong>Do we not ship this?</strong> Rare and it should be available.</li></ol>
<p>Any of those is better than the default, which is that the change lands and the system is permanently slower.</p>
<h2 id="the-thing-to-measure-first">the thing to measure first<a class="anchor" href="#the-thing-to-measure-first" aria-label="link to this section">#</a></h2>
<p>Before setting a budget, get the current numbers and the distribution. p50, p75, p95, p99. On real user hardware and real networks, not on a developer laptop on office wifi.</p>
<p>Then set the budget at roughly where you are, and ratchet it down over time rather than setting an aspirational number you fail immediately.</p>
<p>A budget you exceed on day one gets disabled on day two.</p>]]></content:encoded></item><item><title>Connection pooling, explained properly</title><link>https://readme.news/connection-pooling-explained-properly/</link><guid isPermaLink="true">https://readme.news/connection-pooling-explained-properly/</guid><pubDate>Wed, 15 Jul 2026 09:00:00 +0000</pubDate><description>Why your database has 400 connections, why that is bad, and how to size a pool without guessing.</description><content:encoded><![CDATA[<p>Connection pool sizing is done by copying a number from a blog post, and the number is usually wrong in a specific and expensive way.</p>
<h2 id="why-connections-are-expensive">why connections are expensive<a class="anchor" href="#why-connections-are-expensive" aria-label="link to this section">#</a></h2>
<p>In Postgres specifically, each connection is a separate operating system process with its own memory. A few megabytes of baseline, plus work memory for sorting and hashing, plus its share of shared buffer access.</p>
<p>Four hundred connections means four hundred processes. The scheduler is context switching between them, they are contending for the same locks and buffers, and the memory is largely wasted because most of them are idle.</p>
<p>The counterintuitive result, which has been measured many times: <strong>throughput frequently goes down as connection count goes up, past a fairly low threshold.</strong></p>
<p>More connections does not mean more concurrency. It means more contention.</p>
<h2 id="the-actual-number">the actual number<a class="anchor" href="#the-actual-number" aria-label="link to this section">#</a></h2>
<p>A widely used starting formula:</p>
<div class="code"><pre><code>connections = (core_count × 2) + effective_spindle_count</code></pre></div>
<p>For an 8-core machine with SSD storage, that is somewhere around 16 to 20.</p>
<p>That number seems shockingly low to people running pools of 100 or more. It is correct, and the reasoning is straightforward: a query is either using CPU or waiting on I/O. You need enough connections to keep the cores busy and to have some work queued behind I/O waits. Past that, additional connections are queued at the database instead of queued in your pool, and queueing at the database is worse because it consumes resources.</p>
<p><strong>Test it.</strong> Take your load test, run it at pool sizes of 10, 20, 40, 80, and 160, and plot throughput and p99 latency. The curve rises, flattens, and then degrades. Most people are on the degrading side and have never looked.</p>
<h2 id="the-pooler-layers">the pooler layers<a class="anchor" href="#the-pooler-layers" aria-label="link to this section">#</a></h2>
<p>Three, and they do different things.</p>
<p><strong>Application-side pool.</strong> In-process, reuses connections across requests. Every ORM and database driver has one. This is the minimum.</p>
<p><strong>External pooler</strong> — PgBouncer, pgcat, or a cloud provider's equivalent. Sits between your application and the database and multiplexes many client connections onto few server connections.</p>
<p>This is what you need when you have many application instances. Twenty instances with a pool of 20 each is 400 connections to the database, even if each instance is mostly idle. A pooler collapses that to the number the database actually wants.</p>
<p><strong>The database's own limit.</strong> <code>max_connections</code>. Set it lower than you think — it is a safety valve, and setting it high does not make the database faster, it makes the failure mode worse.</p>
<h2 id="pooling-modes-and-the-one-that-bites">pooling modes, and the one that bites<a class="anchor" href="#pooling-modes-and-the-one-that-bites" aria-label="link to this section">#</a></h2>
<p>External poolers have modes and choosing wrong causes subtle correctness bugs.</p>
<p><strong>Session pooling.</strong> A client gets a server connection for the duration of its session. Safe, and provides little multiplexing benefit.</p>
<p><strong>Transaction pooling.</strong> A server connection is assigned per transaction and returned after commit. This is where the big multiplexing win is, and it is what most people want.</p>
<p><strong>The catch:</strong> anything that depends on session state breaks.</p>
<ul><li>Prepared statements (unless the pooler supports them explicitly, which newer ones do)</li><li><code>SET</code> at the session level</li><li>Session-level advisory locks</li><li><code>LISTEN</code>/<code>NOTIFY</code></li><li>Temporary tables</li><li><code>WITH HOLD</code> cursors</li></ul>
<p>If your ORM uses server-side prepared statements by default — many do — you must either disable them or use a pooler that handles them. This is the single most common transaction-pooling problem and it manifests as intermittent errors under load, which is a miserable thing to debug.</p>
<p><strong>Statement pooling.</strong> A connection per statement. Maximum multiplexing, breaks multi-statement transactions. Almost never what you want.</p>
<h2 id="the-settings-that-matter">the settings that matter<a class="anchor" href="#the-settings-that-matter" aria-label="link to this section">#</a></h2>
<p>Beyond size:</p>
<p><strong>Connection timeout.</strong> How long a request waits for a connection before failing. This should be short — a few seconds. A request waiting thirty seconds for a connection has already been abandoned by its caller.</p>
<p><strong>Idle timeout.</strong> Return connections to the database when not in use. Important with many application instances.</p>
<p><strong>Max lifetime.</strong> Recycle connections periodically. This handles the case where a database failover happens and your pool is holding connections to the old primary — without a max lifetime, some pools hold those indefinitely.</p>
<p><strong>Validation.</strong> Test a connection before handing it out, or handle the failure on first use. Something has to deal with connections that died while idle.</p>
<h2 id="the-diagnostic">the diagnostic<a class="anchor" href="#the-diagnostic" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">sql</span><pre><code class="lang-sql">SELECT state, count(*), max(now() - state_change) AS longest
FROM pg_stat_activity
WHERE backend_type = 'client backend'
GROUP BY state;</code></pre></div>
<p>What you are looking for:</p>
<ul><li><strong>Many <code>idle in transaction</code></strong> — the worst state. A transaction is open and doing nothing, holding locks and preventing vacuum. This is an application bug: a transaction opened and not committed, usually because of an early return or an exception path.</li><li><strong>Many <code>idle</code></strong> — pool is oversized. Harmless but wasteful.</li><li><strong>Many <code>active</code> with long durations</strong> — queries are slow; the pool is not the problem.</li></ul>
<p><code>idle in transaction</code> is the one to alert on. It is always a bug and it causes outages that look like database problems and are application problems.</p>
<h2 id="the-summary">the summary<a class="anchor" href="#the-summary" aria-label="link to this section">#</a></h2>
<p>Size your pool from a load test, not from a blog post. It is smaller than you think. Use an external pooler in transaction mode if you have many application instances, and check your prepared statement behavior when you do.</p>]]></content:encoded></item><item><title>Why your container image is 1.4 gigabytes</title><link>https://readme.news/why-your-container-image-is-14-gigabytes/</link><guid isPermaLink="true">https://readme.news/why-your-container-image-is-14-gigabytes/</guid><pubDate>Fri, 10 Jul 2026 09:00:00 +0000</pubDate><description>It should be forty megabytes. Here is where the rest of it came from and how to get it back.</description><content:encoded><![CDATA[<p>A container image for a compiled service should be tens of megabytes. For an interpreted one, low hundreds. If yours is over a gigabyte, something specific went wrong and it is usually one of six things.</p>
<p>Size matters for real reasons: pull time on cold start, registry cost, deployment speed when you are scaling out under load, and attack surface — every package in the image is something that can have a CVE you have to answer for.</p>
<h2 id="the-six-causes">the six causes<a class="anchor" href="#the-six-causes" aria-label="link to this section">#</a></h2>
<p><strong>1. You shipped the build toolchain.</strong></p>
<p>The compiler, the headers, the package manager cache, the source tree, the test fixtures. All of it needed to build, none of it needed to run.</p>
<p>Multi-stage builds fix this completely:</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile">FROM golang:1.24 AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 go build -ldflags="-s -w" -o /app ./cmd/server

FROM gcr.io/distroless/static-debian12
COPY --from=build /app /app
ENTRYPOINT ["/app"]</code></pre></div>
<p>Final image: the binary, plus CA certificates and timezone data. Tens of megabytes.</p>
<p><strong>2. You started from a full distribution image.</strong></p>
<p><code>FROM ubuntu</code> is roughly 80 MB before you install anything, and it includes a package manager, a shell, and a hundred utilities you will never invoke.</p>
<p>The ladder, from largest to smallest:</p>
<ul><li>Full distribution — 80 MB+</li><li><code>-slim</code> variants — 30–80 MB</li><li>Alpine — 5–10 MB, with musl libc, which will occasionally surprise you</li><li>Distroless — just the runtime, no shell, no package manager</li><li><code>scratch</code> — nothing at all, for static binaries</li></ul>
<p><strong>The Alpine caveat</strong>, since it bites people: musl's allocator and DNS resolver behave differently from glibc's. Python performance in particular can be significantly worse, and some binary wheels do not exist for musl. Test rather than assume.</p>
<p><strong>3. Your layers are ordered wrong.</strong></p>
<p>Each instruction creates a layer. A layer is invalidated when it or anything before it changes.</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile"># bad — any source change reinstalls every dependency
COPY . .
RUN npm ci

# good — dependencies are cached until the lockfile changes
COPY package.json package-lock.json ./
RUN npm ci
COPY . .</code></pre></div>
<p>This does not shrink the final image but it dramatically speeds up builds, which is usually what people actually care about.</p>
<p><strong>4. You deleted things in a later layer.</strong></p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile">RUN apt-get install -y build-essential   # layer 1: +400 MB
RUN apt-get remove -y build-essential    # layer 2: marks deleted, image unchanged</code></pre></div>
<p>Layers are additive. Deleting a file in a later layer hides it and does not remove it. The bytes are still in the image and still transferred on pull.</p>
<p>Everything must happen in one <code>RUN</code>:</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile">RUN apt-get update \
 &amp;&amp; apt-get install -y --no-install-recommends build-essential \
 &amp;&amp; make \
 &amp;&amp; apt-get purge -y build-essential \
 &amp;&amp; apt-get autoremove -y \
 &amp;&amp; rm -rf /var/lib/apt/lists/*</code></pre></div>
<p>Better: use a multi-stage build and do not install the toolchain in the final image at all.</p>
<p><strong>5. You have no <code>.dockerignore</code>.</strong></p>
<p><code>COPY . .</code> copies <code>.git</code>, <code>node_modules</code>, build artifacts, test fixtures, and your local <code>.env</code>.</p>
<div class="code"><pre><code>.git
node_modules
dist
*.log
.env*
**/__pycache__
coverage</code></pre></div>
<p>The <code>.git</code> directory alone is frequently hundreds of megabytes on a mature repository, and it is in a lot of images.</p>
<p><strong>6. Your dependencies are enormous.</strong></p>
<p>Sometimes it is genuinely the dependencies — machine learning stacks with CUDA libraries are legitimately multiple gigabytes.</p>
<p>Check whether you need the GPU variant. <code>torch</code> with CUDA is roughly 2.5 GB; the CPU build is a fraction of that. If you are serving on CPU, you are shipping GPU libraries for nothing.</p>
<h2 id="finding-out-where-it-went">finding out where it went<a class="anchor" href="#finding-out-where-it-went" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">docker history --no-trunc &lt;image&gt;       # size per layer</code></pre></div>
<p>Or use a layer inspection tool that shows you which files are in which layer and how much space is wasted. Ten minutes with one of those tells you exactly what to fix.</p>
<h2 id="the-security-dimension">the security dimension<a class="anchor" href="#the-security-dimension" aria-label="link to this section">#</a></h2>
<p>Every package in the image is potential CVE surface, and your scanner will report all of them regardless of whether the code is reachable.</p>
<p>A distroless image has almost nothing to report, which means the reports you do get are signal rather than noise. That is worth more than the size reduction — a vulnerability report with three entries gets read; one with four hundred does not.</p>
<p>The trade-off: no shell means you cannot <code>docker exec</code> in to debug. Use ephemeral debug containers that attach to the running pod's namespaces instead, which is a better practice anyway because it means your production image is not a debugging toolkit.</p>
<h2 id="the-target">the target<a class="anchor" href="#the-target" aria-label="link to this section">#</a></h2>
<ul><li>Compiled language, static binary: <strong>under 30 MB.</strong></li><li>Interpreted with dependencies: <strong>under 200 MB.</strong></li><li>Anything over a gigabyte without a machine learning stack: something is wrong and it is one of the six above.</li></ul>]]></content:encoded></item><item><title>Compression is underrated</title><link>https://readme.news/compression-is-underrated/</link><guid isPermaLink="true">https://readme.news/compression-is-underrated/</guid><pubDate>Mon, 18 May 2026 09:00:00 +0000</pubDate><description>The cheapest performance win available, ignored because it is not glamorous. Where it pays and which algorithm to pick.</description><content:encoded><![CDATA[<p>Compression trades CPU for bytes. On modern hardware, where CPU is abundant and bandwidth is the constraint at nearly every layer, that trade is favorable far more often than people apply it.</p>
<h2 id="the-layers-where-it-pays">the layers where it pays<a class="anchor" href="#the-layers-where-it-pays" aria-label="link to this section">#</a></h2>
<p><strong>HTTP responses.</strong> Everyone does this and most do it badly. Check yours:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">curl -sI -H 'Accept-Encoding: br, gzip' https://example.com/api/data | grep -i content-encoding</code></pre></div>
<p>If that returns nothing, you are shipping uncompressed JSON, and JSON compresses extraordinarily well — commonly 80–90% for typical API responses, because it is mostly repeated key names.</p>
<p>Brotli beats gzip by roughly 15–20% on text at comparable CPU cost, and is supported everywhere. Use it for static assets at maximum level (precomputed, so the CPU cost is paid once) and at a moderate level for dynamic responses.</p>
<p><strong>Database storage.</strong> Most modern databases support per-table or per-column compression. On a table of text or JSON, it frequently halves the storage — which also halves the I/O, which means more of the working set fits in memory, which is where the real win is.</p>
<p>The CPU cost of decompression is almost always smaller than the I/O cost you avoided.</p>
<p><strong>Logs and telemetry.</strong> Log shipping is often a meaningful fraction of internal network traffic and vendor cost. Compressed batching typically reduces it by an order of magnitude, and the batching itself reduces request overhead.</p>
<p><strong>Backups and object storage.</strong> Storage is cheap and it is not free, and egress definitely is not. Compression at rest is a direct cost reduction with no downside for cold data.</p>
<p><strong>Container images.</strong> Zstandard-compressed layers pull faster than gzip, which matters for cold starts and for anything that scales by launching new instances.</p>
<p><strong>Inter-service traffic.</strong> gRPC and Protocol Buffers are already compact. If you are sending JSON between services — and most people are — compressing it is a large win that costs one configuration line.</p>
<h2 id="which-algorithm">which algorithm<a class="anchor" href="#which-algorithm" aria-label="link to this section">#</a></h2>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">algorithm</th><th style="text-align:left">use for</th></tr></thead><tbody><tr><td style="text-align:left"><strong>zstd</strong></td><td style="text-align:left">almost everything. Wide speed/ratio range, fast decompression.</td></tr><tr><td style="text-align:left"><strong>brotli</strong></td><td style="text-align:left">HTTP text, especially static assets at max level.</td></tr><tr><td style="text-align:left"><strong>gzip</strong></td><td style="text-align:left">compatibility fallback. Never the best choice, always supported.</td></tr><tr><td style="text-align:left"><strong>lz4</strong></td><td style="text-align:left">when speed dominates entirely. In-memory, hot paths, real-time.</td></tr><tr><td style="text-align:left"><strong>xz / lzma</strong></td><td style="text-align:left">archives you compress once and rarely read. Slow, small.</td></tr></tbody></table></div>
<p>The default answer is <strong>zstd</strong>. It has a level parameter spanning from faster-than-lz4 to nearly-as-small-as-xz, decompression is fast at every level, and it has dictionary support.</p>
<h2 id="the-dictionary-trick">the dictionary trick<a class="anchor" href="#the-dictionary-trick" aria-label="link to this section">#</a></h2>
<p>The most underused feature in compression, and the one with the biggest payoff for small messages.</p>
<p>Compression works by finding repetition. A 200-byte JSON message has almost no internal repetition, so compression barely helps — sometimes it makes it bigger.</p>
<p>But across <em>many</em> messages, there is enormous repetition: the same field names, the same enum values, the same URL prefixes.</p>
<p>A shared dictionary trained on representative samples gives the compressor that repetition up front:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">zstd --train samples/*.json -o dict.zst</code></pre></div>
<p>Then compress each message against the dictionary. Small-message ratios that were 1.1× become 3× or better. For any system moving many small similar messages — event streams, queue payloads, cache values — this is a large and nearly free win.</p>
<h2 id="when-not-to-compress">when not to compress<a class="anchor" href="#when-not-to-compress" aria-label="link to this section">#</a></h2>
<p><strong>Already-compressed data.</strong> Images, video, audio, archives. You will spend CPU to make it slightly larger.</p>
<p><strong>Very small payloads without a dictionary.</strong> Under a few hundred bytes, compression overhead can exceed the savings.</p>
<p><strong>When you are CPU-bound and not bandwidth-bound.</strong> Measure. This is rarer than people assume but it does happen.</p>
<p><strong>Encrypted data, compressed after encryption.</strong> Pointless — ciphertext is incompressible. And compressing <em>before</em> encryption can leak information about the plaintext through the ciphertext length, which is the class of attack that includes CRIME and BREACH. If you compress and encrypt, know why it is safe in your context.</p>
<h2 id="the-ten-minute-audit">the ten-minute audit<a class="anchor" href="#the-ten-minute-audit" aria-label="link to this section">#</a></h2>
<ol><li>Check that HTTP responses are compressed, including API responses, not just HTML.</li><li>Check that Brotli is enabled, not just gzip.</li><li>Check your log shipping compresses and batches.</li><li>Check whether your largest database tables support compression and whether it is on.</li><li>If you move many small similar messages, train a dictionary.</li></ol>
<p>That is an afternoon and it routinely produces a larger improvement than a month of application-level optimization, at a fraction of the risk.</p>
<p>The reason it does not happen is that nobody gets promoted for enabling Brotli.</p>]]></content:encoded></item><item><title>Why your build is slow</title><link>https://readme.news/why-your-build-is-slow/</link><guid isPermaLink="true">https://readme.news/why-your-build-is-slow/</guid><pubDate>Wed, 25 Mar 2026 09:00:00 +0000</pubDate><description>Six causes, in order of how often they&#x27;re the actual problem, with the fix for each.</description><content:encoded><![CDATA[<p>Build time is the tax you pay on every single change, and most teams have never measured where it goes. Here are the causes, roughly ordered by how often I find each one to be the actual bottleneck.</p>
<h2 id="1-you-are-not-caching-or-your-cache-never-hits">1. You are not caching, or your cache never hits<a class="anchor" href="#1-you-are-not-caching-or-your-cache-never-hits" aria-label="link to this section">#</a></h2>
<p>By far the most common. Not "we have no cache" — teams have caches. The caches do not hit.</p>
<p>Cache keys that include a timestamp, a branch name, a commit SHA, or anything else that changes every run will produce a 0% hit rate while looking like a working cache. Worse than nothing, because you also pay the upload.</p>
<p><strong>Measure the hit rate first.</strong> Most CI systems report it and almost nobody looks. If it is not above 80% on a typical branch build, fix the key before you do anything else.</p>
<p>The key should be a hash of the inputs that actually determine the output: the lockfile for dependencies, the source files for a compilation unit. Nothing else.</p>
<h2 id="2-you-are-rebuilding-things-that-did-not-change">2. You are rebuilding things that did not change<a class="anchor" href="#2-you-are-rebuilding-things-that-did-not-change" aria-label="link to this section">#</a></h2>
<p>If your build runs everything on every commit, you are paying for the whole repository to verify a one-line change in one module.</p>
<p>The fix is a build graph that knows what depends on what and only rebuilds the affected subtree. This is what Bazel, Buck, Nx, Turborepo, and Gradle's configuration cache exist for.</p>
<p>The cost is real — adopting a graph-aware build system is a project, and for a small repository it is not worth it. The threshold is roughly when your full build exceeds five minutes and most changes touch one small part.</p>
<h2 id="3-your-tests-are-serial-or-parallel-badly">3. Your tests are serial, or parallel badly<a class="anchor" href="#3-your-tests-are-serial-or-parallel-badly" aria-label="link to this section">#</a></h2>
<p>Two separate failures.</p>
<p><strong>Serial:</strong> you have 4,000 tests running one after another on one core. Almost every test framework has a parallel mode and it is frequently off by default because parallel tests expose shared state.</p>
<p>If your tests fail when run in parallel, that is a bug in the tests — shared fixtures, a common database, an order dependency — and fixing it is worth doing regardless of speed.</p>
<p><strong>Parallel badly:</strong> you split tests into eight shards by file count, and one shard takes six minutes while the others take ninety seconds. Your suite takes six minutes.</p>
<p>Split by <em>historical duration</em>, not by count. Most CI systems support this and it frequently halves the wall clock for free.</p>
<h2 id="4-you-are-doing-full-clones">4. You are doing full clones<a class="anchor" href="#4-you-are-doing-full-clones" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">yaml</span><pre><code class="lang-yaml">- uses: actions/checkout@v4
  with:
    fetch-depth: 1</code></pre></div>
<p>A repository with a long history and large binary files can take minutes to clone. Almost no build step needs history. If one does — a <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a> generator, a version-from-tags scheme — give that one job the full clone and shallow-clone everything else.</p>
<h2 id="5-your-runners-are-too-small">5. Your runners are too small<a class="anchor" href="#5-your-runners-are-too-small" aria-label="link to this section">#</a></h2>
<p>The least satisfying answer and frequently the correct one.</p>
<p>Engineer time is expensive. A larger runner that halves your build time pays for itself immediately at any reasonable team size, and the arithmetic is easy to do:</p>
<div class="code"><pre><code>build minutes/day × engineers × wait fraction × hourly cost
    vs.
extra runner cost</code></pre></div>
<p>The extra runner is almost always cheaper. Teams resist this because compute cost is a visible line item and engineer waiting is not.</p>
<h2 id="6-you-are-compiling-in-the-container-build">6. You are compiling in the container build<a class="anchor" href="#6-you-are-compiling-in-the-container-build" aria-label="link to this section">#</a></h2>
<p>If your Dockerfile runs <code>npm install</code> and then a compile, and you are not using BuildKit cache mounts and multi-stage builds properly, you are redoing work on every image build.</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile"># cache mount survives across builds
RUN --mount=type=cache,target=/root/.npm \
    npm ci</code></pre></div>
<p>Layer ordering matters too: copy the lockfile and install before copying source, so a source change does not invalidate the dependency layer. This is well known and consistently gotten wrong.</p>
<h2 id="the-measurement">the measurement<a class="anchor" href="#the-measurement" aria-label="link to this section">#</a></h2>
<p>Before fixing anything, get the breakdown:</p>
<div class="code"><pre><code>queue wait     ← runner capacity or concurrency limits
checkout       ← clone depth, LFS
dependency install ← caching
build          ← incrementality
test           ← parallelism
publish        ← usually fine</code></pre></div>
<p>Most teams find that one stage is 60% of the time and they had assumed it was a different one.</p>
<h2 id="the-target">the target<a class="anchor" href="#the-target" aria-label="link to this section">#</a></h2>
<p>Under ten minutes for the check that gates merge. Under two minutes if you can get there.</p>
<p>The reason for the ten-minute line is behavioral: past it, people stop waiting, start batching, and the whole delivery process degrades in ways that are hard to attribute back to the build.</p>
<p>Not everything has to be in that ten minutes. Split the blocking check from the exhaustive one. Lint, typecheck, and unit tests gate the merge. Integration, e2e, and the full platform matrix run after and page someone on failure.</p>
<p>That single split is usually the largest available improvement and it requires no new tooling.</p>]]></content:encoded></item><item><title>GDC and the engineering nobody outside games learns from</title><link>https://readme.news/gdc-and-the-engineering-nobody-outside-games-learns-from/</link><guid isPermaLink="true">https://readme.news/gdc-and-the-engineering-nobody-outside-games-learns-from/</guid><pubDate>Wed, 04 Mar 2026 09:00:00 +0000</pubDate><description>Frame budgets, determinism, and shipping to a fixed target. Game developers solved problems the rest of us are still arguing about.</description><content:encoded><![CDATA[<p>The Game Developers Conference is running this week, and as usual the technical talks contain more transferable engineering than most software conferences, delivered to an audience that mostly does not overlap with the people who would benefit.</p>
<p>Here is what the rest of us should be stealing.</p>
<h2 id="the-frame-budget">the frame budget<a class="anchor" href="#the-frame-budget" aria-label="link to this section">#</a></h2>
<p>A game running at 60 frames per second has 16.6 milliseconds to do everything: input, simulation, physics, animation, culling, draw call submission, audio, and whatever the game is actually about.</p>
<p>Not on average. Every frame. A frame that takes 20 ms is a visible stutter, and players notice a single dropped frame.</p>
<p>That constraint produces a discipline the rest of software largely lacks:</p>
<p><strong>Budgets are allocated per subsystem, in advance.</strong> Physics gets 3 ms. AI gets 2 ms. If a system exceeds its budget, that is a bug with an owner, not a "performance improvement opportunity for next quarter."</p>
<p>Compare to how most web services handle latency: an aspirational p99 target, no per-component budget, and no mechanism that fails when a component regresses.</p>
<p>Steal this. Give each layer of your request path a latency budget. Assert on it in tests. A change that pushes the database layer from 20 ms to 40 ms should fail CI, not show up on a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> a month later.</p>
<p><strong>Worst case matters more than average.</strong> Nobody in games optimizes mean frame time. They optimize the 99.9th percentile, because that is what the player experiences as jank.</p>
<p>The web equivalent is obvious and widely ignored.</p>
<h2 id="determinism-as-a-feature">determinism as a feature<a class="anchor" href="#determinism-as-a-feature" aria-label="link to this section">#</a></h2>
<p>Multiplayer games with lockstep networking require that the same inputs produce bit-identical outputs on every machine. That is a much stronger constraint than "correct," and achieving it teaches you where nondeterminism actually hides:</p>
<ul><li>Floating point differences across compilers and architectures.</li><li>Iteration order over hash containers.</li><li>Time-based logic.</li><li>Uninitialized memory.</li><li>Anything threaded without a strict ordering.</li></ul>
<p>Every one of those is a source of flaky tests and unreproducible bugs in ordinary software, and most engineers have never had to hunt them systematically.</p>
<p>If your test suite is flaky, the game industry solved that problem decades ago and the answer is: make the system deterministic, then the flakiness is a bug you can find.</p>
<h2 id="data-oriented-design">data-oriented design<a class="anchor" href="#data-oriented-design" aria-label="link to this section">#</a></h2>
<p>The insight: modern CPUs are enormously fast and memory is enormously slow, so performance is dominated by cache behavior rather than instruction count.</p>
<p>Which means the layout of your data matters more than the cleverness of your algorithm, at least until the constant factors stop dominating.</p>
<p>Arrays of structs versus structs of arrays. Contiguous memory over pointer chasing. Processing in batches over per-object virtual dispatch.</p>
<p>This applies directly to any data-processing code and most developers have never been taught to think about it. If you have a hot loop over a collection of objects and it is slower than it should be, the layout is usually why.</p>
<h2 id="shipping-to-a-fixed-target">shipping to a fixed target<a class="anchor" href="#shipping-to-a-fixed-target" aria-label="link to this section">#</a></h2>
<p>Console games ship to hardware that does not change for seven years. You cannot tell the player to buy more RAM. You cannot autoscale.</p>
<p>That produces genuine engineering rather than procurement: measure, budget, optimize, cut scope, ship.</p>
<p>There is a lesson here for cloud-native development, where "we will scale it" has become a substitute for making things efficient. The teams that treat their resource envelope as fixed produce dramatically more efficient systems, and the constraint is a gift.</p>
<h2 id="the-thing-games-do-worse">the thing games do worse<a class="anchor" href="#the-thing-games-do-worse" aria-label="link to this section">#</a></h2>
<p>To be fair: game codebases are frequently a mess, crunch culture is a genuine industry problem, testing practices are often weaker than in other software, and "ship it and patch it" has become normalized in a way that undermines a lot of the above.</p>
<p>Take the engineering discipline. Do not take the working conditions.</p>
<h2 id="the-talks-to-look-for">the talks to look for<a class="anchor" href="#the-talks-to-look-for" aria-label="link to this section">#</a></h2>
<p>The technical postmortems are the good ones — a team explaining what went wrong in a shipped product, in detail, publicly. The industry has a stronger culture of this than almost any other software domain, and the material is excellent regardless of whether you care about games.</p>]]></content:encoded></item><item><title>How to read a benchmark</title><link>https://readme.news/how-to-read-a-benchmark/</link><guid isPermaLink="true">https://readme.news/how-to-read-a-benchmark/</guid><pubDate>Sat, 31 Jan 2026 09:00:00 +0000</pubDate><description>A short guide to the specific ways performance numbers lie, from someone who has published misleading ones.</description><content:encoded><![CDATA[<p>Every vendor publishes benchmarks. Most of them are technically accurate and misleading, and the techniques are consistent enough to enumerate.</p>
<p>I have published misleading benchmarks myself, without intending to. That is the normal case — the incentives select for flattering methodology long before anyone decides to deceive.</p>
<h2 id="the-questions-to-ask-in-order">the questions to ask, in order<a class="anchor" href="#the-questions-to-ask-in-order" aria-label="link to this section">#</a></h2>
<p><strong>1. What was the baseline configured like?</strong></p>
<p>The single most common distortion. The new thing is tuned by experts who built it. The baseline is at defaults, or configured by someone who wanted a specific result.</p>
<p>Look for: was the comparison system's cache warmed? Its connection pool sized? Its JIT given time? Its indexes created? If the writeup does not say, assume the answer is no.</p>
<p><strong>2. What is the workload, and is it yours?</strong></p>
<p>A benchmark showing 10× on sequential scans tells you nothing if your workload is point lookups. A benchmark on 100-byte values tells you nothing about 100 KB values.</p>
<p>The best benchmark is the one you run on your own data. Every published number is a hypothesis about your workload, not a measurement of it.</p>
<p><strong>3. Is it a throughput number or a latency number, and which do you care about?</strong></p>
<p>These trade off. A system can achieve enormous throughput by batching, which destroys latency. A system can achieve excellent latency by not batching, which caps throughput.</p>
<p>Numbers presented without both are presenting the flattering one.</p>
<p><strong>4. What percentile?</strong></p>
<p>An average latency is nearly useless. Users experience the tail. A system with a 5 ms average and a 4-second p99 is a bad system, and its average looks great.</p>
<p>If a benchmark reports only mean latency, that is a signal about the p99.</p>
<p><strong>5. How long did it run?</strong></p>
<p>Short benchmarks miss: garbage collection, <a class="xref" href="/claude-opus-45-and-the-compaction-problem/" title="Claude Opus 4.5 and the compaction problem">compaction</a>, cache eviction, log rotation, connection churn, thermal throttling, and every periodic background process.</p>
<p>A 30-second benchmark measures the best 30 seconds. Anything under ten minutes on a stateful system is measuring the warm-up.</p>
<p><strong>6. How many runs, and what was the variance?</strong></p>
<p>A single run is an anecdote. If there is no error bar, the difference between the two bars may be noise, and you cannot tell.</p>
<p><strong>7. What hardware, and was it the same?</strong></p>
<p>Different instance types, different storage, different network. Cloud instances of the same type vary meaningfully between individual instances — noisy neighbors, different underlying CPU steppings.</p>
<p><strong>8. Who ran it?</strong></p>
<p>Vendor benchmarks favor the vendor. Not usually through dishonesty — through a thousand small decisions about what to measure, made by people who know their own system's strengths.</p>
<p>Independent benchmarks are better and rarer. A benchmark run by the losing vendor's competitor is data; a benchmark where the losing vendor was invited to tune their configuration is much better data.</p>
<h2 id="the-specific-tricks">the specific tricks<a class="anchor" href="#the-specific-tricks" aria-label="link to this section">#</a></h2>
<p><strong>Comparing against a version from two years ago.</strong> Technically accurate, completely useless.</p>
<p><strong>Choosing a metric where you win.</strong> If a system reports "requests per second per core," ask what happens to total requests per second.</p>
<p><strong>The unlabeled log scale.</strong> Makes a 15% difference look like a chasm.</p>
<p><strong>The truncated y-axis.</strong> Same trick, more common.</p>
<p><strong>Different consistency levels.</strong> Comparing an eventually-consistent write to a synchronously-replicated one is not a comparison. This is endemic in database benchmarking.</p>
<p><strong>Excluding the slow path.</strong> "Cached read performance" where the competitor's cache was not warm.</p>
<p><strong>Cost normalization that ignores the cost.</strong> "Performance per dollar" using list price when nobody pays list price.</p>
<h2 id="how-to-publish-an-honest-one">how to publish an honest one<a class="anchor" href="#how-to-publish-an-honest-one" aria-label="link to this section">#</a></h2>
<p>If you are producing benchmarks:</p>
<ul><li>Publish the exact configuration for both systems, and the commands.</li><li>Ask the other project to review your configuration before you publish. If they say your baseline is untuned, fix it.</li><li>Report p50, p95, p99, and max. Report variance across runs.</li><li>Say what workload this represents and what it does not.</li><li>Publish the harness so people can reproduce it.</li><li>State the version numbers and the date.</li></ul>
<p>That is more work and it is the difference between marketing and engineering.</p>
<h2 id="the-practical-rule">the practical rule<a class="anchor" href="#the-practical-rule" aria-label="link to this section">#</a></h2>
<p>For anything you are actually going to depend on: run it yourself, on your data, for at least an hour, with the tail latency recorded.</p>
<p>Published benchmarks tell you which systems are worth testing. They do not tell you which one to use, and treating them as if they do is how you end up migrating twice.</p>]]></content:encoded></item><item><title>Postgres 18's async I/O, four months in</title><link>https://readme.news/postgres-18s-async-io-four-months-in/</link><guid isPermaLink="true">https://readme.news/postgres-18s-async-io-four-months-in/</guid><pubDate>Thu, 15 Jan 2026 09:00:00 +0000</pubDate><description>The biggest storage-layer change in years is in production now. Here&#x27;s what it actually did to real workloads.</description><content:encoded><![CDATA[<p>Postgres 18 shipped with an asynchronous I/O subsystem — the largest change to how Postgres talks to storage in a very long time. Enough people have run it in production now to say something more useful than "it is fast."</p>
<h2 id="what-it-does">what it does<a class="anchor" href="#what-it-does" aria-label="link to this section">#</a></h2>
<p>Historically Postgres read from disk synchronously: a backend process issues a read, blocks, gets the page, continues. On a spinning disk that was fine — the disk was the bottleneck and you could not do better. On modern NVMe, where the device can service many requests in parallel and a single-threaded synchronous reader leaves most of the device idle, it wasted a great deal of available throughput.</p>
<p>Postgres 18 adds an I/O method abstraction:</p>
<div class="code"><pre><code>io_method = sync          # the old behavior
io_method = worker        # dedicated I/O worker processes (default)
io_method = io_uring      # Linux io_uring, if built with support</code></pre></div>
<p>Sequential scans, bitmap heap scans, and vacuum use it. Other paths do not yet; this is an incremental rollout across releases.</p>
<h2 id="what-people-are-seeing">what people are seeing<a class="anchor" href="#what-people-are-seeing" aria-label="link to this section">#</a></h2>
<p>The consistent pattern from the reports I trust:</p>
<p><strong>Large sequential scans on fast NVMe: substantial improvement.</strong> Analytical queries scanning a large table are the clearest win. This is exactly what you would predict — the workload was read-parallel and the code was not.</p>
<p><strong>Vacuum: meaningfully faster.</strong> This matters more than it sounds. Vacuum being slow is the root of a large class of Postgres operational problems, and anything that makes it finish sooner reduces bloat pressure.</p>
<p><strong>OLTP with a good cache hit rate: little change.</strong> If your working set is in shared buffers, you were not doing I/O, so making I/O faster does nothing. Most transactional workloads are in this category and should not expect much.</p>
<p><strong>Cloud block storage: mixed.</strong> Network-attached storage has enough latency and enough of its own queueing that the gains are smaller and less predictable than on local NVMe. Test rather than assume.</p>
<h2 id="the-tuning-notes">the tuning notes<a class="anchor" href="#the-tuning-notes" aria-label="link to this section">#</a></h2>
<p>Two settings matter and the defaults are conservative:</p>
<div class="code"><pre><code>io_combine_limit = 128kB        # how much adjacent I/O to merge into one request
effective_io_concurrency = 16   # how many concurrent reads to issue</code></pre></div>
<p><code>effective_io_concurrency</code> has historically been set based on spindle count, which is meaningless on NVMe. For a modern SSD, values well above the old defaults are appropriate. Measure — the curve flattens and then degrades, and where it turns depends on your device.</p>
<p>If you are on Linux with a recent kernel and can build with io_uring support, it generally outperforms the worker method by a modest margin. The worker method is the safe default and it is genuinely good.</p>
<h2 id="the-other-things-in-18-worth-knowing">the other things in 18 worth knowing<a class="anchor" href="#the-other-things-in-18-worth-knowing" aria-label="link to this section">#</a></h2>
<p><strong>Skip scan for B-tree indexes.</strong> A multi-column index on <code>(a, b)</code> can now be used for a query filtering only on <code>b</code>, when <code>a</code> has low cardinality. This eliminates a category of redundant index that people have been maintaining for years.</p>
<p><strong><code>uuidv7()</code> built in.</strong> Time-ordered UUIDs, natively. If you use UUIDs as primary keys, v7 dramatically improves index locality compared to v4, which scatters inserts across the whole B-tree. This is a real performance issue that a lot of applications have and do not know about.</p>
<p><strong>OAuth authentication support</strong>, which matters for anyone trying to eliminate static database passwords.</p>
<p><strong>Faster major-version upgrades</strong>, with statistics preserved across <code>pg_upgrade</code>. Previously you finished an upgrade with no statistics and a database that planned terribly until you ran <code>ANALYZE</code> across everything. That was a genuine outage risk and it is fixed.</p>
<h2 id="should-you-upgrade">should you upgrade<a class="anchor" href="#should-you-upgrade" aria-label="link to this section">#</a></h2>
<p>Yes, on the normal schedule — after a point release or two, with a tested rollback, having read the release notes for the incompatibilities.</p>
<p>The async I/O is the headline and for most transactional workloads the statistics preservation on upgrade and <code>uuidv7()</code> will matter more day to day.</p>
<p>As always with Postgres: the release is well-tested, the migration is usually boring, and the people who get burned are the ones who skipped four major versions and tried to do it in one jump.</p>]]></content:encoded></item><item><title>Rust 1.90 and the linker nobody talks about</title><link>https://readme.news/rust-190-and-the-linker-nobody-talks-about/</link><guid isPermaLink="true">https://readme.news/rust-190-and-the-linker-nobody-talks-about/</guid><pubDate>Fri, 19 Sep 2025 09:00:00 +0000</pubDate><description>LLD becomes the default linker on x86-64 Linux. Link times drop and nobody notices, which is the point.</description><content:encoded><![CDATA[<p>Rust 1.90 makes LLD the default linker for <code>x86_64-unknown-linux-gnu</code>. This is a build-time performance change with no user-facing API and it is one of the more useful things to ship this year.</p>
<h2 id="why-linking-matters">why linking matters<a class="anchor" href="#why-linking-matters" aria-label="link to this section">#</a></h2>
<p>For a large Rust binary, linking is frequently the dominant cost of an incremental build. You change one line, the compiler recompiles one crate quickly, and then the linker spends several seconds stitching together a hundred megabytes of object files and debug information.</p>
<p>GNU <code>ld</code> is old, single-threaded in the parts that matter, and was designed for a different era of binary sizes. <code>lld</code> is parallel, substantially faster, and has been production-ready for years — it is the default on several other platforms already.</p>
<p>Reported improvements vary by project, with the largest gains on debug builds of large dependency trees. Two to five times faster linking is a common range. If your edit-compile-run loop is four seconds and linking was two of them, you feel this every single time.</p>
<h2 id="how-to-check">how to check<a class="anchor" href="#how-to-check" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">cargo build --timings</code></pre></div>
<p>Opens an HTML report showing where time went, including link time as a distinct phase. Most people have never looked at this and are surprised by what it says.</p>
<p>If you were already opting into <code>lld</code> or <code>mold</code> via <code>.cargo/config.toml</code>, nothing changes for you. If you were not — which is most people, because the configuration was obscure and required knowing it existed — you get the improvement for free on upgrade.</p>
<h2 id="the-general-lesson">the general lesson<a class="anchor" href="#the-general-lesson" aria-label="link to this section">#</a></h2>
<p>Defaults are the most impactful thing a toolchain ships.</p>
<p>A faster linker has been available for years. Anyone could have configured it. The instructions were in a blog post. Almost nobody did, because it required knowing the option existed, knowing it was safe, and caring enough to edit a config file.</p>
<p>Changing the default delivers the improvement to everyone at once. That is worth more than a decade of documentation.</p>
<p>This generalizes. If you maintain a tool and you find yourself writing documentation explaining how to enable a better behavior, ask whether the better behavior should be the default. The answer is usually yes, and the reason it is not is usually caution about breaking a small number of unusual setups — which is a real concern that should be handled with a flag to opt <em>out</em>, not a flag to opt in.</p>
<h2 id="the-rest-of-190">the rest of 1.90<a class="anchor" href="#the-rest-of-190" aria-label="link to this section">#</a></h2>
<ul><li><strong>Cargo workspace publishing</strong> — <code>cargo publish --workspace</code> publishes multiple interdependent crates in the correct order. Anyone maintaining a multi-crate project has written a shell script for this. Now you can delete it.</li><li><strong>Demoted target tiers</strong> for several less-used platforms.</li><li><strong><code>x86_64-apple-darwin</code> demoted to tier 2</strong> with host tools still provided, which reflects where Apple hardware has actually gone.</li><li>The usual const stabilizations and API additions.</li></ul>
<h2 id="the-trend-line">the trend line<a class="anchor" href="#the-trend-line" aria-label="link to this section">#</a></h2>
<p>Rust's reputation for slow compilation is the single most cited objection to adopting it, and it has been steadily addressed for years: incremental compilation, parallel front-end work, the new <a class="xref" href="/rust-184-and-the-quiet-rewrite-underneath/" title="Rust 1.84 and the quiet rewrite underneath">trait solver</a>, better <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a>, and now the linker.</p>
<p>Compile times are meaningfully better than they were three years ago and still slower than Go. That gap is partly fundamental — monomorphization and the amount of optimization Rust does are not free — and partly still addressable.</p>
<p>Anyone who tried Rust in 2021 and bounced off the build times should try again. The numbers are different now.</p>]]></content:encoded></item><item><title>Caching is the only optimization that reliably works</title><link>https://readme.news/caching-is-the-only-optimization-that-reliably-works/</link><guid isPermaLink="true">https://readme.news/caching-is-the-only-optimization-that-reliably-works/</guid><pubDate>Mon, 18 Aug 2025 09:00:00 +0000</pubDate><description>A hierarchy of caches, the two hard problems, and why your p99 is bad in a way profiling won&#x27;t show you.</description><content:encoded><![CDATA[<p>Most performance work is picking up pennies. Rewrite the hot loop, get 15%. Switch the serialization format, get 20%. Move to a faster language, get 3x and spend a quarter doing it.</p>
<p>Caching gets you 1000x, and it does it by not doing the work at all.</p>
<h2 id="the-hierarchy">the hierarchy<a class="anchor" href="#the-hierarchy" aria-label="link to this section">#</a></h2>
<p>Every layer is faster than the one below and every layer has a different invalidation problem. Roughly:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">layer</th><th style="text-align:left">latency</th><th style="text-align:left">invalidation</th></tr></thead><tbody><tr><td style="text-align:left">CPU cache</td><td style="text-align:left">~1 ns</td><td style="text-align:left">the hardware's problem</td></tr><tr><td style="text-align:left">process memory</td><td style="text-align:left">~100 ns</td><td style="text-align:left">yours, and it's per-instance</td></tr><tr><td style="text-align:left">local disk / SSD</td><td style="text-align:left">~100 µs</td><td style="text-align:left">yours, survives restart</td></tr><tr><td style="text-align:left">network cache (Redis)</td><td style="text-align:left">~1 ms</td><td style="text-align:left">yours, shared, coherent</td></tr><tr><td style="text-align:left">CDN edge</td><td style="text-align:left">~10 ms</td><td style="text-align:left">yours, distributed, hard</td></tr><tr><td style="text-align:left">origin</td><td style="text-align:left">~100 ms+</td><td style="text-align:left">not a cache</td></tr></tbody></table></div>
<p>The wins are largest at the top and the invalidation is hardest at the bottom. That trade-off is the entire discipline.</p>
<h2 id="the-part-everyone-skips">the part everyone skips<a class="anchor" href="#the-part-everyone-skips" aria-label="link to this section">#</a></h2>
<p><strong>Measure the hit rate.</strong> Not once — continuously, as a first-class metric with an alert.</p>
<p>A cache with a 50% hit rate is not half as good as one with 99%. It is dramatically worse than that, because the misses dominate the latency distribution:</p>
<div class="code"><pre><code>99% hit rate:  0.99 × 1ms + 0.01 × 100ms  =  1.99 ms average
50% hit rate:  0.50 × 1ms + 0.50 × 100ms  = 50.5  ms average</code></pre></div>
<p>25x worse. And your p99 is entirely determined by the miss path regardless of the hit rate, which is why profiling under normal load tells you nothing about it.</p>
<p>I have seen caches deployed with no hit rate metric, running at 20%, considered a success because latency improved slightly. Somebody added memory pressure, eviction went up, and the "improvement" evaporated silently.</p>
<h2 id="the-failure-mode-nobody-plans-for">the failure mode nobody plans for<a class="anchor" href="#the-failure-mode-nobody-plans-for" aria-label="link to this section">#</a></h2>
<p><strong>Cache stampede.</strong> A popular key expires. A thousand concurrent requests all miss simultaneously. All thousand hit the database. The database falls over. The cache never repopulates because nothing succeeds. Your outage is now self-sustaining.</p>
<p>Three fixes, use all of them:</p>
<ol><li><strong>Request coalescing.</strong> One request per key populates; the rest wait on it. <code>singleflight</code> in Go, <code>Lock</code> around the miss path elsewhere.</li><li><strong>Jittered TTLs.</strong> Never expire a thousand keys at the same instant. Add random ±10%.</li><li><strong>Serve stale while revalidating.</strong> Return the expired value immediately and refresh in the background. Almost always the right choice; almost nobody does it by default.</li></ol>
<h2 id="invalidation">invalidation<a class="anchor" href="#invalidation" aria-label="link to this section">#</a></h2>
<p>Phil Karlton's line about the two hard problems is quoted constantly and rarely acted on. The practical version:</p>
<p><strong>Prefer TTL to explicit invalidation.</strong> A TTL is a bounded staleness guarantee you can reason about. Explicit invalidation is a distributed systems problem where every write path must know every cache key it affects, forever, including the ones added next year.</p>
<p><strong>If you must invalidate explicitly, use key versioning.</strong> Do not delete — change the key.</p>
<div class="code"><pre><code>user:1234:v7:profile</code></pre></div>
<p>Bump the version on write. Old keys age out naturally. No delete storm, no race between invalidate and repopulate, no partially-invalidated state.</p>
<p><strong>Never cache without a TTL.</strong> "It will be invalidated when it changes" is how you get a value from 2023 in production and no way to find out.</p>
<h2 id="the-one-that-is-not-optional">the one that is not optional<a class="anchor" href="#the-one-that-is-not-optional" aria-label="link to this section">#</a></h2>
<p>Cache the expensive read-mostly thing closest to where it is used. If you do nothing else in a performance push, find the query that runs on every request, returns nearly-identical data, and takes 40 ms. Put it in memory with a 60-second TTL.</p>
<p>That single change is worth more than a month of algorithmic optimization on most web applications, and it takes an afternoon.</p>
<h2 id="when-not-to-cache">when not to cache<a class="anchor" href="#when-not-to-cache" aria-label="link to this section">#</a></h2>
<p>Correctness-critical data with strict consistency requirements. Account balances, inventory at the point of sale, permission checks. The staleness window is a correctness bug, not a performance trade.</p>
<p>Cache the permission <em>lookup</em> if you must, with a short TTL and an explicit invalidation on revoke — and understand that you have accepted a window during which a revoked user still has access. Sometimes that is fine. Decide deliberately, and write down which one you chose.</p>]]></content:encoded></item><item><title>Rust 1.89 and the long tail of const generics</title><link>https://readme.news/rust-189-and-the-long-tail-of-const-generics/</link><guid isPermaLink="true">https://readme.news/rust-189-and-the-long-tail-of-const-generics/</guid><pubDate>Fri, 08 Aug 2025 09:00:00 +0000</pubDate><description>Inferred array lengths in const generic position, plus x86 SIMD stabilizations that matter for anyone doing math.</description><content:encoded><![CDATA[<p>Rust 1.89 is out. The headline is explicitly inferred const generic arguments — you can now write <code>_</code> where a const generic parameter would go and let inference figure it out.</p>
<div class="code"><span class="code-lang">rust</span><pre><code class="lang-rust">fn frobnicate&lt;const N: usize&gt;(arr: [u8; N]) -&gt; [u8; N] { /* ... */ }

let data = [1, 2, 3, 4];
let out: [u8; _] = frobnicate(data);   // N inferred as 4</code></pre></div>
<p>Previously you either wrote the number, which duplicated information the compiler already had, or restructured to avoid needing it. Small change, removes a real papercut, and it composes well with the growing amount of const-generic code in the ecosystem.</p>
<h2 id="the-simd-stabilizations">the SIMD stabilizations<a class="anchor" href="#the-simd-stabilizations" aria-label="link to this section">#</a></h2>
<p>A large batch of x86 intrinsics stabilized, including AVX-512 families. This matters for a specific and growing population: people writing hand-tuned kernels in Rust for <a class="xref" href="/compression-is-underrated/" title="Compression is underrated">compression</a>, encoding, cryptography, and increasingly ML inference.</p>
<p>Before this, using these intrinsics required nightly. Which meant that a production library wanting AVX-512 paths had to either require nightly — a non-starter for most consumers — or drop to C via FFI.</p>
<p>That was a genuine competitive disadvantage against C++ for numerical work and it is now mostly gone.</p>
<p>The pattern for using them safely:</p>
<div class="code"><span class="code-lang">rust</span><pre><code class="lang-rust">if is_x86_feature_detected!("avx512f") {
    unsafe { fast_path_avx512(data) }
} else if is_x86_feature_detected!("avx2") {
    unsafe { fast_path_avx2(data) }
} else {
    scalar_path(data)
}</code></pre></div>
<p>Runtime detection, multiple implementations, scalar fallback. Tedious and correct. <code>std::simd</code> remains unstable for the portable abstraction, which is the thing most people actually want and which continues to take a long time.</p>
<h2 id="also-in-the-release">also in the release<a class="anchor" href="#also-in-the-release" aria-label="link to this section">#</a></h2>
<ul><li><strong><code>unsafe</code> attributes</strong> get more coverage, continuing the 2024 edition direction.</li><li><strong><code>Result::flatten</code></strong> stabilizes. Small and frequently wanted.</li><li><strong>Cargo</strong> now supports <code>--compile-time-deps</code> for building only what is needed for procedural macros and build scripts, which speeds up some tooling workflows meaningfully.</li><li><strong>i686 targets</strong> raise their baseline CPU requirement, which will affect approximately nobody and will affect them a lot.</li></ul>
<h2 id="the-const-generics-arc-for-context">the const generics arc, for context<a class="anchor" href="#the-const-generics-arc-for-context" aria-label="link to this section">#</a></h2>
<p>Const generics were stabilized in a minimal form in Rust 1.51, four years ago. Since then the feature has been growing capability incrementally: first only integers, then inference improvements, then better <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a>, now this.</p>
<p>The full feature — const generic expressions, where you can write <code>[T; N * 2]</code> — is still unstable and has been "coming soon" for a long time. The blocker is that arbitrary const expressions in type position require deciding when two type-level expressions are equal, which is undecidable in general and requires drawing a careful line.</p>
<p>Rust's approach to this has consistently been: ship the decidable subset, leave the hard part open, do not paint yourself into a corner. It is slow and it has not produced an unsound <a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">type system</a>, which is more than several languages with similar features can say.</p>]]></content:encoded></item><item><title>Grok 4 and the benchmark that ate the discourse</title><link>https://readme.news/grok-4-and-the-benchmark-that-ate-the-discourse/</link><guid isPermaLink="true">https://readme.news/grok-4-and-the-benchmark-that-ate-the-discourse/</guid><pubDate>Thu, 10 Jul 2025 09:00:00 +0000</pubDate><description>xAI claims frontier results with heavy test-time compute. The number that matters is the one nobody quotes.</description><content:encoded><![CDATA[<p>xAI released Grok 4 last night with claimed state-of-the-art results across several benchmarks, including a striking number on Humanity's Last Exam.</p>
<p>The launch also included a "Heavy" tier that runs multiple agents in parallel and selects among their answers, and a $300/month subscription for it.</p>
<h2 id="reading-the-numbers-correctly">reading the numbers correctly<a class="anchor" href="#reading-the-numbers-correctly" aria-label="link to this section">#</a></h2>
<p>The headline HLE figure comes from the Heavy configuration with tools enabled. That is a legitimate configuration and it is not comparable to a single-sample number from a competitor, which is how it was presented in most coverage.</p>
<p>There are at least four distinct things being reported as "the score":</p>
<ol><li>Single sample, no tools.</li><li>Single sample, with tools (search, code execution).</li><li>Consensus of N samples, no tools.</li><li>Multi-agent parallel with tools and selection.</li></ol>
<p>Number four can be five or ten times the cost of number one. Comparing across these without stating which is being used is not a small methodological quibble. It is the difference between "our model is better" and "we spent more money at inference time."</p>
<p>To be clear: xAI disclosed their configurations. The disclosure was in the livestream and the fine print. The number that traveled was the big one, with no configuration attached, and that is now just how model launches work.</p>
<h2 id="the-parallel-agents-technique">the parallel-agents technique<a class="anchor" href="#the-parallel-agents-technique" aria-label="link to this section">#</a></h2>
<p>Worth taking seriously on its own merits, independent of the marketing.</p>
<p>Running N independent attempts and selecting the best is a well-established test-time scaling method, and it works better than one long chain for a specific reason: independent samples have independent errors, so selection can filter them, while a single chain compounds its errors with no mechanism to recover.</p>
<p>The hard part is selection. If you can verify — the tests pass, the proof checks, the code compiles — selection is easy and this technique is enormously powerful. If you cannot verify, you need a judge, and the judge has the same <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> as the generator.</p>
<p>This is why the technique works so much better on math and code than on open-ended reasoning. Verifiability is the whole game.</p>
<p><strong>The practical version for your own work:</strong> if you have a verifier, sample multiple times and filter. You do not need a special model tier for this. Three samples through a cheap model with a test-suite check will frequently outperform one sample through an expensive one, for less money.</p>
<h2 id="the-other-thing">the other thing<a class="anchor" href="#the-other-thing" aria-label="link to this section">#</a></h2>
<p>Grok's public-facing behavior in the weeks before this launch included a series of incidents that xAI attributed to a system prompt change. I am not going to recount them; they are well documented and they were bad.</p>
<p>The engineering lesson worth extracting: a system prompt is production configuration. It should be version controlled, reviewed, tested against an adversarial eval suite, and rolled out gradually. Treating it as a text box someone can edit is how you get an incident that makes international news.</p>
<p>If your product has a system prompt in a config file that anyone can change without review, fix that this week.</p>
<h2 id="the-state-of-play">the state of play<a class="anchor" href="#the-state-of-play" aria-label="link to this section">#</a></h2>
<p>Grok 4 is a competitive <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a>. xAI went from founded to frontier in roughly two years, which remains the most remarkable thing about the company and is mostly a story about capital and urgency rather than research insight.</p>
<p>The models are converging. The differentiation is moving to price, latency, tooling, and trust, and xAI's position on that last one is self-inflicted.</p>]]></content:encoded></item><item><title>The Illusion of Thinking, and the argument about what reasoning is</title><link>https://readme.news/the-illusion-of-thinking-and-the-argument-about-what-reasoning-is/</link><guid isPermaLink="true">https://readme.news/the-illusion-of-thinking-and-the-argument-about-what-reasoning-is/</guid><pubDate>Fri, 06 Jun 2025 09:00:00 +0000</pubDate><description>An Apple paper finds reasoning models collapse past a complexity threshold. The rebuttals are as instructive as the paper.</description><content:encoded><![CDATA[<p>Apple researchers published "The Illusion of Thinking," evaluating reasoning models on controllable puzzle environments — Tower of Hanoi, river crossing, blocks world — where difficulty can be scaled precisely.</p>
<p>The headline finding: past a certain complexity, accuracy collapses to zero, and counterintuitively the models <em>reduce</em> their <a class="xref" href="/openai-ships-open-weights-for-the-first-time-since-gpt-2/" title="OpenAI ships open weights for the first time since GPT-2">reasoning effort</a> as problems get harder, despite having budget remaining.</p>
<p>The paper is good, the reaction was overheated in both directions, and the rebuttals are worth reading alongside it.</p>
<h2 id="what-the-paper-found">what the paper found<a class="anchor" href="#what-the-paper-found" aria-label="link to this section">#</a></h2>
<p>Three regimes:</p>
<ol><li><strong>Low complexity</strong> — standard models match or beat reasoning models. The extra thinking is wasted and sometimes harmful, because the model overthinks its way past a correct early answer.</li><li><strong>Medium complexity</strong> — reasoning models win clearly. This is the regime the benchmarks live in.</li><li><strong>High complexity</strong> — both collapse to zero accuracy.</li></ol>
<p>The reduced-effort finding at high complexity is the genuinely interesting one. The models emit fewer reasoning tokens on harder problems, which is exactly backwards, and suggests something like learned giving-up rather than a compute limit.</p>
<h2 id="the-rebuttals">the rebuttals<a class="anchor" href="#the-rebuttals" aria-label="link to this section">#</a></h2>
<p>Several, and they land differently.</p>
<p><strong>The output length objection.</strong> Tower of Hanoi with N disks requires 2^N − 1 moves. At N=15 that is 32,767 moves. If the model must enumerate every move in its output, it hits the token limit before it hits a reasoning limit. Several researchers showed models explicitly stating they would not enumerate all moves due to length — and being scored as failures. That is measuring output capacity, not reasoning.</p>
<p>Ask instead for a <em>program</em> that generates the solution and the models do fine. That is a meaningful distinction: knowing the algorithm versus executing it by hand.</p>
<p><strong>The unsolvable instances objection.</strong> Some river-crossing configurations in the evaluation set have no solution. Models were penalized for failing to solve them, which is not a reasoning failure, it is a benchmark bug.</p>
<p><strong>The framing objection.</strong> "Reasoning models cannot reason" was the headline everywhere. The paper does not claim that. It claims a specific scaling limitation on a specific class of problem.</p>
<h2 id="what-survives">what survives<a class="anchor" href="#what-survives" aria-label="link to this section">#</a></h2>
<p>After the corrections, the real finding is narrower and still important:</p>
<p>Current reasoning models do not reliably execute long deterministic procedures. They can identify the right algorithm and fail to carry it out over many steps. Error rates compound; there is no self-correction mechanism strong enough to catch accumulated drift over hundreds of steps.</p>
<p>That is a genuine limitation with direct practical consequences. If your task requires exact multi-step execution — a data migration, a complex refactor across many files, a financial calculation — the model should be writing code that does it, not doing it token by token.</p>
<h2 id="the-practical-rule">the practical rule<a class="anchor" href="#the-practical-rule" aria-label="link to this section">#</a></h2>
<p><strong>Use the model to produce the procedure. Use a computer to execute it.</strong></p>
<p>This is not a workaround, it is the correct architecture. Deterministic execution is what computers are for. A model that writes a correct script and runs it is strictly better than a model that simulates the script in its head, and it is verifiable, repeatable, and debuggable.</p>
<p>Every production agent architecture I have seen work well converges on this. Every one that tries to do arithmetic in the reasoning trace eventually produces a number that is wrong in a way nobody catches.</p>
<h2 id="the-meta-lesson">the meta-lesson<a class="anchor" href="#the-meta-lesson" aria-label="link to this section">#</a></h2>
<p>A paper from a large company with a competitive interest, on a contested topic, with a provocative title, will be read as a position statement regardless of its contents. The authors probably knew that.</p>
<p>Read the methodology section. The methodology section is where papers are true or false.</p>]]></content:encoded></item><item><title>Llama 4 arrives, and the leaderboard problem gets a name</title><link>https://readme.news/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/</link><guid isPermaLink="true">https://readme.news/llama-4-arrives-and-the-leaderboard-problem-gets-a-name/</guid><pubDate>Mon, 07 Apr 2025 09:00:00 +0000</pubDate><description>Scout and Maverick ship with a 10M-token context claim and an arena entry that wasn&#x27;t the released model.</description><content:encoded><![CDATA[<p>Meta released Llama 4 over the weekend: <strong>Scout</strong> (17B active parameters, 16 experts, a claimed 10-million-token context window) and <strong>Maverick</strong> (17B active, 128 experts), both mixture-of-experts, both released under the Llama community license. A larger <strong>Behemoth</strong> was described as still training.</p>
<p>Within forty-eight hours the release turned into a story about benchmark integrity instead.</p>
<h2 id="what-happened">what happened<a class="anchor" href="#what-happened" aria-label="link to this section">#</a></h2>
<p>Maverick posted a very strong score on LMArena, the human-preference leaderboard. It then emerged that the model evaluated on the arena was an "experimental chat version" tuned for conversationality — not the checkpoint released to the public. LMArena updated its policies and published the disputed comparison. Meta's response was that experimental variants are normal and the arena version was labeled.</p>
<p>Both of those things can be true and the outcome is still bad, because the number that traveled was attached to a model nobody could download.</p>
<h2 id="why-this-keeps-happening">why this keeps happening<a class="anchor" href="#why-this-keeps-happening" aria-label="link to this section">#</a></h2>
<p>Leaderboards are the only shared vocabulary the field has, and they are being asked to carry weight they cannot bear.</p>
<ul><li><strong>Human preference arenas</strong> measure whether people like the answer. That correlates with quality and also with formatting, length, confidence, and sycophancy. A model tuned to be agreeable climbs.</li><li><strong>Static benchmarks</strong> leak into training data. Every popular benchmark is on the internet and every <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> has read the internet. Contamination is not always deliberate and it is essentially always present.</li><li><strong>Vendor-run evaluations</strong> use vendor-chosen settings. Consensus-of-64 for yours, single-sample for theirs. Both numbers are real and the comparison is not.</li></ul>
<p>There is no fix that survives contact with commercial incentives. The only durable answer is that you have to run your own evaluation on your own task.</p>
<h2 id="the-10-million-token-claim">the 10 million token claim<a class="anchor" href="#the-10-million-token-claim" aria-label="link to this section">#</a></h2>
<p>Scout's context window is stated at 10M tokens, achieved through an interleaved attention scheme without positional embeddings in some layers, plus inference-time temperature scaling on attention. It was trained on far shorter sequences and generalizes upward.</p>
<p>Take this as an upper bound on what the architecture accepts, not on what it usefully processes. Independent long-context evaluations found substantial degradation well before that number. That is not unique to Llama 4 — it is true of every long-context claim — but 10M is a big enough number that the gap between accepted and useful is enormous.</p>
<h2 id="what-is-actually-good-here">what is actually good here<a class="anchor" href="#what-is-actually-good-here" aria-label="link to this section">#</a></h2>
<p>The MoE architecture with 17B active parameters is a real efficiency story. Maverick's quality-per-active-parameter is strong, and active parameters are what determine inference cost. A model you can serve at 17B economics with much better quality than a 17B dense model is a genuinely useful thing to have.</p>
<p>Native multimodality — images in the base model rather than bolted on via an adapter — is also the right architecture and is where everyone is heading.</p>
<p>The license is still not open source. "Llama community license" has an MAU threshold and naming requirements. Call it open weights, which it is, and do not call it open source, which it is not.</p>
<h2 id="the-practical-advice">the practical advice<a class="anchor" href="#the-practical-advice" aria-label="link to this section">#</a></h2>
<p>Build an eval harness with fifty examples from your actual domain. Run every candidate model against it. Store the results in your repo next to your tests.</p>
<p>It will take a day and it will make you immune to this entire genre of news.</p>]]></content:encoded></item>
</channel>
</rss>
