<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — infrastructure</title>
<link>https://readme.news/tags/infrastructure/</link>
<atom:link href="https://readme.news/tags/infrastructure/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged infrastructure.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>Least privilege, actually applied</title><link>https://readme.news/least-privilege-actually-applied/</link><guid isPermaLink="true">https://readme.news/least-privilege-actually-applied/</guid><pubDate>Fri, 25 Sep 2026 09:00:00 +0000</pubDate><description>Everyone agrees with the principle. Almost nobody has checked what their service can actually reach.</description><content:encoded><![CDATA[<p>Least privilege is one of those principles nobody argues with and few implement, because implementing it means doing the tedious work of finding out what permissions are actually used.</p>
<p>Here is the tedious work, in a manageable order.</p>
<h2 id="the-audit-in-one-afternoon">the audit, in one afternoon<a class="anchor" href="#the-audit-in-one-afternoon" aria-label="link to this section">#</a></h2>
<p>For each service, answer four questions. Write the answers down.</p>
<p><strong>1. What identity does it run as?</strong> Not "a service account" — which one. A surprising number of services run as something shared, or as a human's credentials that were used for the initial deploy and never replaced.</p>
<p><strong>2. What can that identity do?</strong> Enumerate the actual permissions. In most cloud platforms this is one API call. The answer is usually much broader than anyone expects, because permissions accumulate and nothing removes them.</p>
<p><strong>3. What does it actually use?</strong> This is the gap. Cloud providers log every authorised call — CloudTrail, audit logs, equivalents elsewhere. Ninety days of logs tell you precisely which permissions were exercised.</p>
<p><strong>4. What is the delta?</strong> Everything granted and never used is a candidate for removal, and that set is typically most of the grant.</p>
<p>That fourth number is the finding. In every audit I have seen, a service uses somewhere between 5% and 20% of what it is permitted to do.</p>
<h2 id="the-specific-over-grants-to-look-for">the specific over-grants to look for<a class="anchor" href="#the-specific-over-grants-to-look-for" aria-label="link to this section">#</a></h2>
<p><strong>Wildcards.</strong> <code>s3:*</code> on <code>*</code>. Almost always the result of "it wasn't working so I broadened it until it did," which is a debugging technique that becomes a permanent security posture.</p>
<p><strong>Write access where only read is used.</strong> Extremely common for anything reading configuration or reference data.</p>
<p><strong>Access to everything of a type.</strong> A service that needs one bucket having access to all buckets. Scope to the resource, and where possible to a prefix within it.</p>
<p><strong>Permissions for a feature that was removed.</strong> The code went; the grant stayed.</p>
<p><strong>Human roles used by machines.</strong> A deploy pipeline running as a role designed for an engineer's console access — which usually includes the ability to change permissions, and that is the one that turns a compromise into a takeover.</p>
<p><strong><code>iam:*</code> or its equivalent.</strong> The permission to grant permissions. Anything with this is effectively an administrator regardless of what else it has. Treat it as its own category and audit it separately.</p>
<h2 id="the-direction-to-move">the direction to move<a class="anchor" href="#the-direction-to-move" aria-label="link to this section">#</a></h2>
<p><strong>Short-lived over long-lived.</strong> Workload identity — the container proves what it is and receives a token that expires in minutes — removes the stealable credential entirely. This is the single highest-value change on this list, and it is architectural rather than a matter of tightening a policy.</p>
<p><strong>Scoped over broad.</strong> One bucket and one prefix, not the account.</p>
<p><strong>Read over write.</strong> Split the identity if the same service does both, so the read path cannot write.</p>
<p><strong>Deny at the boundary.</strong> A service in a network that cannot reach the internet cannot exfiltrate, whatever its IAM permissions say. Network policy and identity policy are independent layers and both matter.</p>
<h2 id="the-thing-that-makes-it-stick">the thing that makes it stick<a class="anchor" href="#the-thing-that-makes-it-stick" aria-label="link to this section">#</a></h2>
<p>An audit is a snapshot. Permissions creep back the next time something does not work at 6 p.m. on a Friday.</p>
<p>Two mechanisms hold the line:</p>
<p><strong>Permissions in code, reviewed like code.</strong> If a grant requires a pull request, the broad one gets a comment. If it is a console click, it does not.</p>
<p><strong>A scheduled review.</strong> Quarterly, using the same used-versus-granted delta. Twenty minutes per service, and it catches both the emergency grant nobody reverted and the feature that was deleted.</p>
<h2 id="the-honest-framing-for-a-sceptical-audience">the honest framing for a sceptical audience<a class="anchor" href="#the-honest-framing-for-a-sceptical-audience" aria-label="link to this section">#</a></h2>
<p>Least privilege does not prevent compromise. It bounds what a compromise costs.</p>
<p>That is the argument to make when someone asks why it is worth the effort: the question is not whether a credential will leak — a dependency, a laptop, a log file, a misconfigured bucket, eventually one will. The question is whether the answer to "what could they do with it" is "read one bucket" or "anything."</p>
<p>Those two answers are the difference between an incident report and a breach notification, and the work that separates them is an afternoon per service.</p>]]></content:encoded></item><item><title>Secrets management, practically</title><link>https://readme.news/secrets-management-practically/</link><guid isPermaLink="true">https://readme.news/secrets-management-practically/</guid><pubDate>Wed, 08 Jul 2026 09:00:00 +0000</pubDate><description>Not a survey of vaults. The specific practices that actually reduce risk, in order of what to do first.</description><content:encoded><![CDATA[<p>Most secret management advice is a product comparison. Here is the practice instead, ordered by what to do first.</p>
<h2 id="1-stop-long-lived-credentials-existing">1. stop long-lived credentials existing<a class="anchor" href="#1-stop-long-lived-credentials-existing" aria-label="link to this section">#</a></h2>
<p>The single highest-value change, and it is architectural rather than a tool.</p>
<p>A static credential can be stolen, leaked, committed, logged, or exfiltrated from a developer laptop. A credential that lives for fifteen minutes and is scoped to one operation cannot be usefully stolen.</p>
<p><strong>Workload identity.</strong> Your CI job, your container, your function proves its identity to the cloud provider and receives a short-lived token. No stored secret at all.</p>
<div class="code"><span class="code-lang">yaml</span><pre><code class="lang-yaml"># GitHub Actions with OIDC — no stored cloud credentials
permissions:
  id-token: write
steps:
  - uses: aws-actions/configure-aws-credentials@v4
    with:
      role-to-assume: arn:aws:iam::123456789012:role/deploy
      aws-region: us-east-1</code></pre></div>
<p>That eliminates the most commonly stolen class of credential entirely. If you do one thing from this article, do this one.</p>
<p><strong>Database credentials</strong> can work the same way — IAM authentication, or dynamically generated credentials from a secret manager with a short lease.</p>
<h2 id="2-keep-them-out-of-the-places-they-end-up">2. keep them out of the places they end up<a class="anchor" href="#2-keep-them-out-of-the-places-they-end-up" aria-label="link to this section">#</a></h2>
<p><strong>Not in version control.</strong> Obvious, universally violated. Run a scanner in pre-commit and in CI. Both — pre-commit catches it before it happens, CI catches it when someone skipped the hook.</p>
<p><strong>Once committed, assume compromised.</strong> Rewriting history does not help; the object is in every clone and in every fork. Rotate, then clean up.</p>
<p><strong>Not in the container image.</strong> Layers are inspectable. <code>docker history</code> and a layer extraction tool will find it.</p>
<p><strong>Not in environment variables, ideally.</strong> This is more controversial. Environment variables leak: into crash dumps, into child processes, into <code>/proc</code>, into logs when someone prints the environment for debugging, into error tracking services that capture context.</p>
<p>Files with restrictive permissions, mounted at runtime, are better. Environment variables are convenient and are the pragmatic choice for many systems — just know what you are accepting.</p>
<p><strong>Not in logs.</strong> Redact at the logger, with a deny-list of key names, not at each call site. Somebody will forget at a call site. The logger never forgets.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">REDACT = {"password", "token", "secret", "api_key", "authorization", "cookie"}

def scrub(d):
    return {k: ("***" if k.lower() in REDACT else v) for k, v in d.items()}</code></pre></div>
<p><strong>Not in error tracking.</strong> Most error trackers capture local variables and request headers by default. Configure the redaction, then verify it by triggering a test error and reading what arrived.</p>
<h2 id="3-rotate-and-test-the-rotation">3. rotate, and test the rotation<a class="anchor" href="#3-rotate-and-test-the-rotation" aria-label="link to this section">#</a></h2>
<p>Rotation is only real if it has been executed. A rotation procedure that has never run is a document, not a control.</p>
<p><strong>The property that makes rotation painless: support two valid credentials at once.</strong> Add the new one, deploy, verify, remove the old one. Without overlap, rotation requires downtime, so it does not happen.</p>
<p>Design for this when you create the credential, not when you need to rotate it under incident pressure.</p>
<h2 id="4-scope-narrowly">4. scope narrowly<a class="anchor" href="#4-scope-narrowly" aria-label="link to this section">#</a></h2>
<p>A credential should permit exactly what its holder needs.</p>
<ul><li>Read-only where writes are not needed.</li><li>One bucket, one prefix, not the account.</li><li>One database, one schema, one set of tables.</li><li>Time-bounded where possible.</li></ul>
<p>The test: if this credential leaked, what could an attacker do? If the answer is "anything," the scope is wrong regardless of how well you protect it.</p>
<h2 id="5-know-what-you-have">5. know what you have<a class="anchor" href="#5-know-what-you-have" aria-label="link to this section">#</a></h2>
<p>An inventory of every secret: what it is, where it lives, what it grants, who owns it, when it was last rotated.</p>
<p>Most organizations cannot produce this, which means they cannot answer "what do we rotate" during an incident, which turns a two-hour response into a two-day one.</p>
<h2 id="the-incident-procedure">the incident procedure<a class="anchor" href="#the-incident-procedure" aria-label="link to this section">#</a></h2>
<p>When a secret is exposed, in this order:</p>
<ol><li><strong>Rotate first.</strong> Before investigating, before cleanup, before the postmortem. Every minute the credential is valid is a minute of exposure.</li><li><strong>Then assess what it could reach</strong>, and check logs for use.</li><li><strong>Then clean up</strong> the exposure.</li><li><strong>Then work out how it happened.</strong></li></ol>
<p>The common mistake is investigating first. While you investigate, the credential is live.</p>
<h2 id="the-tooling-note">the tooling note<a class="anchor" href="#the-tooling-note" aria-label="link to this section">#</a></h2>
<p>Every major cloud has a secret manager. They are all adequate. The dedicated tools add dynamic credential generation, fine-grained policy, and audit — genuinely useful at scale and not where to start.</p>
<p><strong>Start with: OIDC workload identity for machine-to-machine, a cloud secret manager for what remains, scanning in CI, redaction in the logger, and a tested rotation procedure.</strong></p>
<p>That covers the overwhelming majority of real-world credential compromise, and it is a week of work rather than a platform migration.</p>]]></content:encoded></item><item><title>The crawler tolls, one year on</title><link>https://readme.news/the-crawler-tolls-one-year-on/</link><guid isPermaLink="true">https://readme.news/the-crawler-tolls-one-year-on/</guid><pubDate>Wed, 01 Jul 2026 09:00:00 +0000</pubDate><description>Default blocking and pay-per-crawl changed who can read the web. An assessment of what actually happened.</description><content:encoded><![CDATA[<p>A year ago today a CDN sitting in front of a large fraction of the web flipped its default: AI crawlers blocked unless explicitly allowed, with a marketplace for charging per crawl.</p>
<p>Enough time has passed to say what actually happened rather than what everyone predicted.</p>
<h2 id="what-changed">what changed<a class="anchor" href="#what-changed" aria-label="link to this section">#</a></h2>
<p><strong>The norm inverted.</strong> Before, crawling was permitted by default and <code>robots.txt</code> was a request. Now, for a large share of the web, crawling is denied by default and access is a negotiation.</p>
<p>That is a genuine structural change to how the web works and it happened through one company's configuration default rather than through any standards process, legislation, or public deliberation.</p>
<p><strong>Licensing deals concentrated.</strong> Large AI companies negotiated bulk access with large publishers. That was always the likely outcome: the parties with lawyers and leverage made arrangements, and the arrangements are private.</p>
<p><strong>Small publishers got very little.</strong> The pay-per-crawl mechanism works, technically. The revenue for a site with modest traffic is negligible — the arithmetic never supported anything else. The publishers who most needed a new economic model got the one that pays least.</p>
<p><strong>Non-commercial crawling got harder.</strong> Academic researchers, archivists, and independent search projects have no licensing budget and no negotiating position. Carve-outs exist and are discretionary, which means the ability to study the web is now something you apply for.</p>
<p>That is the outcome I was most worried about and it is the one that materialized most clearly.</p>
<h2 id="what-did-not-change">what did not change<a class="anchor" href="#what-did-not-change" aria-label="link to this section">#</a></h2>
<p><strong>Training data supply.</strong> The frontier labs have enormous existing corpora, licensed sources, and synthetic generation. The marginal value of newly crawled web text was already declining. Restricting it did not create the leverage publishers hoped for.</p>
<p><strong>Traffic.</strong> Referral traffic from search to publishers continued its decline, driven by AI answers in search results, which is a completely separate mechanism from training crawlers and which blocking crawlers does nothing about.</p>
<p>This is the part that was most misunderstood at the time. The traffic problem and the training problem have different causes and blocking crawlers only addresses one of them — the one with less economic impact.</p>
<h2 id="the-thing-to-actually-take-from-it">the thing to actually take from it<a class="anchor" href="#the-thing-to-actually-take-from-it" aria-label="link to this section">#</a></h2>
<p><strong>Infrastructure defaults are policy.</strong> A configuration change at a chokepoint reshaped access to a large fraction of the web, with no process and no appeal.</p>
<p>That is not a criticism of the specific decision, which was popular and defensible. It is an observation about where power actually sits, and it generalizes: the entities that can change the web's behavior are the ones with concentration at a layer everyone depends on, and there are about five of them.</p>
<p><strong>For your own site</strong>, the decision remains yours and it is worth making deliberately rather than accepting a default:</p>
<ul><li><strong>Documentation sites</strong> frequently want to be in the training data. Being the thing the model knows about is worth more than the pageview you did not get.</li><li><strong>Original reporting and analysis</strong> has a stronger case for restriction.</li><li><strong>Anything you want found</strong> should still permit search crawlers, which are a different category and are frequently blocked by accident when people configure this.</li></ul>
<p>Check what you are actually blocking. A meaningful number of sites blocked their own search indexing in the first months of this and did not notice for weeks.</p>
<h2 id="the-unresolved-thing">the unresolved thing<a class="anchor" href="#the-unresolved-thing" aria-label="link to this section">#</a></h2>
<p>The web's economic model — publish freely, get traffic, monetize traffic — is breaking, and nothing has replaced it.</p>
<p>Crawler tolls are not the replacement; the arithmetic does not work at the scale of the actual web, where most content is made by people with no ability to negotiate anything.</p>
<p>Licensing deals are not the replacement either; they work for a few hundred large publishers and for nobody else.</p>
<p>I do not know what the replacement is. I am increasingly convinced that nobody does, and that the interval between the old model failing and a new one existing is going to be long and is going to be bad for the open web.</p>
<p>That is a genuinely pessimistic conclusion and I have not found a way around it in a year of thinking about it.</p>]]></content:encoded></item><item><title>The World Cup is the largest load test ever run</title><link>https://readme.news/the-world-cup-is-the-largest-load-test-ever-run/</link><guid isPermaLink="true">https://readme.news/the-world-cup-is-the-largest-load-test-ever-run/</guid><pubDate>Wed, 10 Jun 2026 09:00:00 +0000</pubDate><description>A month of synchronized global demand across three countries, sixteen cities, and every streaming platform at once.</description><content:encoded><![CDATA[<p>The tournament starts tomorrow: forty-eight teams, three host countries, sixteen venues, and a match schedule designed so that a very large fraction of the planet is watching the same thing at the same moment, repeatedly, for a month.</p>
<p>From an engineering perspective this is the most demanding recurring event in consumer computing, and almost nothing about how it works gets written up.</p>
<h2 id="the-shape-of-the-demand">the shape of the demand<a class="anchor" href="#the-shape-of-the-demand" aria-label="link to this section">#</a></h2>
<p><strong>Synchronized, not smooth.</strong> Streaming traffic for a scheduled match is a step function. Millions of concurrent sessions establish within a two-minute window around kickoff, and the ones that fail to establish are a product failure with no recovery — the user missed the start.</p>
<p><strong>Multi-peak within a session.</strong> A goal produces a spike in social traffic, in betting platforms, in messaging, in news sites, and in the streams of people switching from another match. These arrive within seconds of each other and are correlated across completely unrelated companies.</p>
<p><strong>Correlated across the industry.</strong> This is the part that is genuinely unusual. Every CDN, every mobile network, every payment processor, and every messaging platform experiences the peak simultaneously. There is no borrowing capacity from a quiet neighbor, because there is no quiet neighbor.</p>
<p><strong>Multi-region with different profiles.</strong> Matches in three countries across several <a class="xref" href="/time-zones-and-why-your-calendar-code-is-wrong/" title="Time zones, and why your calendar code is wrong">time zones</a> means the traffic profile shifts across the tournament in ways that capacity planning must anticipate.</p>
<h2 id="what-breaks-historically">what breaks, historically<a class="anchor" href="#what-breaks-historically" aria-label="link to this section">#</a></h2>
<p><strong>Payment processing at scale.</strong> Betting and merchandise platforms see enormous transaction spikes at specific moments. Payment processors are a shared dependency and their capacity is a hard constraint nobody downstream controls.</p>
<p><strong>Mobile networks in venue areas.</strong> Sixty thousand people in one place, all trying to upload video. This is a well-understood problem with an expensive solution (temporary cell capacity) and it still degrades.</p>
<p><strong>Authentication systems.</strong> Everyone logs in at once. Login is frequently the least scaled part of a streaming stack because it is not on the hot path during normal operation.</p>
<p><strong>The thundering herd on recovery.</strong> Something fails, comes back, and every client reconnects simultaneously — causing a second failure. Retry jitter is the single most important line of code in a system like this and it is routinely absent.</p>
<p><strong>Ad insertion.</strong> Server-side ad insertion personalizes streams per viewer while keeping segments cacheable. It is a hard constraint and it is a disproportionate source of live-stream incidents.</p>
<h2 id="what-the-good-operators-do">what the good operators do<a class="anchor" href="#what-the-good-operators-do" aria-label="link to this section">#</a></h2>
<p><strong>Pre-scale, do not autoscale.</strong> Autoscaling responds to load after it arrives. Instance startup is measured in tens of seconds; the demand arrives in one. For a known event at a known time, capacity is provisioned in advance and the autoscaler is a safety net, not the mechanism.</p>
<p><strong>Load shed by feature, not by user.</strong> When capacity is short, disable the recommendations, the comments, the statistics overlay — keep the video. A degraded experience for everyone beats a perfect experience for 80% and nothing for the rest.</p>
<p><strong>Multi-CDN with active steering.</strong> Not for capacity alone — for the fact that any single CDN will have a bad region on a given day. A client-side or DNS-level steering layer that measures real performance and shifts traffic is the difference between a degraded minute and an outage.</p>
<p><strong>Rehearse.</strong> Group stage matches are the rehearsal for the knockout rounds. The teams that treat early matches as production load tests, with instrumentation and a retrospective after each one, are the ones that survive the final.</p>
<p><strong>A war room with authority.</strong> Not a monitoring <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> — a room with the people who can make decisions, including the decision to turn features off, without an approval chain.</p>
<h2 id="the-generalizable-lesson">the generalizable lesson<a class="anchor" href="#the-generalizable-lesson" aria-label="link to this section">#</a></h2>
<p>The interesting property here is <strong>demand you cannot smooth, shed, or refuse</strong>.</p>
<p>Most systems get to spread load over time, queue it, or degrade gracefully by making users wait. None of those work when the product is a live event — a queue means the user misses the goal.</p>
<p>What is left is: provision for peak, make every layer redundant, degrade by feature rather than by user, and rehearse.</p>
<p>That is expensive, unglamorous, and the only thing that works. It is worth remembering the next time someone proposes autoscaling as the answer to a spike that arrives faster than a machine can boot.</p>
<p>Some capacity you have to already own.</p>]]></content:encoded></item><item><title>Edge computing, honestly</title><link>https://readme.news/edge-computing-honestly/</link><guid isPermaLink="true">https://readme.news/edge-computing-honestly/</guid><pubDate>Fri, 29 May 2026 09:00:00 +0000</pubDate><description>Running code close to users is a real win for a narrow set of workloads and a complication for everything else.</description><content:encoded><![CDATA[<p>Edge computing has been sold as a general architectural improvement: run your code in hundreds of locations near users, everything gets faster.</p>
<p>The physics is real. The applicability is narrower than the marketing, and the reason is data.</p>
<h2 id="the-physics">the physics<a class="anchor" href="#the-physics" aria-label="link to this section">#</a></h2>
<p>Light in fiber travels roughly 200,000 km/s. New York to London and back is about 55 ms of pure propagation, before any processing. Add TLS handshakes, TCP setup, and real-world routing that is not a great circle, and a cross-Atlantic round trip is frequently 100 ms or more.</p>
<p>Running code 20 ms from the user instead of 120 ms is a genuine improvement and it is not achievable any other way.</p>
<h2 id="the-problem">the problem<a class="anchor" href="#the-problem" aria-label="link to this section">#</a></h2>
<p>Your data is not at the edge. It is in a database, in one region, and if your edge function needs it, you have moved the compute closer to the user and left the round trip in place — plus added a hop.</p>
<div class="code"><pre><code>user → edge (5ms) → origin database (120ms) → edge → user</code></pre></div>
<p>That is slower than the user talking to the origin directly, because you added a hop to a path that was always dominated by the database call.</p>
<p>This is the single most common edge computing mistake and it is easy to make, because the architecture diagram looks right.</p>
<h2 id="what-edge-is-genuinely-good-for">what edge is genuinely good for<a class="anchor" href="#what-edge-is-genuinely-good-for" aria-label="link to this section">#</a></h2>
<p><strong>Anything that needs no origin data:</strong></p>
<ul><li><strong>Redirects and rewrites.</strong> URL normalization, locale routing, legacy path mapping.</li><li><strong>Authentication token validation.</strong> A signed JWT can be verified with a public key at the edge, and an invalid request never reaches your origin. This is a real win — you reject bad traffic at the perimeter.</li><li><strong>A/B test assignment.</strong> Deterministic hash of a cookie into a bucket. No state required.</li><li><strong>Header manipulation.</strong> Security headers, CORS, feature policy.</li><li><strong>Bot filtering and <a class="xref" href="/rate-limiting-the-four-algorithms-and-when-each-is-wrong/" title="Rate limiting: the four algorithms and when each is wrong">rate limiting</a>.</strong> Reject at the edge, before the request costs you anything.</li><li><strong>Personalization of cached content.</strong> Fetch the cached page, inject the user's name from a cookie, return. The expensive part stays cached.</li></ul>
<p><strong>Anything where the data is genuinely replicated to the edge:</strong></p>
<p>Several platforms now offer edge-replicated key-value and SQL storage. If your data is small, read-heavy, and tolerant of replication lag — configuration, <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">feature flags</a>, product catalogs, translations — this works well and the latency win is real.</p>
<p>The constraints are real too: writes go to a primary, replication is eventual, and storage per location is limited.</p>
<h2 id="what-edge-is-bad-for">what edge is bad for<a class="anchor" href="#what-edge-is-bad-for" aria-label="link to this section">#</a></h2>
<p><strong>Anything write-heavy.</strong> Writes need coordination. Coordination needs a primary. The primary is in one place.</p>
<p><strong>Anything requiring strong consistency.</strong> By definition, this needs coordination, which needs round trips, which is what you were trying to avoid.</p>
<p><strong>Anything with a large working set.</strong> You cannot replicate a terabyte to three hundred locations.</p>
<p><strong>Anything computationally heavy.</strong> Edge runtimes have tight CPU and memory limits. They are designed for milliseconds of work per request.</p>
<p><strong>Anything that needs a specific runtime.</strong> Most edge platforms run a constrained JavaScript or <a class="xref" href="/the-component-model-and-the-plugin-problem/" title="The component model and the plugin problem">WebAssembly</a> environment. Your Python dependency with a C extension is not going there.</p>
<h2 id="the-architecture-that-works">the architecture that works<a class="anchor" href="#the-architecture-that-works" aria-label="link to this section">#</a></h2>
<p>Layered, with each layer doing what it is good at:</p>
<div class="code"><pre><code>edge      → auth check, rate limit, routing, cached content, header work
regional  → application logic, caching, session state
origin    → the database, the writes, the truth</code></pre></div>
<p>Most requests are answered at the edge from cache. Some go to a regional application tier. Few reach the origin.</p>
<p>That is a CDN with programmability, which is what edge computing actually is, and framing it that way produces much better decisions than framing it as "serverless everywhere."</p>
<h2 id="the-thing-to-measure-first">the thing to measure first<a class="anchor" href="#the-thing-to-measure-first" aria-label="link to this section">#</a></h2>
<p>Before adopting any of this: <strong>where does your latency actually go?</strong></p>
<p>Break down a typical request:</p>
<ul><li>DNS</li><li>TLS handshake</li><li>Network round trip</li><li>Time to first byte at origin</li><li>Origin processing</li><li>Database time within that</li><li>Response transfer</li></ul>
<p>If origin processing is 400 ms and network is 40 ms, moving compute to the edge addresses 40 ms of a 440 ms problem. Fix the 400 first.</p>
<p>This is the most common reason edge adoption disappoints: it was applied to a latency problem that was not a network problem.</p>
<h2 id="the-honest-summary">the honest summary<a class="anchor" href="#the-honest-summary" aria-label="link to this section">#</a></h2>
<p>Edge is a very good CDN with programmability, and that is genuinely valuable — it lets you do real work at the perimeter that used to require an origin request.</p>
<p>It is not a general application platform, and the platforms selling it as one are selling the constraint as a feature.</p>
<p>Use it for the perimeter. Keep your data where it can be consistent.</p>]]></content:encoded></item><item><title>Compression is underrated</title><link>https://readme.news/compression-is-underrated/</link><guid isPermaLink="true">https://readme.news/compression-is-underrated/</guid><pubDate>Mon, 18 May 2026 09:00:00 +0000</pubDate><description>The cheapest performance win available, ignored because it is not glamorous. Where it pays and which algorithm to pick.</description><content:encoded><![CDATA[<p>Compression trades CPU for bytes. On modern hardware, where CPU is abundant and bandwidth is the constraint at nearly every layer, that trade is favorable far more often than people apply it.</p>
<h2 id="the-layers-where-it-pays">the layers where it pays<a class="anchor" href="#the-layers-where-it-pays" aria-label="link to this section">#</a></h2>
<p><strong>HTTP responses.</strong> Everyone does this and most do it badly. Check yours:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">curl -sI -H 'Accept-Encoding: br, gzip' https://example.com/api/data | grep -i content-encoding</code></pre></div>
<p>If that returns nothing, you are shipping uncompressed JSON, and JSON compresses extraordinarily well — commonly 80–90% for typical API responses, because it is mostly repeated key names.</p>
<p>Brotli beats gzip by roughly 15–20% on text at comparable CPU cost, and is supported everywhere. Use it for static assets at maximum level (precomputed, so the CPU cost is paid once) and at a moderate level for dynamic responses.</p>
<p><strong>Database storage.</strong> Most modern databases support per-table or per-column compression. On a table of text or JSON, it frequently halves the storage — which also halves the I/O, which means more of the working set fits in memory, which is where the real win is.</p>
<p>The CPU cost of decompression is almost always smaller than the I/O cost you avoided.</p>
<p><strong>Logs and telemetry.</strong> Log shipping is often a meaningful fraction of internal network traffic and vendor cost. Compressed batching typically reduces it by an order of magnitude, and the batching itself reduces request overhead.</p>
<p><strong>Backups and object storage.</strong> Storage is cheap and it is not free, and egress definitely is not. Compression at rest is a direct cost reduction with no downside for cold data.</p>
<p><strong>Container images.</strong> Zstandard-compressed layers pull faster than gzip, which matters for cold starts and for anything that scales by launching new instances.</p>
<p><strong>Inter-service traffic.</strong> gRPC and Protocol Buffers are already compact. If you are sending JSON between services — and most people are — compressing it is a large win that costs one configuration line.</p>
<h2 id="which-algorithm">which algorithm<a class="anchor" href="#which-algorithm" aria-label="link to this section">#</a></h2>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">algorithm</th><th style="text-align:left">use for</th></tr></thead><tbody><tr><td style="text-align:left"><strong>zstd</strong></td><td style="text-align:left">almost everything. Wide speed/ratio range, fast decompression.</td></tr><tr><td style="text-align:left"><strong>brotli</strong></td><td style="text-align:left">HTTP text, especially static assets at max level.</td></tr><tr><td style="text-align:left"><strong>gzip</strong></td><td style="text-align:left">compatibility fallback. Never the best choice, always supported.</td></tr><tr><td style="text-align:left"><strong>lz4</strong></td><td style="text-align:left">when speed dominates entirely. In-memory, hot paths, real-time.</td></tr><tr><td style="text-align:left"><strong>xz / lzma</strong></td><td style="text-align:left">archives you compress once and rarely read. Slow, small.</td></tr></tbody></table></div>
<p>The default answer is <strong>zstd</strong>. It has a level parameter spanning from faster-than-lz4 to nearly-as-small-as-xz, decompression is fast at every level, and it has dictionary support.</p>
<h2 id="the-dictionary-trick">the dictionary trick<a class="anchor" href="#the-dictionary-trick" aria-label="link to this section">#</a></h2>
<p>The most underused feature in compression, and the one with the biggest payoff for small messages.</p>
<p>Compression works by finding repetition. A 200-byte JSON message has almost no internal repetition, so compression barely helps — sometimes it makes it bigger.</p>
<p>But across <em>many</em> messages, there is enormous repetition: the same field names, the same enum values, the same URL prefixes.</p>
<p>A shared dictionary trained on representative samples gives the compressor that repetition up front:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">zstd --train samples/*.json -o dict.zst</code></pre></div>
<p>Then compress each message against the dictionary. Small-message ratios that were 1.1× become 3× or better. For any system moving many small similar messages — event streams, queue payloads, cache values — this is a large and nearly free win.</p>
<h2 id="when-not-to-compress">when not to compress<a class="anchor" href="#when-not-to-compress" aria-label="link to this section">#</a></h2>
<p><strong>Already-compressed data.</strong> Images, video, audio, archives. You will spend CPU to make it slightly larger.</p>
<p><strong>Very small payloads without a dictionary.</strong> Under a few hundred bytes, compression overhead can exceed the savings.</p>
<p><strong>When you are CPU-bound and not bandwidth-bound.</strong> Measure. This is rarer than people assume but it does happen.</p>
<p><strong>Encrypted data, compressed after encryption.</strong> Pointless — ciphertext is incompressible. And compressing <em>before</em> encryption can leak information about the plaintext through the ciphertext length, which is the class of attack that includes CRIME and BREACH. If you compress and encrypt, know why it is safe in your context.</p>
<h2 id="the-ten-minute-audit">the ten-minute audit<a class="anchor" href="#the-ten-minute-audit" aria-label="link to this section">#</a></h2>
<ol><li>Check that HTTP responses are compressed, including API responses, not just HTML.</li><li>Check that Brotli is enabled, not just gzip.</li><li>Check your log shipping compresses and batches.</li><li>Check whether your largest database tables support compression and whether it is on.</li><li>If you move many small similar messages, train a dictionary.</li></ol>
<p>That is an afternoon and it routinely produces a larger improvement than a month of application-level optimization, at a fraction of the risk.</p>
<p>The reason it does not happen is that nobody gets promoted for enabling Brotli.</p>]]></content:encoded></item><item><title>The queue is the architecture</title><link>https://readme.news/the-queue-is-the-architecture/</link><guid isPermaLink="true">https://readme.news/the-queue-is-the-architecture/</guid><pubDate>Fri, 15 May 2026 09:00:00 +0000</pubDate><description>Most scaling problems are solved by making something asynchronous. Most reliability problems are caused by doing it badly.</description><content:encoded><![CDATA[<p>The single most effective architectural move available to most systems is: take the slow thing out of the request path and put it in a queue.</p>
<p>It is also the move that introduces the most subtle <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a>, and the gap between "we added a queue" and "we added a queue correctly" is large.</p>
<h2 id="what-it-buys">what it buys<a class="anchor" href="#what-it-buys" aria-label="link to this section">#</a></h2>
<p><strong>Latency.</strong> The user gets a response when the work is accepted, not when it is done. A checkout that returns in 80 ms and sends the confirmation email asynchronously is a much better product than one that returns in 900 ms.</p>
<p><strong>Absorbing spikes.</strong> A queue is a buffer. Traffic that would overwhelm a synchronous system accumulates and drains. This is the difference between a slow period and an outage.</p>
<p><strong>Isolation.</strong> If the email provider is down, checkout still works. The messages accumulate and send later.</p>
<p><strong>Retry for free.</strong> A failed message goes back on the queue. A failed synchronous call is a user-visible error.</p>
<h2 id="what-it-costs">what it costs<a class="anchor" href="#what-it-costs" aria-label="link to this section">#</a></h2>
<p><strong>Eventual consistency, everywhere.</strong> The user completed checkout and the confirmation has not arrived. The record exists and the search index does not have it. Every asynchronous boundary introduces a window where the system is inconsistent, and your UI has to be honest about it.</p>
<p><strong>Debugging across the boundary.</strong> A synchronous stack trace tells you the whole story. An asynchronous failure requires correlating a producer, a broker, and a consumer, possibly hours apart.</p>
<p><strong>Ordering.</strong> Most <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> do not guarantee it, or guarantee it only within a partition. If your consumer must process events in order, that is a design constraint that reaches back into how you partition.</p>
<p><strong>Duplicate delivery.</strong> Almost all queues are at-least-once. Your consumer <em>will</em> receive the same message twice. If that is not safe, you have a bug that appears under load, weeks after launch.</p>
<h2 id="the-rules">the rules<a class="anchor" href="#the-rules" aria-label="link to this section">#</a></h2>
<p><strong>1. Consumers must be idempotent. Non-negotiable.</strong></p>
<p>At-least-once delivery means duplicates. The consumer must produce the same result whether it processes a message once or five times.</p>
<p>The usual implementation: a natural <a class="xref" href="/idempotency-is-the-only-distributed-systems-concept-you-need/" title="Idempotency is the only distributed systems concept you need">idempotency</a> key, and a record of processed keys.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def handle(msg):
    key = msg.idempotency_key
    with tx():
        if already_processed(key):
            return
        do_the_work(msg)
        mark_processed(key)</code></pre></div>
<p>The <code>mark_processed</code> must be in the same transaction as the work, or you have moved the race rather than eliminated it.</p>
<p><strong>2. Every queue needs a dead letter queue, and someone must watch it.</strong></p>
<p>A message that fails repeatedly must go somewhere. A DLQ nobody monitors is a place where data goes to be silently lost, which is worse than an error, because errors are visible.</p>
<p>Alert on DLQ depth. Not on it being non-zero — on it growing.</p>
<p><strong>3. Retry with backoff and jitter, and cap the attempts.</strong></p>
<p>Immediate retry on a failing downstream is a denial of service you are performing against yourself. Exponential backoff with jitter, a maximum attempt count, then the DLQ.</p>
<p><strong>4. Monitor queue depth and age, not just throughput.</strong></p>
<p>Throughput looks healthy right up until it does not. The metrics that tell you something is wrong:</p>
<ul><li><strong>Depth</strong> — how many messages are waiting.</li><li><strong>Oldest message age</strong> — the most useful single metric. If it is growing, your consumers cannot keep up, and you know how far behind you are in time rather than in count.</li></ul>
<p><strong>5. Decide what happens when the queue is full.</strong></p>
<p>It will be. Reject the producer, drop messages, or block? Each is right in different cases and the default is usually wrong for you. An unbounded queue is not a solution; it is a memory leak with extra steps.</p>
<p><strong>6. Keep the payload small and the reference stable.</strong></p>
<p>Put an ID in the message, not the whole object. The consumer fetches current state. This avoids stale data in the message and keeps the broker fast.</p>
<p>The exception: if you need the state <em>as it was</em> when the event occurred, put it in the message deliberately, and say so.</p>
<p><strong>7. Version your message schema from day one.</strong></p>
<p>Producers and consumers deploy independently. Old consumers will see new messages. Include a version field. Make additive changes only, or handle both shapes.</p>
<h2 id="the-thing-that-surprises-people">the thing that surprises people<a class="anchor" href="#the-thing-that-surprises-people" aria-label="link to this section">#</a></h2>
<p><strong>Queues do not reduce load. They defer it.</strong></p>
<p>If your consumer processes 100 messages per second and you produce 150, you do not have a working system with a buffer. You have a system that is failing slowly, and the queue depth graph is a countdown.</p>
<p>A queue absorbs <em>bursts</em>. It does not fix a sustained capacity deficit, and the failure mode when you use it that way is a queue that grows for six hours and then an incident where you are simultaneously behind and unable to catch up.</p>
<p>Alert on the age, watch the trend, and size the consumers for the sustained rate.</p>]]></content:encoded></item><item><title>You still do not need Kubernetes</title><link>https://readme.news/you-still-do-not-need-kubernetes/</link><guid isPermaLink="true">https://readme.news/you-still-do-not-need-kubernetes/</guid><pubDate>Mon, 13 Apr 2026 09:00:00 +0000</pubDate><description>It&#x27;s excellent software solving a real problem that most teams do not have. The honest threshold, and what to do below it.</description><content:encoded><![CDATA[<p>Kubernetes is genuinely good software. It solves a real problem well. It has an enormous ecosystem and a large pool of people who know it.</p>
<p>It is also, for a majority of the teams running it, a substantial amount of complexity in exchange for benefits they do not receive, and saying so is still mildly heretical.</p>
<h2 id="the-problem-it-actually-solves">the problem it actually solves<a class="anchor" href="#the-problem-it-actually-solves" aria-label="link to this section">#</a></h2>
<p>Kubernetes was built for: many services, many teams, heterogeneous workloads, on a fleet of machines, where you want bin-packing efficiency and declarative self-healing, and where the platform is operated by people whose job that is.</p>
<p>If you have all of those, it is the right answer and there is no close second.</p>
<h2 id="what-it-costs">what it costs<a class="anchor" href="#what-it-costs" aria-label="link to this section">#</a></h2>
<p><strong>A permanent learning tax.</strong> Pods, deployments, services, ingresses, configmaps, secrets, persistent volume claims, storage classes, service accounts, roles, network policies, resource quotas, and a YAML dialect for each. Every engineer who deploys anything must learn a meaningful fraction of it.</p>
<p><strong>Operational surface.</strong> Control plane upgrades, node upgrades, CNI plugin, CSI driver, ingress controller, cert manager, metrics server, log shipper. Each is a component that can break and that must be upgraded on someone else's schedule.</p>
<p><strong>Debugging distance.</strong> "Why is my service not reachable" has a dozen possible answers across five layers, and diagnosing it requires understanding all of them.</p>
<p><strong>Cost, frequently.</strong> A managed control plane plus nodes sized for the platform's own overhead plus the observability stack it needs is often more than the equivalent capacity on simpler infrastructure.</p>
<p><strong>Resume-driven adoption.</strong> This is real and worth naming. Kubernetes on your CV is worth money. That is a genuine incentive pointed away from the simplest solution that works.</p>
<h2 id="the-honest-threshold">the honest threshold<a class="anchor" href="#the-honest-threshold" aria-label="link to this section">#</a></h2>
<p>You probably want Kubernetes if:</p>
<ul><li>More than roughly fifteen to twenty distinct services, deployed independently.</li><li>More than a handful of teams that need to deploy without coordinating.</li><li>You have someone whose job includes operating the platform, not as a side task.</li><li>Genuinely heterogeneous workloads with different scaling characteristics.</li><li>Multi-tenancy requirements with real isolation needs.</li></ul>
<p>You probably do not if:</p>
<ul><li>Under ten services.</li><li>One or two teams.</li><li>Nobody owns the platform.</li><li>Traffic is predictable.</li><li>You are running one application with a database.</li></ul>
<h2 id="what-to-do-below-the-threshold">what to do below the threshold<a class="anchor" href="#what-to-do-below-the-threshold" aria-label="link to this section">#</a></h2>
<p>The options are better than they were, and all of them are boring:</p>
<p><strong>A platform-as-a-service.</strong> Push code, it runs. This is the correct answer for a very large number of applications and the reason people avoid it is usually aesthetic.</p>
<p><strong>Containers on a managed container service</strong> without the orchestrator — the various "run this container, scale it, load balance it" products every cloud offers. You get containers, autoscaling, and rolling deploys without the platform.</p>
<p><strong>A couple of servers and a process manager.</strong> Systemd units, a reverse proxy with automatic certificates, and a deploy script. This runs an enormous amount of traffic, is trivially debuggable, and every engineer already understands it.</p>
<p><strong>Docker Compose on one machine.</strong> For staging, for internal tools, for anything where a single host is enough. Unfashionable, works.</p>
<h2 id="the-migration-path-argument">the migration-path argument<a class="anchor" href="#the-migration-path-argument" aria-label="link to this section">#</a></h2>
<p>"We will need Kubernetes eventually, so we should start now."</p>
<p>This is the most common justification and it is usually wrong, for two reasons.</p>
<p>First, the complexity cost is paid every day from now until then, and "eventually" frequently never arrives.</p>
<p>Second, containerizing your application is the actual hard part of a future migration, and you can do that without an orchestrator. A containerized application running under a simple process manager can move to Kubernetes later in a couple of weeks.</p>
<p>Build the container. Skip the platform until you have the problem it solves.</p>
<h2 id="the-position-i-will-defend">the position I will defend<a class="anchor" href="#the-position-i-will-defend" aria-label="link to this section">#</a></h2>
<p>The teams I have seen most successfully run Kubernetes are large organizations with dedicated <a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">platform teams</a>, where it is genuinely the right tool.</p>
<p>The teams I have seen most damaged by it are small ones where a single engineer set it up, that engineer left, and the remaining team is operating a system nobody understands and is afraid to touch.</p>
<p>That second failure is common, expensive, and entirely predictable from the staffing at adoption time. If nobody's job is going to be operating the platform, you should not have a platform.</p>]]></content:encoded></item><item><title>Infrastructure as code, ten years of lessons</title><link>https://readme.news/infrastructure-as-code-ten-years-of-lessons/</link><guid isPermaLink="true">https://readme.news/infrastructure-as-code-ten-years-of-lessons/</guid><pubDate>Wed, 08 Apr 2026 09:00:00 +0000</pubDate><description>State files, drift, modules that became frameworks, and the one practice that separates teams that like their IaC from teams that don&#x27;t.</description><content:encoded><![CDATA[<p>Infrastructure as code won the argument. Nobody clicks through a console to provision production anymore, or admits to it.</p>
<p>What did not get settled is how to do it without producing something everyone hates. Here is what a decade of watching this has produced.</p>
<h2 id="the-state-file-is-the-whole-problem">the state file is the whole problem<a class="anchor" href="#the-state-file-is-the-whole-problem" aria-label="link to this section">#</a></h2>
<p>Declarative infrastructure tools work by comparing three things: your configuration, a recorded state, and reality. Every hard problem comes from those three disagreeing.</p>
<p><strong>State drift.</strong> Someone changed something by hand. Now state says one thing and reality says another, and the next apply will either revert their change or fail confusingly.</p>
<p>The fix is process, not tooling: <strong>nobody has write access to production infrastructure except the pipeline.</strong> Break-glass access exists, is audited, and triggers a drift check afterward. Without this, drift is continuous and IaC becomes theater.</p>
<p><strong>State locking.</strong> Two applies at once corrupt state. Every backend supports locking. Verify yours is actually enabled — this is a default that people leave off and discover during an incident.</p>
<p><strong>State as a blast radius.</strong> One giant state file means every apply touches everything and a corrupted state loses everything. Split by lifecycle and blast radius: networking, data stores, and applications should not share state.</p>
<p>Rule of thumb: if a plan takes more than a couple of minutes, your state is too big.</p>
<h2 id="modules-that-became-frameworks">modules that became frameworks<a class="anchor" href="#modules-that-became-frameworks" aria-label="link to this section">#</a></h2>
<p>The pattern is universal. A team writes a module to wrap a resource with sensible defaults. It grows parameters. It grows conditionals. Three years later it has forty inputs, nested conditional logic, and nobody can tell what it creates without running a plan.</p>
<p>At that point the abstraction is costing more than the duplication it prevented.</p>
<p><strong>The guidance that holds:</strong></p>
<ul><li><strong>Wrap only when there is real shared policy</strong> — tagging, naming, security baselines. Not to save typing.</li><li><strong>Modules should be shallow.</strong> A module that takes one resource and adds organizational defaults is good. A module that composes twelve resources with conditional branches is a framework, and frameworks need maintainers.</li><li><strong>Prefer duplication over the wrong abstraction</strong>, more so here than in application code, because infrastructure changes less and the cost of a bad abstraction is higher.</li><li><strong>No conditionals that change what resources exist.</strong> <code>count = var.enabled ? 1 : 0</code> is fine once. Nested versions of it produce configurations nobody can reason about.</li></ul>
<h2 id="the-practice-that-separates-teams">the practice that separates teams<a class="anchor" href="#the-practice-that-separates-teams" aria-label="link to this section">#</a></h2>
<p><strong>Plan output is reviewed as part of the pull request.</strong></p>
<p>Not "run apply and see." The plan — the exact set of creates, changes, and destroys — posted to the pull request, read by a human, before merge.</p>
<p>Teams that do this catch the accidental destroy, the security group opened to the world, the database replacement that would have caused an outage. Teams that do not find out during apply.</p>
<p>This is a one-day CI setup and it is the single highest-value practice in the whole category.</p>
<p>The corollary: <strong>pay attention to every destroy in a plan.</strong> Most infrastructure incidents caused by IaC are a resource being replaced when the author expected it to be updated in place. Some attribute changes force replacement. The plan tells you. Nobody reads it.</p>
<h2 id="the-things-to-keep-out">the things to keep out<a class="anchor" href="#the-things-to-keep-out" aria-label="link to this section">#</a></h2>
<p><strong>Secrets.</strong> Never in the configuration, never in state. State files contain resource attributes in plaintext, including some you would not expect. Use a secret manager and reference it, and encrypt the state backend regardless.</p>
<p><strong>Application deployment.</strong> Infrastructure tools are bad at deploying applications — they are declarative and convergent, and deployment is imperative and sequenced. Use them to provision the platform; use something else to deploy onto it.</p>
<p><strong>Anything with a fast change cadence.</strong> If a value changes weekly, it should be configuration read at runtime, not infrastructure applied through a pipeline.</p>
<h2 id="the-fork-question">the fork question<a class="anchor" href="#the-fork-question" aria-label="link to this section">#</a></h2>
<p>The OpenTofu fork means there is now a genuinely open-source implementation with a foundation behind it, alongside the commercial one. Both work. The configuration language is largely compatible.</p>
<p>The practical guidance: this is a lower-stakes decision than it feels. The <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a> is in your modules and your provider usage, not in the binary. Pick based on your licensing requirements and your appetite for the respective governance models, and know that migrating between them is a smaller project than it sounds.</p>
<h2 id="the-thing-i-would-tell-a-team-starting-today">the thing I would tell a team starting today<a class="anchor" href="#the-thing-i-would-tell-a-team-starting-today" aria-label="link to this section">#</a></h2>
<p>Start smaller than you think. One state file per environment per major system. Plain resources before modules. Plan review in CI from day one. No manual changes, ever, enforced by permissions.</p>
<p>The teams that hate their infrastructure code all made the same mistake: they built an abstraction layer before they understood the domain, and then they were maintaining two things.</p>]]></content:encoded></item><item><title>The state of self-hosting</title><link>https://readme.news/the-state-of-self-hosting/</link><guid isPermaLink="true">https://readme.news/the-state-of-self-hosting/</guid><pubDate>Fri, 20 Mar 2026 09:00:00 +0000</pubDate><description>Running your own infrastructure got dramatically easier while the industry was arguing about the cloud. A practical assessment.</description><content:encoded><![CDATA[<p>The default answer to "where should this run" has been "the cloud" for fifteen years, and for most of that time it was correct.</p>
<p>Several things changed and the answer is now more nuanced than the reflex suggests.</p>
<h2 id="what-changed-in-favor-of-self-hosting">what changed in favor of self-hosting<a class="anchor" href="#what-changed-in-favor-of-self-hosting" aria-label="link to this section">#</a></h2>
<p><strong>Machines got enormous.</strong> A single server you can rent for a few hundred dollars a month has more cores, more memory, and dramatically more I/O than a rack of hardware from 2012. A very large number of applications fit on one machine with room to spare.</p>
<p><strong>The tooling got good.</strong> Configuration management, container runtimes, reverse proxies with automatic certificates, backup tooling. What required a team a decade ago requires a competent person and a weekend.</p>
<p><strong>Cloud egress pricing did not fall.</strong> Compute prices came down. Bandwidth pricing at the major clouds is still a large multiple of what it costs, and for bandwidth-heavy applications it dominates the bill.</p>
<p><strong>Managed service prices are high relative to the alternative.</strong> A managed database costs several times what the equivalent instance costs, for operational convenience that is real and is not always worth the multiple.</p>
<h2 id="what-changed-against-it">what changed against it<a class="anchor" href="#what-changed-against-it" aria-label="link to this section">#</a></h2>
<p><strong>Security expectations rose.</strong> Patching, hardening, monitoring, incident response. Self-hosting means you own all of it, and the threat environment is worse than it was.</p>
<p><strong>Compliance frameworks assume cloud controls.</strong> SOC 2, ISO 27001, and their relatives are achievable self-hosted and the evidence collection is more work.</p>
<p><strong>The talent assumption inverted.</strong> A decade ago every team had someone who knew Linux systems administration. Now a lot of teams do not, and hiring for it is harder than hiring for cloud skills.</p>
<h2 id="the-honest-decision-framework">the honest decision framework<a class="anchor" href="#the-honest-decision-framework" aria-label="link to this section">#</a></h2>
<p><strong>Self-host when:</strong></p>
<ul><li>Your workload is steady rather than spiky. Cloud's core value proposition is elasticity, and you are paying for elasticity you do not use.</li><li>Bandwidth is a large share of your bill.</li><li>You have or can hire operational competence.</li><li>Data locality or sovereignty is a requirement.</li><li>You are at a scale where the cloud premium is a meaningful number — which starts lower than most people assume.</li></ul>
<p><strong>Use the cloud when:</strong></p>
<ul><li>Traffic is spiky or unpredictable.</li><li>You are early and optimizing for speed of iteration over unit economics.</li><li>You need global presence and do not want to operate it.</li><li>Your team's time is better spent on the product, which for most early-stage companies it is.</li><li>Compliance requirements are easier to satisfy with a provider's attestations.</li></ul>
<p><strong>The hybrid that most people should consider:</strong> run the steady baseline on owned or rented hardware, burst to cloud for peaks, keep object storage and CDN with a provider. This captures most of the cost advantage without giving up elasticity where it matters.</p>
<h2 id="the-middle-option-nobody-talks-about">the middle option nobody talks about<a class="anchor" href="#the-middle-option-nobody-talks-about" aria-label="link to this section">#</a></h2>
<p>Between "hyperscaler" and "rack in a colo" there is a large market of dedicated server providers and mid-size clouds: a real machine, in a real datacenter, with network and power handled, for a monthly fee.</p>
<p>You get root, predictable performance without noisy neighbors, and bandwidth allowances that are not priced as a profit center. You do not get managed databases, autoscaling, or a hundred adjacent services.</p>
<p>For a very large number of applications this is the correct answer and it is under-considered because the discourse is binary.</p>
<h2 id="the-operational-minimum">the operational minimum<a class="anchor" href="#the-operational-minimum" aria-label="link to this section">#</a></h2>
<p>If you self-host, these are non-negotiable:</p>
<ul><li><strong>Automated, tested restores.</strong> Not backups — <em>restores</em>. A backup you have never restored is a hypothesis. Test it quarterly, on a schedule, with a timer running.</li><li><strong>Unattended security updates</strong>, at least for the OS.</li><li><strong>Monitoring with alerting that reaches a human.</strong> Disk full is the most common self-hosted outage and it is entirely preventable.</li><li><strong><a class="xref" href="/infrastructure-as-code-ten-years-of-lessons/" title="Infrastructure as code, ten years of lessons">Infrastructure as code</a>.</strong> The machine must be reproducible. If rebuilding it requires someone's memory, you have a single point of failure that is a person.</li><li><strong>A documented runbook</strong> for the failures you expect: disk, certificate expiration, service crash, host failure.</li></ul>
<p>That is a weekend of setup and a few hours a month. If nobody on the team will own those hours, use the cloud — that is a legitimate reason and it is the actual deciding factor more often than cost is.</p>
<h2 id="the-thing-that-changed-my-mind">the thing that changed my mind<a class="anchor" href="#the-thing-that-changed-my-mind" aria-label="link to this section">#</a></h2>
<p>I used to treat "we run our own servers" as a red flag. I now treat "we are on the cloud and have never modeled the alternative" as an equal one.</p>
<p>Both are defaults applied without analysis. The analysis takes an afternoon and the answer is frequently not what the reflex says.</p>]]></content:encoded></item><item><title>HTTP/3 and QUIC, five years in</title><link>https://readme.news/http3-and-quic-five-years-in/</link><guid isPermaLink="true">https://readme.news/http3-and-quic-five-years-in/</guid><pubDate>Fri, 13 Mar 2026 09:00:00 +0000</pubDate><description>It shipped, it works, most of the web uses it, and almost nobody understands what changed. A practical review.</description><content:encoded><![CDATA[<p>HTTP/3 is now the majority protocol for a large fraction of web traffic, supported by every major browser and CDN. Most developers have never thought about it, which is the correct outcome for a transport protocol.</p>
<p>It is still worth understanding what it actually changed, because a few of the consequences affect how you build.</p>
<h2 id="the-problem-it-solved">the problem it solved<a class="anchor" href="#the-problem-it-solved" aria-label="link to this section">#</a></h2>
<p>HTTP/2 introduced multiplexing: many logical streams over one TCP connection. That fixed head-of-line blocking at the HTTP layer.</p>
<p>It did not fix it at the TCP layer. TCP delivers bytes in order. If one packet is lost, everything behind it waits, including data for streams that were completely unaffected. On a lossy connection — mobile, congested wifi — HTTP/2 could be worse than HTTP/1.1 with six connections, because one loss stalled everything instead of one sixth of everything.</p>
<p>QUIC moves the transport to UDP and implements reliability, ordering, and congestion control per stream. A lost packet stalls only the stream it belonged to.</p>
<h2 id="what-else-came-with-it">what else came with it<a class="anchor" href="#what-else-came-with-it" aria-label="link to this section">#</a></h2>
<p><strong>Encryption is mandatory and integrated.</strong> TLS 1.3 is part of the protocol rather than a layer on top. The handshake is one round trip, or zero for a resumed connection.</p>
<p><strong>Connection migration.</strong> A QUIC connection is identified by a connection ID, not by the four-tuple of IP addresses and ports. Change networks — wifi to cellular — and the connection survives. Your download does not restart.</p>
<p>This is the feature users notice without knowing why. Walking out of a building while a video plays used to stall it.</p>
<p><strong>Better loss recovery.</strong> QUIC distinguishes between packet loss and reordering more accurately than TCP, and its acknowledgment format carries more information. Recovery is faster.</p>
<p><strong>Evolvability.</strong> TCP is implemented in kernels and middleboxes and cannot change, because the internet is full of devices that will drop anything unfamiliar. QUIC is in userspace and encrypted, so its internals are invisible to middleboxes and can actually be updated.</p>
<p>That last point is arguably the most important long-term consequence. Transport protocol ossification was a genuine crisis and QUIC is the <a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">escape hatch</a>.</p>
<h2 id="the-practical-consequences-for-you">the practical consequences for you<a class="anchor" href="#the-practical-consequences-for-you" aria-label="link to this section">#</a></h2>
<p><strong>Domain sharding is now actively harmful.</strong> Splitting assets across <code>static1.example.com</code> and <code>static2.example.com</code> was a workaround for HTTP/1.1's connection limit. Under HTTP/2 it was pointless. Under HTTP/3 it is worse than pointless, because each domain requires a separate connection with a separate handshake and separate congestion state.</p>
<p>One origin. If you still have sharding from a 2014 optimization guide, remove it.</p>
<p><strong>Concatenating and spriting are counterproductive.</strong> Same reasoning. Many small files multiplex fine and cache better individually.</p>
<p><strong>Priority matters and is under-configured.</strong> HTTP/3 has an extensible priority scheme. Most servers use defaults. If you have a page where certain resources are critical, priority hints (<code>fetchpriority</code>) are worth setting and are widely supported.</p>
<p><strong>UDP blocking is real but small.</strong> Some corporate networks block UDP/443. Clients fall back to HTTP/2 automatically, so this is a performance issue rather than a correctness one. Do not build anything that requires HTTP/3.</p>
<p><strong>Your observability may not see it.</strong> A lot of network monitoring tooling was built for TCP. Check whether your tools actually understand QUIC or are silently reporting nothing.</p>
<h2 id="the-parts-that-were-harder-than-expected">the parts that were harder than expected<a class="anchor" href="#the-parts-that-were-harder-than-expected" aria-label="link to this section">#</a></h2>
<p><strong>CPU cost.</strong> QUIC's userspace implementation and per-packet encryption use more CPU than kernel TCP. This has improved substantially with offload support and better implementations, and it is a real cost at high volume.</p>
<p><strong>Middlebox hostility.</strong> Some networks throttle or block UDP because it looks like something they should throttle. This is improving as QUIC becomes normal traffic.</p>
<p><strong>Debugging is harder.</strong> You cannot read a QUIC connection with tcpdump the way you could read HTTP/1.1. <code>qlog</code> and browser devtools help. The tooling is younger.</p>
<h2 id="the-assessment">the assessment<a class="anchor" href="#the-assessment" aria-label="link to this section">#</a></h2>
<p>For a user on a good connection, HTTP/3 is roughly a wash. For a user on a bad connection — mobile, congested, high latency, lossy — it is a substantial improvement, and those users are a large fraction of the world.</p>
<p>That is exactly the right kind of improvement: invisible to the people who were already fine, meaningful to the people who were not.</p>
<p>Enable it, remove your HTTP/1.1-era workarounds, and go back to not thinking about the transport layer.</p>]]></content:encoded></item><item><title>GTC 2026 and the rack that draws a megawatt</title><link>https://readme.news/gtc-2026-and-the-rack-that-draws-a-megawatt/</link><guid isPermaLink="true">https://readme.news/gtc-2026-and-the-rack-that-draws-a-megawatt/</guid><pubDate>Mon, 09 Mar 2026 09:00:00 +0000</pubDate><description>New silicon, new interconnect, and an industry where the unit of purchase is a room rather than a card.</description><content:encoded><![CDATA[<p>Nvidia's annual conference happened this week and the shape of the announcements confirms what has been obvious for two years: the product is no longer a chip.</p>
<h2 id="the-unit-of-sale">the unit of sale<a class="anchor" href="#the-unit-of-sale" aria-label="link to this section">#</a></h2>
<p>The thing being sold is a <a class="xref" href="/gtc-2025-a-roadmap-to-2027-and-a-warning-about-power/" title="GTC 2025: a roadmap to 2027 and a warning about power">rack-scale</a> system — dozens of accelerators, a co-designed interconnect, integrated liquid cooling, and a power distribution architecture — that arrives as a unit and is installed by people who specialize in installing it.</p>
<p>That is a fundamentally different business from selling PCIe cards, and it is the moat. Competing on a single accelerator's FLOPs per watt is a fight several companies can have. Competing on an integrated rack with a proprietary high-bandwidth interconnect between every device in it is a fight almost nobody can have, because the interconnect is where a decade of engineering lives.</p>
<h2 id="the-power-numbers">the power numbers<a class="anchor" href="#the-power-numbers" aria-label="link to this section">#</a></h2>
<p>The per-rack power figures continue climbing. The industry has moved from "how many kilowatts per rack" to "how many racks per megawatt," which is a different mental model.</p>
<p>The consequences, which are the actual story:</p>
<p><strong>Existing datacenter shells are mostly unusable.</strong> A facility designed for 10 kW racks cannot host these regardless of floor space. The buildout is new construction, near new substations, on a permitting timeline.</p>
<p><strong>Liquid cooling is not optional and not new.</strong> Direct-to-chip liquid cooling is now the baseline, and the interesting engineering has moved to facility-level heat rejection — where does the heat go, and can it be sold to someone.</p>
<p><strong>Higher-voltage DC distribution</strong> is being adopted to reduce conversion losses and copper mass at these current levels. This is a genuine architectural change to how datacenters distribute power and it is being driven entirely by this workload.</p>
<h2 id="the-software-announcements">the software announcements<a class="anchor" href="#the-software-announcements" aria-label="link to this section">#</a></h2>
<p>The inference serving stack continues to be where the practical developer value is. Disaggregated serving — running prefill and decode on separate hardware pools because they have completely different compute and memory characteristics — is now mature enough to be the default recommendation rather than an optimization.</p>
<p>If you operate inference at any scale and have not looked at disaggregation, that remains the largest single efficiency win available. Prefill is compute-bound and parallelizable. Decode is memory-bandwidth-bound and sequential. Running them on one homogeneous pool wastes a substantial fraction of your silicon on whichever phase is not currently limiting.</p>
<h2 id="the-competitive-picture">the competitive picture<a class="anchor" href="#the-competitive-picture" aria-label="link to this section">#</a></h2>
<p>Every hyperscaler now has its own accelerator in volume production. The software gap versus CUDA remains the deciding factor, and it is narrowing slowly rather than quickly, because CUDA's advantage is fifteen years of accumulated libraries, kernels, tooling, and institutional knowledge rather than any single technical feature.</p>
<p>The realistic near-term outcome is not displacement. It is a market where hyperscalers run their own silicon for their own high-volume internal workloads and buy Nvidia for everything else, which is enough to constrain pricing without threatening the position.</p>
<h2 id="what-a-developer-should-take-from-this">what a developer should take from this<a class="anchor" href="#what-a-developer-should-take-from-this" aria-label="link to this section">#</a></h2>
<p><strong>Almost none of it directly.</strong> You will not buy one of these. You will rent time on one.</p>
<p>What matters to you:</p>
<ul><li><strong>Capacity comes in steps</strong>, tied to construction schedules. Plan around availability, not just price.</li><li><strong>Efficiency work compounds.</strong> Every optimization that reduces tokens, batches better, or caches more is worth more in an environment where the underlying resource is physically constrained.</li><li><strong>The abstraction layer is your friend.</strong> The hardware underneath your API calls will change several times. If your code cares, that is a design problem you created.</li></ul>
<h2 id="the-sentence-that-summarizes-the-era">the sentence that summarizes the era<a class="anchor" href="#the-sentence-that-summarizes-the-era" aria-label="link to this section">#</a></h2>
<p>The bottleneck on artificial intelligence is currently electrical engineering and civil construction, and it has been for about two years.</p>
<p>Every discussion about model capability that ignores this is discussing a hypothetical.</p>]]></content:encoded></item><item><title>Streaming an Olympics: the hardest scaling problem nobody discusses</title><link>https://readme.news/streaming-an-olympics-the-hardest-scaling-problem-nobody-discusses/</link><guid isPermaLink="true">https://readme.news/streaming-an-olympics-the-hardest-scaling-problem-nobody-discusses/</guid><pubDate>Fri, 06 Feb 2026 09:00:00 +0000</pubDate><description>Synchronized global demand, hard latency requirements, and no option to shed load. A look at how it actually works.</description><content:encoded><![CDATA[<p>The Winter Olympics start today, which makes this a good moment to talk about a category of engineering problem that gets almost no coverage despite being one of the hardest in production computing.</p>
<p>Live streaming a global sporting event has a combination of constraints that almost nothing else has.</p>
<h2 id="why-it-is-hard">why it is hard<a class="anchor" href="#why-it-is-hard" aria-label="link to this section">#</a></h2>
<p><strong>The demand is synchronized.</strong> Web traffic is normally a smooth curve. A live event is a step function — millions of people join within the same sixty seconds, at a moment you know in advance and cannot move.</p>
<p><strong>You cannot shed load.</strong> The standard answer to overload is to degrade: serve cached content, drop non-essential features, queue people. None of those work when the product is a live video feed. A queue means the user misses the thing.</p>
<p><strong>Latency is a correctness constraint.</strong> If your stream is forty seconds behind broadcast, viewers learn the result from their phone before they see it. That is a product failure with no technical symptom — every metric is green and the experience is ruined.</p>
<p><strong>Peak is unpredictable within a predictable window.</strong> You know the event starts at 8 p.m. You do not know that a particular race will be close and that traffic will spike 3× in the final ninety seconds.</p>
<p><strong>You cannot test at scale.</strong> There is no staging environment with fifty million concurrent viewers. Load tests approximate. The real event is the test.</p>
<h2 id="the-architecture-roughly">the architecture, roughly<a class="anchor" href="#the-architecture-roughly" aria-label="link to this section">#</a></h2>
<p><strong>Ingest</strong> — camera feeds encoded at the venue, sent over dedicated circuits, with redundant paths because a fiber cut during a final is a career event.</p>
<p><strong>Transcode</strong> — one source becomes a ladder of bitrates and resolutions, times several codecs, times audio tracks and languages. That is a large multiplication and it happens in real time with no room to fall behind.</p>
<p><strong>Packaging</strong> — segmented into chunks, typically 2 to 6 seconds, in HLS and DASH. Low-latency variants use much smaller chunks or chunked transfer encoding to cut the delay, at the cost of more requests and worse cache behavior.</p>
<p><strong>Distribution</strong> — multiple CDNs, always. Not for capacity alone but because CDNs have bad days, and a client-side switching layer that measures performance and moves traffic is the difference between a degraded minute and an outage.</p>
<p><strong>The client</strong> — adaptive bitrate logic that decides which quality to request based on measured bandwidth and buffer level. This is where a surprising amount of the perceived quality difference between services actually lives.</p>
<h2 id="the-parts-that-are-counterintuitive">the parts that are counterintuitive<a class="anchor" href="#the-parts-that-are-counterintuitive" aria-label="link to this section">#</a></h2>
<p><strong>Cache hit ratio is everything, and low-latency streaming ruins it.</strong> A 6-second segment requested by a million people is one origin fetch and a million edge hits. Cut to 1-second segments for lower latency and you have six times the requests, each with a shorter window to accumulate hits. Latency and efficiency are in direct tension and every service picks a point on that curve.</p>
<p><strong>The last mile is not yours and dominates the experience.</strong> Home wifi, congested cell towers, and oversubscribed ISP links cause most of the buffering users blame on you. The only lever you have is graceful adaptation — dropping quality smoothly rather than stalling.</p>
<p><strong>Advertising insertion is a distributed systems problem.</strong> Server-side ad insertion means personalizing a stream per viewer while keeping segments cacheable, which is a genuinely hard constraint and is where a lot of live-stream failures actually originate.</p>
<p><strong>The failure mode that matters is the thundering herd on recovery.</strong> When something breaks and comes back, every client reconnects at once. Without jitter on the retry, recovery causes a second outage. This is the single most common way these events go badly.</p>
<h2 id="the-lesson-that-generalizes">the lesson that generalizes<a class="anchor" href="#the-lesson-that-generalizes" aria-label="link to this section">#</a></h2>
<p>Almost every hard part here is about <strong>synchronized demand you cannot smooth and cannot refuse</strong>.</p>
<p>Most systems get to spread load over time, shed it, or queue it. When you cannot do any of those, you are left with provisioning for peak, redundancy at every layer, and graceful degradation that preserves the core experience.</p>
<p>That is expensive and it is the only thing that works. Which is worth remembering when someone proposes autoscaling as the answer to a spike that arrives faster than an instance can boot.</p>
<p>Some problems you solve with capacity you already own.</p>]]></content:encoded></item><item><title>The electricity bill arrives</title><link>https://readme.news/the-electricity-bill-arrives/</link><guid isPermaLink="true">https://readme.news/the-electricity-bill-arrives/</guid><pubDate>Sat, 17 Jan 2026 09:00:00 +0000</pubDate><description>Datacenter power demand is showing up in residential rates, and the regulatory fight is now the main constraint on compute supply.</description><content:encoded><![CDATA[<p>The AI infrastructure buildout has reached the phase where it shows up on other people's bills, and the resulting political process is now a more important constraint on compute supply than anything happening in a fab.</p>
<h2 id="the-mechanism">the mechanism<a class="anchor" href="#the-mechanism" aria-label="link to this section">#</a></h2>
<p>A datacenter campus needs a lot of power at one point on the grid. Delivering it requires transmission upgrades, substation construction, and often new generation. Those cost billions and take years.</p>
<p>Who pays is decided by state public utility commissions in rate cases. The options:</p>
<ul><li><strong>The datacenter pays</strong>, through special contracts with minimum-take provisions and upfront infrastructure contributions.</li><li><strong>All ratepayers pay</strong>, spread across the customer base, which is how transmission has traditionally been socialized.</li><li><strong>Some blend</strong>, which is what actually happens.</li></ul>
<p>Utilities generally prefer the blend, because socialized costs are easier to recover and because they earn a regulated return on capital investment — which means building more infrastructure is directly profitable for them regardless of who uses it.</p>
<p>Consumer advocates object. Datacenter operators object to bearing the full cost. Commissions split the difference, differently in every state.</p>
<h2 id="why-this-is-now-the-binding-constraint">why this is now the binding constraint<a class="anchor" href="#why-this-is-now-the-binding-constraint" aria-label="link to this section">#</a></h2>
<p>Chips are available. Capital is available. What is not available is a grid interconnection on a useful timeline.</p>
<p>Interconnection <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> in most US markets run years. Transformer lead times remain long. Transmission line construction requires siting approval, which requires public process, which is where projects die.</p>
<p>So the sequence for new capacity is: secure a site with grid access, sign a power agreement, wait. The waiting is the schedule. Everything else can be compressed with money; this cannot.</p>
<h2 id="what-it-means-for-compute-prices">what it means for compute prices<a class="anchor" href="#what-it-means-for-compute-prices" aria-label="link to this section">#</a></h2>
<p>Two effects pulling in opposite directions.</p>
<p><strong>Regional differentiation increases.</strong> Regions with available power and cooperative regulators get capacity. Regions without do not. That produces real price differences between cloud regions for the same instance type, larger than the historical spread.</p>
<p><strong>Off-peak becomes genuinely cheaper.</strong> Grid economics are about peak demand. Workloads that can shift to off-peak hours are worth real money to operators, and that value will get passed through as pricing. Batch inference, training runs, and anything asynchronous is a candidate.</p>
<h2 id="what-to-actually-do">what to actually do<a class="anchor" href="#what-to-actually-do" aria-label="link to this section">#</a></h2>
<p><strong>Check regional pricing before you pick a region by habit.</strong> The default region in your organization was chosen years ago for reasons that may no longer hold. For a large workload the difference is material.</p>
<p><strong>Make batch work time-flexible.</strong> If your nightly job can run in a four-hour window instead of at a fixed time, you can take spot capacity and off-peak pricing. This is a small engineering change with a large cost effect and almost nobody does it.</p>
<p><strong>Measure your actual utilization.</strong> The cheapest watt is the one you do not use. A very large fraction of provisioned cloud compute runs at low utilization, and in an environment where power is the constraint, that waste is now expensive rather than merely inelegant.</p>
<p><strong>Watch your provider's regional capacity announcements</strong> if you are planning anything large. Capacity comes online in steps tied to substation energization dates, and knowing the schedule is worth something.</p>
<h2 id="the-part-that-is-not-about-your-bill">the part that is not about your bill<a class="anchor" href="#the-part-that-is-not-about-your-bill" aria-label="link to this section">#</a></h2>
<p>There is a legitimate public policy question here that the industry mostly does not want to engage with: whether the cost of connecting datacenters should be borne by households whose electricity bills are rising.</p>
<p>The industry's answer is usually that datacenters bring jobs and tax revenue, which is true and is a smaller number than the capital involved. The honest version is that the buildout is happening faster than the institutions that allocate its costs can deliberate, and the allocations being made now will be argued about for a decade.</p>
<p>Engineers do not decide this. Engineers do decide how much power their systems consume, and that has stopped being an abstract concern.</p>]]></content:encoded></item><item><title>Post-quantum migration is a 2026 project</title><link>https://readme.news/post-quantum-migration-is-a-2026-project/</link><guid isPermaLink="true">https://readme.news/post-quantum-migration-is-a-2026-project/</guid><pubDate>Sun, 11 Jan 2026 09:00:00 +0000</pubDate><description>Harvest-now-decrypt-later makes this urgent for anything with a long confidentiality horizon. The tooling is finally ready.</description><content:encoded><![CDATA[<p>The post-quantum transition has been discussed as a distant problem for a decade. It is now a project with a schedule, standardized algorithms, and shipping implementations, and the reason to start is not that quantum computers exist.</p>
<h2 id="harvest-now-decrypt-later">harvest now, decrypt later<a class="anchor" href="#harvest-now-decrypt-later" aria-label="link to this section">#</a></h2>
<p>The threat model that makes this urgent has nothing to do with when a cryptographically relevant quantum computer arrives.</p>
<p>An adversary records your encrypted traffic today and stores it. When they can break the key exchange — in five years, in fifteen — they decrypt the archive.</p>
<p>So the question is not "when will quantum computers work." It is: <strong>how long does your data need to stay confidential?</strong></p>
<ul><li>Session tokens: minutes. Do not care.</li><li>Customer PII: years to decades. Care a lot.</li><li>Medical records: a lifetime. Care enormously.</li><li>Government and defense: generational.</li><li>Source code and trade secrets: depends, usually longer than you think.</li></ul>
<p>If anything you transmit has a confidentiality horizon past roughly 2035, traffic you send today is already at risk. That is the whole argument and it does not depend on any prediction about quantum hardware.</p>
<h2 id="what-is-standardized">what is standardized<a class="anchor" href="#what-is-standardized" aria-label="link to this section">#</a></h2>
<p>NIST finalized the core standards:</p>
<ul><li><strong>ML-KEM</strong> (FIPS 203), formerly Kyber — key encapsulation. This is the one that matters for TLS.</li><li><strong>ML-DSA</strong> (FIPS 204), formerly Dilithium — digital signatures.</li><li><strong>SLH-DSA</strong> (FIPS 205), formerly SPHINCS+ — hash-based signatures, conservative fallback with larger signatures.</li></ul>
<p>The guidance across national security agencies is consistent: begin migration now, complete it well before 2035, prioritize by confidentiality horizon.</p>
<h2 id="what-is-already-shipping">what is already shipping<a class="anchor" href="#what-is-already-shipping" aria-label="link to this section">#</a></h2>
<p>More than most people realize.</p>
<p><strong>TLS key exchange.</strong> Hybrid X25519 plus ML-KEM is deployed by default in major browsers and supported by major CDNs and load balancers. A meaningful fraction of web traffic is already post-quantum protected for key exchange, and most people running those services did not do anything to enable it.</p>
<p>Check yours:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">openssl s_client -connect example.com:443 -groups X25519MLKEM768 &lt;/dev/null 2&gt;&amp;1 \
  | grep -i "negotiated\|group"</code></pre></div>
<p><strong>SSH.</strong> OpenSSH has shipped post-quantum key exchange by default for several releases. If your servers are current, your SSH sessions are already hybrid.</p>
<p><strong>Signatures are lagging</strong>, and that is the harder half. Certificate chains, code signing, and firmware verification all involve long-lived trust anchors and ecosystem-wide coordination. ML-DSA signatures and public keys are substantially larger than ECDSA, which has real consequences for handshake size, embedded devices, and anything with a fixed-size signature field.</p>
<h2 id="what-to-actually-do-this-year">what to actually do this year<a class="anchor" href="#what-to-actually-do-this-year" aria-label="link to this section">#</a></h2>
<p><strong>1. Inventory your cryptography.</strong> Where do you use asymmetric crypto, with what algorithm, and what is the confidentiality or integrity horizon? Most organizations cannot answer this and the exercise is more valuable than any individual migration step.</p>
<p><strong>2. Get to hybrid key exchange.</strong> For most people this means: update TLS libraries, update the load balancer, verify the negotiated group. It is largely free and it addresses the harvest-now threat, which is the urgent one.</p>
<p><strong>3. Build crypto-agility.</strong> The lasting fix is not "migrate to ML-KEM." It is "be able to change algorithms without a rewrite." Hard-coded algorithm choices scattered through a codebase are the actual problem, and they will be the problem again for whatever comes after this.</p>
<p>Centralize crypto in one module. Make the algorithm a configuration value. Version your protocol so you can negotiate.</p>
<p><strong>4. Ask your vendors.</strong> Every SaaS provider, every library, every hardware security module. The answers will be uneven and asking creates the pressure that fixes it.</p>
<p><strong>5. Do not roll your own.</strong> Use the vetted implementations. The failure mode for post-quantum crypto is a side-channel in a hand-written implementation of a lattice operation, and that failure is silent.</p>
<h2 id="the-thing-that-makes-this-hard">the thing that makes this hard<a class="anchor" href="#the-thing-that-makes-this-hard" aria-label="link to this section">#</a></h2>
<p>It is a migration with no visible benefit. Nothing gets faster. No feature ships. The success condition is that in fifteen years, nothing bad happens.</p>
<p>That is the hardest kind of project to fund, and it is why the organizations that do it well will be the ones that started when it was still early enough to be cheap.</p>]]></content:encoded></item><item><title>CES 2026: the power supply is the product</title><link>https://readme.news/ces-2026-the-power-supply-is-the-product/</link><guid isPermaLink="true">https://readme.news/ces-2026-the-power-supply-is-the-product/</guid><pubDate>Tue, 06 Jan 2026 09:00:00 +0000</pubDate><description>Another year of AI in appliances, and one genuine trend hiding underneath: everything is now thermally constrained.</description><content:encoded><![CDATA[<p>The annual Las Vegas exercise in stapling language models to inventory has concluded. As always, the interesting signal is in the parts nobody put on a keynote slide.</p>
<h2 id="the-actual-trend">the actual trend<a class="anchor" href="#the-actual-trend" aria-label="link to this section">#</a></h2>
<p>Every category at this show is now constrained by thermals and power delivery rather than by compute.</p>
<p><strong>Laptops</strong> with neural accelerators that cannot sustain their rated throughput for more than a few minutes in a thin chassis. The TOPS number on the sticker is a peak, and the peak lasts as long as the thermal mass does.</p>
<p><strong>Handhelds</strong> where battery life is the entire product decision and every added capability is a subtraction from it.</p>
<p><strong>Home devices</strong> that want to run models locally and cannot, because the power envelope of something that sits on a shelf is a few watts.</p>
<p><strong>Desktops</strong> where the power supply recommendations for a current-generation GPU have crept past what a lot of household circuits comfortably deliver alongside everything else in the room.</p>
<p>This is what a computing era looks like when it hits a physical wall. The interesting engineering for the next several years is efficiency, not capability — which is historically when the best engineering happens.</p>
<h2 id="what-to-actually-take-from-the-show">what to actually take from the show<a class="anchor" href="#what-to-actually-take-from-the-show" aria-label="link to this section">#</a></h2>
<p><strong><a class="xref" href="/pixel-10-and-the-on-device-model-as-a-platform-feature/" title="Pixel 10 and the on-device model as a platform feature">On-device</a> inference is a spec sheet lie in the thin-and-light category.</strong> If you are shipping software that assumes a local NPU, test on a machine that has been running for twenty minutes, not on a cold one. The difference is large and nobody benchmarks it.</p>
<p><strong>Unified memory keeps winning.</strong> Every serious local-AI machine announced has shared CPU/GPU memory at high capacity. Discrete VRAM is a bad fit for a workload where capacity matters more than peak bandwidth, and the industry has figured that out.</p>
<p><strong>The robot demos are still teleoperated.</strong> They have been for four years. Watch for the operator's hands, or the suspiciously smooth trajectory, or the fact that the demo never deviates from the script. There is real progress in robotics and it is happening in warehouses, not on stages.</p>
<h2 id="the-accessory-economy">the accessory economy<a class="anchor" href="#the-accessory-economy" aria-label="link to this section">#</a></h2>
<p>A large fraction of the floor was accessories for devices that do not exist yet: cases, docks, and mounts for AI wearables. That is a leading indicator of nothing except that the accessory industry moves fast and takes risks cheaply.</p>
<h2 id="the-one-thing-i-would-actually-buy">the one thing I would actually buy<a class="anchor" href="#the-one-thing-i-would-actually-buy" aria-label="link to this section">#</a></h2>
<p>The mundane one: displays got better and cheaper. High-refresh, high-resolution, accurate-color monitors are at prices that were flagship prices three years ago.</p>
<p>If you have been running the same monitor since 2020 and you spend eight hours a day looking at text on it, that is the highest-value hardware purchase available to you, and it will not be mentioned in a single keynote.</p>
<h2 id="the-meta-observation">the meta-observation<a class="anchor" href="#the-meta-observation" aria-label="link to this section">#</a></h2>
<p>CES has been a bad predictor of what matters for about a decade. The genuinely important hardware of the last ten years — the M1, the H100, TPUs, the shift to ARM in servers — was announced at industry events or in blog posts, to audiences who understood it, without a stage show.</p>
<p>The consumer electronics show is now mostly a trade event for retail buyers, covered as if it were a technology forecast. Read the specs. Skip the narrative.</p>]]></content:encoded></item><item><title>re:Invent and the year of the boring cloud announcement</title><link>https://readme.news/reinvent-and-the-year-of-the-boring-cloud-announcement/</link><guid isPermaLink="true">https://readme.news/reinvent-and-the-year-of-the-boring-cloud-announcement/</guid><pubDate>Tue, 02 Dec 2025 09:00:00 +0000</pubDate><description>Silicon, agents, and a keynote that mostly described things that already existed. That&#x27;s a healthy sign.</description><content:encoded><![CDATA[<p>AWS re:Invent is underway and the announcement pace is, as always, absurd. Most of it will not matter to you. Here is the part that will.</p>
<h2 id="silicon">silicon<a class="anchor" href="#silicon" aria-label="link to this section">#</a></h2>
<p>Amazon continues pushing Trainium and Inferentia as the alternative to buying Nvidia. The pitch is price-performance for customers who can tolerate a different software stack.</p>
<p>The honest state of it: the hardware is competitive on paper, the software ecosystem is meaningfully behind CUDA, and the gap is closing slowly. If your workload runs through PyTorch with standard operations, the port is manageable. If you have custom kernels, it is a project.</p>
<p>The strategic point is the same for every hyperscaler: reducing dependency on a single supplier with enormous pricing power. Whether the customer benefit materializes depends entirely on whether the savings get passed through.</p>
<h2 id="the-agent-announcements">the agent announcements<a class="anchor" href="#the-agent-announcements" aria-label="link to this section">#</a></h2>
<p>Every cloud vendor is now shipping agent infrastructure: runtimes, memory services, gateways for tool access, identity for agents, observability for agent traces.</p>
<p>This category is real. Running agents in production has genuine infrastructure requirements that are different from running services:</p>
<ul><li><strong>Long-lived sessions</strong> with state that outlives a request.</li><li><strong>Non-deterministic execution paths</strong> that make traditional tracing awkward.</li><li><strong>Cost per invocation that varies by orders of magnitude.</strong></li><li><strong>Identity and permission scoping</strong> for a thing acting on a user's behalf.</li><li><strong>Human approval gates</strong> in the middle of automated flows.</li></ul>
<p>Those are real problems and the tooling is early everywhere. Evaluate on whether it solves a problem you actually have, not on whether the demo was good.</p>
<h2 id="the-pattern-i-would-push-back-on">the pattern I would push back on<a class="anchor" href="#the-pattern-i-would-push-back-on" aria-label="link to this section">#</a></h2>
<p>Every vendor's agent framework wants to be the place your orchestration lives. That is a <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a> position, and orchestration is the layer most likely to be absorbed by the models themselves — it has been happening steadily for two years.</p>
<p>Keep your orchestration portable. Use the managed pieces for the genuinely hard infrastructure — identity, secure tool access, session storage — and keep the logic in your own code.</p>
<h2 id="the-underrated-announcements">the underrated announcements<a class="anchor" href="#the-underrated-announcements" aria-label="link to this section">#</a></h2>
<p>The ones nobody writes about and everybody uses:</p>
<ul><li>Incremental improvements to S3 consistency and performance.</li><li>Networking latency reductions.</li><li>Cost management tooling that is slightly less bad.</li><li>Database engine version updates.</li></ul>
<p>These are worth more to most organizations than any AI announcement, and they get four slides between two hours of agent demos.</p>
<h2 id="the-meta-observation">the meta-observation<a class="anchor" href="#the-meta-observation" aria-label="link to this section">#</a></h2>
<p>Cloud conferences have gotten less interesting, and that is good. It means the platform is mature. The exciting years of a platform are the years when fundamental things are missing.</p>
<p>The interesting question for AWS is not what they announced. It is whether the operational excellence that justified the premium is still there after a year that included a major regional outage. That is an execution question and it does not get answered at a conference.</p>
<h2 id="what-to-actually-do-with-this">what to actually do with this<a class="anchor" href="#what-to-actually-do-with-this" aria-label="link to this section">#</a></h2>
<p>Skip the keynote. Read the "what's new" feed filtered to the services you actually use. Look for the deprecations, which are the announcements that will cost you time and which are never on stage.</p>
<p>And check your bill. The single highest-value hour available to most engineering organizations is someone competent looking at the AWS bill line by line, and almost nobody does it.</p>]]></content:encoded></item><item><title>Nvidia's quarter and the question nobody can answer</title><link>https://readme.news/nvidias-quarter-and-the-question-nobody-can-answer/</link><guid isPermaLink="true">https://readme.news/nvidias-quarter-and-the-question-nobody-can-answer/</guid><pubDate>Fri, 21 Nov 2025 09:00:00 +0000</pubDate><description>Another enormous beat, another set of concerns about circular financing. Both facts are real.</description><content:encoded><![CDATA[<p>Nvidia reported another quarter far above expectations, with data center revenue continuing to grow at a rate that would be implausible in any other context.</p>
<p>The stock's reaction was muted relative to the beat, which tells you the debate has moved from "is demand real" to "is demand <em>sustainable</em>, and how much of it is funded by Nvidia."</p>
<h2 id="the-bull-case">the bull case<a class="anchor" href="#the-bull-case" aria-label="link to this section">#</a></h2>
<p>Straightforward and well supported:</p>
<ul><li>Every hyperscaler raised capital expenditure guidance again.</li><li>Inference demand is growing faster than training demand, and inference is the recurring workload rather than the one-time one.</li><li>Reasoning models consume dramatically more inference compute per request than their predecessors, and adoption is rising.</li><li>Supply remains the constraint. Lead times are long. Customers are queueing.</li><li>The <a class="xref" href="/gtc-2025-a-roadmap-to-2027-and-a-warning-about-power/" title="GTC 2025: a roadmap to 2027 and a warning about power">rack-scale</a> systems business has a moat that individual chip competition does not touch — the interconnect is the product.</li></ul>
<h2 id="the-bear-case">the bear case<a class="anchor" href="#the-bear-case" aria-label="link to this section">#</a></h2>
<p>Also straightforward:</p>
<ul><li>A meaningful share of revenue traces to customers Nvidia has invested in or financed, which makes the demand signal less independent.</li><li>Depreciation schedules on AI hardware are assumed at five to six years. If the useful life is closer to three — which some operators argue, given the pace of generational improvement — reported profitability across the sector is overstated.</li><li>Hyperscalers are building their own silicon. Google's TPUs are mature, Amazon's Trainium is shipping in volume, and every one of those deployments is a substituted Nvidia sale.</li><li>Model efficiency improvements keep arriving. If capability-per-FLOP keeps improving as fast as it has, required FLOPs for a given capability fall.</li><li>The financing environment for the buildout depends on continued access to cheap debt.</li></ul>
<h2 id="what-nobody-knows">what nobody knows<a class="anchor" href="#what-nobody-knows" aria-label="link to this section">#</a></h2>
<p>Whether AI application revenue will eventually justify the infrastructure spend.</p>
<p>Current annualized revenue across the AI application layer is a fraction of annual AI capital expenditure. That gap can close — infrastructure is built ahead of demand in every capital cycle, and railroads, fiber, and cloud all looked insane at the equivalent stage.</p>
<p>It can also not close. Fiber overbuild in 2000 was followed by a decade of dark fiber and a lot of bankruptcies, and the eventual users of that fiber were not the companies that laid it.</p>
<p>Both patterns are real. The people confidently predicting which one applies here are pattern-matching, not analyzing, and that includes the ones I agree with.</p>
<h2 id="why-an-engineer-should-care">why an engineer should care<a class="anchor" href="#why-an-engineer-should-care" aria-label="link to this section">#</a></h2>
<p>Not for investing advice. For planning.</p>
<p><strong>Compute pricing is not going to fall smoothly.</strong> If the capex cycle continues, capacity comes online in steps and prices drift down. If it contracts, capacity tightens and prices firm. Do not build a business model that requires a specific trajectory.</p>
<p><strong>Efficiency work has enduring value.</strong> Whatever happens to the capex cycle, using less compute for the same result is good. Prompt <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a>, model routing, smaller models for routine tasks, batch processing where latency permits — all of it pays regardless of the macro environment.</p>
<p><strong>Multi-provider capability is cheap insurance.</strong> The cost of abstracting your model calls is a day. The cost of being locked to a provider whose pricing or availability changes is much larger.</p>
<h2 id="the-sentence-i-keep-coming-back-to">the sentence I keep coming back to<a class="anchor" href="#the-sentence-i-keep-coming-back-to" aria-label="link to this section">#</a></h2>
<p>The infrastructure being built is real, the demand today is real, and whether they match at the scale being assumed is genuinely unknown to everyone including the people spending the money.</p>
<p>That is an uncomfortable place to be, and pretending otherwise — in either direction — is the main thing to avoid.</p>]]></content:encoded></item><item><title>Cloudflare falls over because of a config file</title><link>https://readme.news/cloudflare-falls-over-because-of-a-config-file/</link><guid isPermaLink="true">https://readme.news/cloudflare-falls-over-because-of-a-config-file/</guid><pubDate>Thu, 20 Nov 2025 09:00:00 +0000</pubDate><description>A permissions change doubles the size of a generated feature file, which overflows a fixed-size buffer, which 500s a fifth of the web.</description><content:encoded><![CDATA[<p>Cloudflare had a significant outage on Tuesday, returning 5xx errors across a large portion of its network for several hours. Their published postmortem is detailed and worth reading in full.</p>
<p>The chain of events is a small masterpiece of the genre.</p>
<h2 id="what-happened">what happened<a class="anchor" href="#what-happened" aria-label="link to this section">#</a></h2>
<p>A database permissions change caused a query that generates a Bot Management feature configuration file to return <strong>duplicate rows</strong>. The query had been returning one row per feature; after the permissions change it returned rows from multiple underlying schemas.</p>
<p>The generated file therefore roughly doubled in size.</p>
<p>The proxy that consumes this file preallocates a fixed-size buffer sized against a limit of 200 features — comfortably above the ~60 actually in use. The doubled file exceeded that limit.</p>
<p>The Rust code handling this hit an unrecoverable error path and the proxy panicked rather than degrading. Because the configuration file propagates network-wide every few minutes, the failure propagated network-wide within minutes.</p>
<p>Recovery was complicated by the fact that the bad file kept regenerating and redeploying.</p>
<h2 id="the-lessons-which-are-old">the lessons, which are old<a class="anchor" href="#the-lessons-which-are-old" aria-label="link to this section">#</a></h2>
<p><strong>Configuration is code and needs the same rigor.</strong> This was not a code deploy. It was a data change that propagated to production automatically with no staging, no canary, and no validation. Config deployment pipelines are consistently held to a lower standard than code deployment pipelines, and config causes a large share of major outages.</p>
<p><strong>Fixed-size limits need to fail soft.</strong> The 200-feature limit was reasonable. Panicking when exceeded was not. The correct behavior for a proxy encountering an oversized config is to log loudly, alert, and continue with the previous known-good version.</p>
<p><strong>Validate generated artifacts before propagation.</strong> A size check, a schema check, a sanity check on row count against the previous version — any of these would have caught it. Generated files should be validated as rigorously as user input, because the generator can be wrong.</p>
<p><strong>Blast radius follows deployment speed.</strong> Config that propagates globally in minutes is a feature until it propagates a bad config globally in minutes. Staged rollout applies to configuration too.</p>
<p><strong>Have a <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">kill switch</a> for automated pipelines.</strong> Much of the recovery time went to stopping the thing that kept redeploying the bad file. Every automated deployment path needs a way to stop it that does not require fixing the underlying problem first.</p>
<h2 id="the-rust-note">the Rust note<a class="anchor" href="#the-rust-note" aria-label="link to this section">#</a></h2>
<p>This will be used as an argument about Rust, and it should not be.</p>
<p>The panic was a deliberate choice at that call site — the code used a construct that terminates on error rather than propagating it. That is a design decision about error handling, available in any language. In C the equivalent code would have written past the buffer, which is worse.</p>
<p>The actual lesson is about where you choose to make errors fatal. In a proxy handling live traffic, almost nothing should be fatal. Degrade, alert, continue. "Fail fast" is good advice for a batch job and bad advice for a load balancer.</p>
<h2 id="the-credit-due">the credit due<a class="anchor" href="#the-credit-due" aria-label="link to this section">#</a></h2>
<p>Cloudflare published a detailed technical postmortem within a day, named the specific code path, and did not hide behind "an issue with a third-party provider."</p>
<p>That is how it should be done and it is rarer than it should be. A company that publishes real postmortems earns more trust than one that never has visible incidents, because the second one is not telling you about them.</p>]]></content:encoded></item><item><title>us-east-1 goes down and takes a large chunk of the internet with it</title><link>https://readme.news/us-east-1-goes-down-and-takes-a-large-chunk-of-the-internet-with-it/</link><guid isPermaLink="true">https://readme.news/us-east-1-goes-down-and-takes-a-large-chunk-of-the-internet-with-it/</guid><pubDate>Tue, 21 Oct 2025 09:00:00 +0000</pubDate><description>A DNS race condition in DynamoDB&#x27;s automation cascades across dozens of AWS services. The lesson is about coupling, not DNS.</description><content:encoded><![CDATA[<p>AWS's us-east-1 region suffered a multi-hour outage yesterday that affected a very large number of services and, through them, a very large fraction of consumer internet applications.</p>
<p>AWS's public summary attributes the trigger to a latent race condition in the automation that manages DynamoDB's DNS records, which resulted in an empty record set for a regional endpoint and no automatic recovery path.</p>
<h2 id="the-cascade">the cascade<a class="anchor" href="#the-cascade" aria-label="link to this section">#</a></h2>
<p>The failure sequence is the interesting part.</p>
<p>DynamoDB's endpoint became unresolvable. That alone would be bad. What made it a regional event is that <strong>an enormous number of AWS's own services use DynamoDB internally.</strong> The EC2 instance launch path, IAM's control plane, Lambda's invocation machinery, and dozens of others depend on it.</p>
<p>So the failure propagated: DynamoDB down means new EC2 instances cannot launch, which means autoscaling cannot replace failing capacity, which means load shedding, which means more failures. Network Load Balancer health checks destabilized. The recovery itself was slowed by the backlog of queued work that had accumulated.</p>
<p>This is textbook <strong>metastable failure</strong>: a system that is stable under normal load and stable under no load, but which, once pushed past a threshold, sustains its own failure through retry amplification and queue buildup even after the original trigger is fixed.</p>
<h2 id="why-us-east-1">why us-east-1<a class="anchor" href="#why-us-east-1" aria-label="link to this section">#</a></h2>
<p>It is the oldest region, the largest, and the default in approximately every tutorial ever written. Several global AWS control planes are homed there — IAM, CloudFront configuration, Route 53's control plane, and others. That means a us-east-1 event has global blast radius even for customers with no resources in the region.</p>
<p>That architecture is a historical artifact. It is also extremely hard to change now, which is a lesson about early decisions in systems that grow.</p>
<h2 id="the-honest-customer-takeaway">the honest customer takeaway<a class="anchor" href="#the-honest-customer-takeaway" aria-label="link to this section">#</a></h2>
<p>The reflexive response is "multi-region." Before you spend a year on that, do the arithmetic.</p>
<p><strong>Multi-region active-active is genuinely hard.</strong> Data consistency across regions, failover testing that actually works, doubled infrastructure cost, and a substantially more complex system that fails in new ways. Many organizations that attempt it end up with a system that is <em>less</em> reliable overall because the complexity introduces more <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> than the regional risk it removes.</p>
<p><strong>The dependency you cannot escape.</strong> If your multi-region architecture depends on a global control plane that lives in us-east-1, you did not achieve independence. Check this specifically. A lot of people discovered it yesterday.</p>
<p><strong>What is actually worth doing, in order:</strong></p>
<ol><li><strong>Know your dependencies.</strong> Most teams cannot enumerate what their service requires to start. Write it down. The exercise is revealing.</li><li><strong>Static stability.</strong> Design so existing capacity keeps serving when the control plane is unavailable. If your service needs to call an API to keep running, it will stop when that API stops. Cache aggressively, fail open where safe, and do not require a control plane call on the request path.</li><li><strong>Graceful degradation.</strong> Decide in advance which features can be turned off. A checkout that works without recommendations is much better than a site that is down.</li><li><strong>Exponential backoff with jitter, and circuit breakers.</strong> Retry storms are what turns an incident into an outage. This is the single highest-leverage code change available.</li><li><strong>Multi-region for the tier that genuinely warrants it.</strong> Which is usually not everything.</li></ol>
<h2 id="the-industry-level-observation">the industry-level observation<a class="anchor" href="#the-industry-level-observation" aria-label="link to this section">#</a></h2>
<p>A meaningful fraction of the world's software depends on a small number of regions operated by a small number of companies. That concentration produces excellent reliability most of the time and correlated failure occasionally.</p>
<p>There is no individual fix. Every company independently choosing the most reliable provider produces exactly this concentration. It is a collective action problem, and the only actors who can address it are regulators thinking about systemic risk, who are — belatedly — starting to.</p>]]></content:encoded></item><item><title>OpenAI and Nvidia sign a circular deal</title><link>https://readme.news/openai-and-nvidia-sign-a-circular-deal/</link><guid isPermaLink="true">https://readme.news/openai-and-nvidia-sign-a-circular-deal/</guid><pubDate>Tue, 23 Sep 2025 09:00:00 +0000</pubDate><description>Up to $100 billion of investment tied to gigawatts of deployment. The financing structures are getting interesting.</description><content:encoded><![CDATA[<p>Nvidia and OpenAI announced a letter of intent under which Nvidia would invest up to $100 billion in OpenAI, staged against the deployment of at least 10 gigawatts of Nvidia systems.</p>
<p>Read that structure carefully, because it is the interesting part.</p>
<h2 id="the-circularity">the circularity<a class="anchor" href="#the-circularity" aria-label="link to this section">#</a></h2>
<p>Nvidia invests in OpenAI. OpenAI uses the money to buy Nvidia systems. The purchase is recognized as Nvidia revenue. The investment is staged against deployment milestones.</p>
<p>This is not fraud and it is not unusual in capital-intensive industries — vendor financing has been standard in telecom, aviation, and semiconductor equipment for decades. A supplier finances a customer's purchase because the supplier has the balance sheet and wants the volume.</p>
<p>It does deserve scrutiny for a specific reason: it makes the demand signal less informative. When a supplier funds its customer's purchases, revenue growth no longer cleanly indicates independent market demand. Some portion of it is the supplier's own capital cycling through.</p>
<p>Analysts have been tracking a widening set of these arrangements across the AI sector — investments in customers, prepayments, equity stakes in companies that are also large purchasers. Individually each is defensible. Collectively they make the sector's growth figures harder to interpret.</p>
<h2 id="the-gigawatt-as-a-unit">the gigawatt as a unit<a class="anchor" href="#the-gigawatt-as-a-unit" aria-label="link to this section">#</a></h2>
<p>Note what is being measured. Not chips, not dollars, not FLOPs. <strong>Gigawatts.</strong></p>
<p>Ten gigawatts is on the order of the electricity consumption of a large metropolitan area. It is roughly ten large nuclear reactors' worth of continuous generation.</p>
<p>The industry has converged on power as the natural unit because power is the binding constraint. You can order chips. You cannot order a substation and have it next quarter.</p>
<p>The consequences flow outward: electricity prices in datacenter-heavy regions, grid interconnection <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a>, transmission buildout, and the political economy of who pays for it. Several US utility regulators are now handling rate cases that are effectively about whether residential customers subsidize datacenter connections.</p>
<p>That fight is going to define a lot of the next five years and it is being had in public utility commission hearings that nobody in tech reads.</p>
<h2 id="what-it-means-for-you">what it means for you<a class="anchor" href="#what-it-means-for-you" aria-label="link to this section">#</a></h2>
<p>If you are building on AI APIs, the practical questions are:</p>
<p><strong>Is my provider's capacity growing?</strong> <a class="xref" href="/rate-limits-are-a-product-decision-not-an-infrastructure-one/" title="Rate limits are a product decision, not an infrastructure one">Rate limits</a> and availability during demand spikes are the observable symptom. Announcements like this are a positive signal for capacity.</p>
<p><strong>Am I exposed to a single provider's economics?</strong> If the financing environment tightens, pricing changes. Multi-provider capability is cheap insurance and you should have built it anyway for reliability reasons.</p>
<p><strong>Are my costs actually falling?</strong> Per-token prices have fallen consistently. Per <em>task</em> costs have not fallen as much, because reasoning models consume more tokens. Measure the thing you pay for.</p>
<h2 id="the-honest-uncertainty">the honest uncertainty<a class="anchor" href="#the-honest-uncertainty" aria-label="link to this section">#</a></h2>
<p>Nobody knows whether the capex cycle is correctly sized. The bull case is that inference demand compounds and every gigawatt gets used. The bear case is that efficiency improvements outrun demand and a lot of concrete is stranded.</p>
<p>Both are held sincerely by smart people with access to the same information. That is what genuine uncertainty looks like, and anyone expressing confidence in either direction is telling you about their position, not about the world.</p>]]></content:encoded></item><item><title>Cloudflare flips the default and starts charging crawlers</title><link>https://readme.news/cloudflare-flips-the-default-and-starts-charging-crawlers/</link><guid isPermaLink="true">https://readme.news/cloudflare-flips-the-default-and-starts-charging-crawlers/</guid><pubDate>Tue, 01 Jul 2025 09:00:00 +0000</pubDate><description>AI bots blocked unless allowed, plus a marketplace for per-crawl payment. A fifth of the web changes its robots policy at once.</description><content:encoded><![CDATA[<p>Cloudflare announced today that new domains on its network will block AI crawlers by default, and launched a pay-per-crawl marketplace letting site operators charge for access.</p>
<p>Cloudflare sits in front of roughly a fifth of the web. A default change at that position is not a product launch, it is a policy change for the internet.</p>
<h2 id="the-mechanism">the mechanism<a class="anchor" href="#the-mechanism" aria-label="link to this section">#</a></h2>
<p>Two pieces.</p>
<p><strong>Default blocking.</strong> New zones get AI crawler blocking on unless the operator opts out. Cloudflare maintains the bot classification — separating search crawlers, which drive traffic back, from training crawlers, which do not.</p>
<p><strong>Pay-per-crawl.</strong> A site sets a price. A crawler that wants the content gets an HTTP 402 Payment Required with terms. Cloudflare handles settlement.</p>
<p>HTTP 402 has been "reserved for future use" since 1997. It is genuinely funny that this is what activated it.</p>
<h2 id="why-the-old-system-failed">why the old system failed<a class="anchor" href="#why-the-old-system-failed" aria-label="link to this section">#</a></h2>
<p><code>robots.txt</code> is a request, not a control. It works because well-behaved crawlers choose to honor it, and that consensus held for thirty years because search engines had an incentive to be well-behaved — they needed publishers to not block them.</p>
<p>AI training crawlers have no such incentive. The content is valuable to them and the traffic they return is zero or nearly so. Multiple studies found training crawlers ignoring <code>robots.txt</code>, rotating user agents, and using residential proxy pools. Once a norm has no enforcement and no incentive, it stops being a norm.</p>
<p>Cloudflare's move replaces a request with a control. That is the actual innovation and it required no new technology at all — just someone at a chokepoint deciding to enforce.</p>
<h2 id="the-case-against">the case against<a class="anchor" href="#the-case-against" aria-label="link to this section">#</a></h2>
<p>Concentrating the ability to gate the web at one CDN is not obviously good, even if this specific use of the power is popular.</p>
<p>The precedent is: an infrastructure company can unilaterally change how content is accessed for a large fraction of the internet, and the mechanism generalizes to things other than AI crawlers. Cloudflare has been thoughtful and has taken public positions on not being an arbiter, and the concentration is still real.</p>
<p>There is also a smaller-player problem. Large AI companies can negotiate licensing deals directly. Researchers, startups, the Internet Archive, and academic crawlers cannot. A tollbooth is regressive: it is a rounding error for the incumbents and a barrier for everyone else. Cloudflare has carve-outs for some of these and the carve-outs are discretionary, which is the point.</p>
<h2 id="for-developers">for developers<a class="anchor" href="#for-developers" aria-label="link to this section">#</a></h2>
<p>Two practical items.</p>
<p><strong>If you run a site</strong>, decide deliberately. Blocking training crawlers is now the default; that may not be what you want. Documentation sites in particular may prefer to be in the training data, because being the thing the model knows about is worth more than the pageview you lost.</p>
<p><strong>If you build anything that crawls</strong>, expect 402s and expect your user agent to matter. Identify honestly, respect the directives, and set up billing if you need paid access. The era of scraping quietly is ending, and the enforcement is technical now rather than legal.</p>
<h2 id="the-bigger-shift">the bigger shift<a class="anchor" href="#the-bigger-shift" aria-label="link to this section">#</a></h2>
<p>The web's economic model was: publish freely, get traffic, monetize traffic. AI answers break the second step, which breaks the third, which will eventually break the first.</p>
<p>Pay-per-crawl is one proposed replacement. Licensing deals are another. Neither is obviously going to work at the scale of the actual web, where most content is made by people with no ability to negotiate anything.</p>
<p>What replaces it is genuinely unresolved, and 2025 is the year everybody stopped pretending otherwise.</p>]]></content:encoded></item><item><title>Google Cloud Next: Ironwood, A2A, and a very cheap Gemini</title><link>https://readme.news/google-cloud-next-ironwood-a2a-and-a-very-cheap-gemini/</link><guid isPermaLink="true">https://readme.news/google-cloud-next-ironwood-a2a-and-a-very-cheap-gemini/</guid><pubDate>Thu, 10 Apr 2025 09:00:00 +0000</pubDate><description>A seventh-gen TPU aimed squarely at inference, plus a protocol for agents talking to other agents.</description><content:encoded><![CDATA[<p>Google Cloud Next happened this week. Three things from it will still matter in a year.</p>
<h2 id="ironwood">Ironwood<a class="anchor" href="#ironwood" aria-label="link to this section">#</a></h2>
<p>The seventh-generation TPU, and the first Google has explicitly positioned for inference rather than training. Large HBM capacity per chip, high interconnect bandwidth, deployable in pods of thousands.</p>
<p>The strategic point is not the specs. It is that Google is the only hyperscaler with a mature, decade-old, production-proven alternative to buying Nvidia. AWS has Trainium and Inferentia, which are real but younger. Microsoft has Maia, which is very young. Google has been running TPUs in production since 2015 and has an entire compiler stack (XLA) and framework story (JAX) built around them.</p>
<p>That translates directly into pricing freedom. When Google prices Gemini aggressively, they are not eating a margin on someone else's silicon.</p>
<h2 id="a2a">A2A<a class="anchor" href="#a2a" aria-label="link to this section">#</a></h2>
<p>Agent2Agent: a protocol for agents from different vendors to discover each other and collaborate on tasks. Announced with a long list of partner companies.</p>
<p>The mental model is that MCP connects an agent to <em>tools and data</em>, while A2A connects an agent to <em>other agents</em>. An agent publishes an "agent card" at a well-known URL describing its capabilities; other agents discover it and delegate tasks over a JSON-RPC-ish protocol with support for long-running work and streaming updates.</p>
<p>My honest read: the problem is real, the timing is early, and the number of partner logos on the announcement slide is inversely correlated with how much production usage a protocol has on day one. Multi-agent systems in 2025 mostly do not work well enough to need a standard for interoperating — the hard part is that agents fail in the middle of tasks, not that they cannot find each other.</p>
<p>But standards need to exist before they are needed, and having the conversation now beats having it in 2027 with four incompatible implementations.</p>
<h2 id="gemini-25-flash">Gemini 2.5 Flash<a class="anchor" href="#gemini-25-flash" aria-label="link to this section">#</a></h2>
<p>The cost-efficient reasoning model, with a controllable <a class="xref" href="/gemini-25-goes-generally-available-with-a-thinking-dial/" title="Gemini 2.5 goes generally available with a thinking dial">thinking budget</a>. You set how many tokens the model may spend reasoning, including zero.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">config = types.GenerateContentConfig(
    thinking_config=types.ThinkingConfig(thinking_budget=1024)
)</code></pre></div>
<p>This is the right API shape and I expect everyone to converge on it. The decision of how much to think is workload-specific, it is the primary cost lever in a reasoning model, and hiding it inside the model's own judgment takes control away from the person paying the bill.</p>
<p>Set it to zero for classification. Set it high for planning. Measure the quality difference on your own task, because the curve is steeper for some tasks than others and there is no general answer.</p>
<h2 id="the-rest">the rest<a class="anchor" href="#the-rest" aria-label="link to this section">#</a></h2>
<p>A great deal of "agentic" branding applied to existing products, an agent development kit, and a marketplace. Cloud vendor conferences have a genre and this one was firmly in it.</p>
<p>The signal-to-slide ratio was better than most.</p>]]></content:encoded></item><item><title>GTC 2025: a roadmap to 2027 and a warning about power</title><link>https://readme.news/gtc-2025-a-roadmap-to-2027-and-a-warning-about-power/</link><guid isPermaLink="true">https://readme.news/gtc-2025-a-roadmap-to-2027-and-a-warning-about-power/</guid><pubDate>Wed, 19 Mar 2025 09:00:00 +0000</pubDate><description>Blackwell Ultra this year, Rubin next, Feynman after. Nvidia is now publishing a schedule like a foundry.</description><content:encoded><![CDATA[<p>Jensen Huang did two hours at the San Jose arena and the substance came down to one slide: a named product cadence stretching to 2027.</p>
<ul><li><strong>Blackwell Ultra (GB300)</strong> — second half of 2025, more HBM per package.</li><li><strong>Vera Rubin</strong> — 2026, new CPU (Vera) and new GPU (Rubin), HBM4.</li><li><strong>Rubin Ultra</strong> — 2027, a much larger rack-scale configuration.</li><li><strong>Feynman</strong> — 2028, named, not detailed.</li></ul>
<p>Publishing a multi-year roadmap at this granularity is a foundry move, not a product-company move. The audience is not developers. It is the utilities, the construction firms, the HBM suppliers, and the CFOs who need to plan capital around it.</p>
<h2 id="the-number-that-should-worry-you">the number that should worry you<a class="anchor" href="#the-number-that-should-worry-you" aria-label="link to this section">#</a></h2>
<p>The power figures for the rack-scale systems are the part of this keynote that will still matter in five years. Current NVL72 racks draw on the order of 120 kW. The Rubin Ultra generation is being discussed in the hundreds of kilowatts per rack.</p>
<p>A traditional enterprise datacenter rack is provisioned for 5 to 15 kW. Air cooling tops out somewhere around 40 kW with heroic effort. Everything past that is liquid, and everything past about 150 kW is liquid plus a fundamentally different power distribution architecture — which is why the roadmap includes 800 VDC distribution.</p>
<p>Translation: existing datacenter shells are largely unusable for this. The buildout is not "install new servers," it is "build new buildings near new substations." That is a multi-year, capital-intensive, permit-bound process, and it is the actual rate limiter on AI capacity through the rest of the decade.</p>
<h2 id="the-software-announcements">the software announcements<a class="anchor" href="#the-software-announcements" aria-label="link to this section">#</a></h2>
<p><strong>Dynamo</strong>, an open-source inference serving framework, is the developer-relevant release. It handles disaggregated serving — running the prefill phase and the decode phase on different hardware pools, because they have completely different compute and memory characteristics. Prefill is compute-bound and parallel; decode is memory-bandwidth-bound and sequential. Running them on the same homogeneous pool wastes a lot of silicon.</p>
<p>If you operate inference at any scale, disaggregation is probably the largest single efficiency win available to you right now, and having a maintained open implementation lowers the bar considerably.</p>
<p><strong>NIM microservices</strong> continue to be Nvidia's attempt to own the deployment layer. Reasonable, containerized, and a <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a> vector you should evaluate with clear eyes.</p>
<h2 id="the-strategic-read">the strategic read<a class="anchor" href="#the-strategic-read" aria-label="link to this section">#</a></h2>
<p>Nvidia's actual product is no longer a chip. It is a rack, plus the networking between racks, plus the software stack on top. Selling a GPU means competing with AMD and with every hyperscaler's internal silicon team. Selling an integrated rack-scale system with a co-designed interconnect means competing with nobody, because nobody else has the interconnect.</p>
<p>That is the moat, it is being widened deliberately, and the roadmap slide is the announcement of that strategy rather than a product list.</p>]]></content:encoded></item><item><title>Stargate is a $500 billion bet on a bottleneck</title><link>https://readme.news/stargate-is-a-500-billion-bet-on-a-bottleneck/</link><guid isPermaLink="true">https://readme.news/stargate-is-a-500-billion-bet-on-a-bottleneck/</guid><pubDate>Thu, 23 Jan 2025 09:00:00 +0000</pubDate><description>OpenAI, Oracle, SoftBank and MGX announce an infrastructure vehicle. The real constraint isn&#x27;t chips — it&#x27;s power and steel.</description><content:encoded><![CDATA[<p>The Stargate Project was announced from the White House on Tuesday: a joint venture between OpenAI, Oracle, SoftBank and MGX, with a stated intent to deploy $500 billion into US AI infrastructure over four years, $100 billion of it "immediately." Construction is already underway in Abilene, Texas.</p>
<p>Two things are true at once and you need to hold both.</p>
<h2 id="the-number-is-not-a-number">the number is not a number<a class="anchor" href="#the-number-is-not-a-number" aria-label="link to this section">#</a></h2>
<p>$500 billion is not committed capital. It is an aspiration with a press release attached. The initial equity is a fraction of it, the rest is contingent on debt markets, vendor financing, offtake agreements, and continued demand that nobody can underwrite four years out. Announcements at this scale are partly a coordination device — you say the number so suppliers, utilities and lenders plan around it.</p>
<p>Elon Musk publicly said the money isn't there. He is a hostile witness with obvious motives, and he is also not obviously wrong about the funding gap.</p>
<h2 id="the-constraint-is-not-gpus">the constraint is not GPUs<a class="anchor" href="#the-constraint-is-not-gpus" aria-label="link to this section">#</a></h2>
<p>Here is the part developers should actually internalize, because it changes how you should think about compute pricing for the next several years.</p>
<p>The binding constraint on AI datacenter buildout is no longer semiconductor supply. It is:</p>
<ul><li><strong>Power.</strong> A gigawatt-class campus needs an interconnection agreement, and US interconnection <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> are measured in years. This is why the sites are going where they are going — West Texas has wind, gas, and a grid operator that can move faster than most.</li><li><strong>Transformers and switchgear.</strong> Lead times for high-voltage transformers ran past two years. You cannot software your way around a transformer.</li><li><strong>Cooling and water rights.</strong> Liquid cooling at rack densities north of 100 kW is a plumbing problem, and plumbing has a supply chain.</li><li><strong>Electricians.</strong> There is a genuine national shortage of people qualified to terminate high-voltage cable, and you cannot fine-tune one.</li></ul>
<p>The Abilene site is the tell. They started building before the financing was finalized because the long pole is not money, it is queue position.</p>
<h2 id="what-it-means-for-your-bill">what it means for your bill<a class="anchor" href="#what-it-means-for-your-bill" aria-label="link to this section">#</a></h2>
<p>If you are trying to model API costs, the useful frame is: capacity comes online in step functions, tied to substation energization dates, not to Nvidia's quarterly shipments. Prices per token will keep falling on a per-capability basis because of model efficiency gains, not because compute is getting cheap. Compute is not getting cheap. Compute is getting <em>more available</em>, in lumps, eighteen months after somebody signs a power purchase agreement.</p>
<h2 id="the-second-order-thing">the second-order thing<a class="anchor" href="#the-second-order-thing" aria-label="link to this section">#</a></h2>
<p>Every one of these campuses is a bet that inference demand keeps compounding. If model efficiency improves faster than demand — if a 2027 model gets today's quality at a tenth of the FLOPs — a lot of this concrete is stranded.</p>
<p>That is not a reason to think the buildout is stupid. It is a reason to notice that "AI capex" and "AI capability" are two different bets, and only one of them is being made by the people writing these checks.</p>]]></content:encoded></item>
</channel>
</rss>
