<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — tooling</title>
<link>https://readme.news/tags/tooling/</link>
<atom:link href="https://readme.news/tags/tooling/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged tooling.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>Shell scripts that outlive you</title><link>https://readme.news/shell-scripts-that-outlive-you/</link><guid isPermaLink="true">https://readme.news/shell-scripts-that-outlive-you/</guid><pubDate>Sat, 22 Aug 2026 09:00:00 +0000</pubDate><description>Six lines at the top of a bash script are the difference between a tool and a trap.</description><content:encoded><![CDATA[<p>Every codebase has a <code>scripts/</code> directory. Most of the files in it were written in ten minutes, work correctly on exactly one machine, and fail in ways that produce no error and no output.</p>
<p>Shell is a fine language for gluing programs together. It is a terrible language for doing it <em>safely</em> unless you tell it to be, and telling it to be takes six lines.</p>
<h2 id="the-preamble">the preamble<a class="anchor" href="#the-preamble" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">#!/usr/bin/env bash
set -euo pipefail
IFS=$'\n\t'</code></pre></div>
<p>What each one prevents:</p>
<p><strong><code>set -e</code></strong> — exit on any command that fails. Without it, a script continues merrily after <code>cd /nonexistent</code> and then runs the rest of its commands in the wrong directory. This is how a cleanup script deletes the wrong thing.</p>
<p><strong><code>set -u</code></strong> — error on an undefined variable. Without it, <code>rm -rf "$BUILD_DIR/"</code> with an unset <code>BUILD_DIR</code> expands to <code>rm -rf /</code>. This has happened to real people, at real companies, more than once.</p>
<p><strong><code>set -o pipefail</code></strong> — a pipeline fails if <em>any</em> stage fails, not just the last one. Without it, <code>curl bad-url | jq .</code> succeeds, because <code>jq</code> was happy with the empty input.</p>
<p><strong><code>IFS=$'\n\t'</code></strong> — stop splitting on spaces. This is what makes filenames with spaces stop being a source of bugs.</p>
<p>Four lines. They convert an entire category of silent wrong behaviour into loud failure.</p>
<h2 id="the-next-four-things">the next four things<a class="anchor" href="#the-next-four-things" aria-label="link to this section">#</a></h2>
<p><strong>Quote every expansion.</strong> <code>"$var"</code>, not <code>$var</code>. Always, including inside <code>[[ ]]</code> where it usually does not matter, because "usually" is not a rule anyone remembers correctly.</p>
<p><strong>Use <code>"${var:?message}"</code> for required inputs.</strong></p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">: "${DATABASE_URL:?DATABASE_URL is required}"</code></pre></div>
<p>One line, fails immediately with a useful message rather than three steps later with a confusing one.</p>
<p><strong>Make it idempotent.</strong> A script that is safe to run twice is a script that is safe to run at all. <code>mkdir -p</code>, <code>rm -f</code>, check-before-create. The second run is the one that happens during an incident when nobody is sure whether the first one worked.</p>
<p><strong>Clean up with <code>trap</code>.</strong></p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT</code></pre></div>
<p>Now the temporary directory is removed whether the script succeeds, fails, or is interrupted.</p>
<h2 id="the-usability-part">the usability part<a class="anchor" href="#the-usability-part" aria-label="link to this section">#</a></h2>
<p>A script that other people run needs the same courtesy as any other <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">usage() {
  cat &lt;&lt;'EOF'
Usage: deploy.sh ENVIRONMENT [--dry-run]

  ENVIRONMENT   staging | production
  --dry-run     print what would happen, change nothing

Requires: awscli &gt;= 2, jq. Reads DEPLOY_ROLE from the environment.
EOF
}
[[ $# -eq 0 || "${1:-}" == "-h" ]] &amp;&amp; { usage; exit 0; }</code></pre></div>
<p><strong>Add a dry-run mode to anything destructive.</strong> It costs one conditional and it is the difference between a script people trust and a script people read three times before running.</p>
<p><strong>Echo what you are about to do.</strong> <code>set -x</code> is the crude version and it is better than silence. A script that prints "deleting 4 objects from s3://bucket/prefix/" before doing it lets a human catch the mistake.</p>
<h2 id="when-to-stop-using-shell">when to stop using shell<a class="anchor" href="#when-to-stop-using-shell" aria-label="link to this section">#</a></h2>
<p>Shell is right for: calling other programs in sequence, moving files, gluing a pipeline together. Under about a hundred lines.</p>
<p>Switch to a real language when you need:</p>
<ul><li><strong>Data structures.</strong> Bash arrays are a trap and associative arrays are worse.</li><li><strong>Any arithmetic beyond counting.</strong></li><li><strong>Error handling with recovery</strong>, rather than exit-on-failure.</li><li><strong>Parsing anything structured.</strong> If you are pulling JSON apart with <code>sed</code>, stop.</li><li><strong>Tests.</strong> You can test shell, and almost nobody does, which tells you something.</li></ul>
<p>Python or Go for anything past that line. The rewrite is an hour and it pays back the first time somebody has to change it.</p>
<h2 id="the-check-that-costs-nothing">the check that costs nothing<a class="anchor" href="#the-check-that-costs-nothing" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">shellcheck scripts/*.sh</code></pre></div>
<p>Run it in CI. It catches unquoted expansions, useless <code>cat</code>, subshell variable scoping, and roughly a dozen other things that produce silent wrong behaviour.</p>
<p>Every shell script I have ever run it against had at least one real finding. It takes five minutes to add and it is the highest-value linting available for the least-linted language in most repositories.</p>]]></content:encoded></item><item><title>Why your container image is 1.4 gigabytes</title><link>https://readme.news/why-your-container-image-is-14-gigabytes/</link><guid isPermaLink="true">https://readme.news/why-your-container-image-is-14-gigabytes/</guid><pubDate>Fri, 10 Jul 2026 09:00:00 +0000</pubDate><description>It should be forty megabytes. Here is where the rest of it came from and how to get it back.</description><content:encoded><![CDATA[<p>A container image for a compiled service should be tens of megabytes. For an interpreted one, low hundreds. If yours is over a gigabyte, something specific went wrong and it is usually one of six things.</p>
<p>Size matters for real reasons: pull time on cold start, registry cost, deployment speed when you are scaling out under load, and attack surface — every package in the image is something that can have a CVE you have to answer for.</p>
<h2 id="the-six-causes">the six causes<a class="anchor" href="#the-six-causes" aria-label="link to this section">#</a></h2>
<p><strong>1. You shipped the build toolchain.</strong></p>
<p>The compiler, the headers, the package manager cache, the source tree, the test fixtures. All of it needed to build, none of it needed to run.</p>
<p>Multi-stage builds fix this completely:</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile">FROM golang:1.24 AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 go build -ldflags="-s -w" -o /app ./cmd/server

FROM gcr.io/distroless/static-debian12
COPY --from=build /app /app
ENTRYPOINT ["/app"]</code></pre></div>
<p>Final image: the binary, plus CA certificates and timezone data. Tens of megabytes.</p>
<p><strong>2. You started from a full distribution image.</strong></p>
<p><code>FROM ubuntu</code> is roughly 80 MB before you install anything, and it includes a package manager, a shell, and a hundred utilities you will never invoke.</p>
<p>The ladder, from largest to smallest:</p>
<ul><li>Full distribution — 80 MB+</li><li><code>-slim</code> variants — 30–80 MB</li><li>Alpine — 5–10 MB, with musl libc, which will occasionally surprise you</li><li>Distroless — just the runtime, no shell, no package manager</li><li><code>scratch</code> — nothing at all, for static binaries</li></ul>
<p><strong>The Alpine caveat</strong>, since it bites people: musl's allocator and DNS resolver behave differently from glibc's. Python performance in particular can be significantly worse, and some binary wheels do not exist for musl. Test rather than assume.</p>
<p><strong>3. Your layers are ordered wrong.</strong></p>
<p>Each instruction creates a layer. A layer is invalidated when it or anything before it changes.</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile"># bad — any source change reinstalls every dependency
COPY . .
RUN npm ci

# good — dependencies are cached until the lockfile changes
COPY package.json package-lock.json ./
RUN npm ci
COPY . .</code></pre></div>
<p>This does not shrink the final image but it dramatically speeds up builds, which is usually what people actually care about.</p>
<p><strong>4. You deleted things in a later layer.</strong></p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile">RUN apt-get install -y build-essential   # layer 1: +400 MB
RUN apt-get remove -y build-essential    # layer 2: marks deleted, image unchanged</code></pre></div>
<p>Layers are additive. Deleting a file in a later layer hides it and does not remove it. The bytes are still in the image and still transferred on pull.</p>
<p>Everything must happen in one <code>RUN</code>:</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile">RUN apt-get update \
 &amp;&amp; apt-get install -y --no-install-recommends build-essential \
 &amp;&amp; make \
 &amp;&amp; apt-get purge -y build-essential \
 &amp;&amp; apt-get autoremove -y \
 &amp;&amp; rm -rf /var/lib/apt/lists/*</code></pre></div>
<p>Better: use a multi-stage build and do not install the toolchain in the final image at all.</p>
<p><strong>5. You have no <code>.dockerignore</code>.</strong></p>
<p><code>COPY . .</code> copies <code>.git</code>, <code>node_modules</code>, build artifacts, test fixtures, and your local <code>.env</code>.</p>
<div class="code"><pre><code>.git
node_modules
dist
*.log
.env*
**/__pycache__
coverage</code></pre></div>
<p>The <code>.git</code> directory alone is frequently hundreds of megabytes on a mature repository, and it is in a lot of images.</p>
<p><strong>6. Your dependencies are enormous.</strong></p>
<p>Sometimes it is genuinely the dependencies — machine learning stacks with CUDA libraries are legitimately multiple gigabytes.</p>
<p>Check whether you need the GPU variant. <code>torch</code> with CUDA is roughly 2.5 GB; the CPU build is a fraction of that. If you are serving on CPU, you are shipping GPU libraries for nothing.</p>
<h2 id="finding-out-where-it-went">finding out where it went<a class="anchor" href="#finding-out-where-it-went" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">docker history --no-trunc &lt;image&gt;       # size per layer</code></pre></div>
<p>Or use a layer inspection tool that shows you which files are in which layer and how much space is wasted. Ten minutes with one of those tells you exactly what to fix.</p>
<h2 id="the-security-dimension">the security dimension<a class="anchor" href="#the-security-dimension" aria-label="link to this section">#</a></h2>
<p>Every package in the image is potential CVE surface, and your scanner will report all of them regardless of whether the code is reachable.</p>
<p>A distroless image has almost nothing to report, which means the reports you do get are signal rather than noise. That is worth more than the size reduction — a vulnerability report with three entries gets read; one with four hundred does not.</p>
<p>The trade-off: no shell means you cannot <code>docker exec</code> in to debug. Use ephemeral debug containers that attach to the running pod's namespaces instead, which is a better practice anyway because it means your production image is not a debugging toolkit.</p>
<h2 id="the-target">the target<a class="anchor" href="#the-target" aria-label="link to this section">#</a></h2>
<ul><li>Compiled language, static binary: <strong>under 30 MB.</strong></li><li>Interpreted with dependencies: <strong>under 200 MB.</strong></li><li>Anything over a gigabyte without a machine learning stack: something is wrong and it is one of the six above.</li></ul>]]></content:encoded></item><item><title>Static analysis is finally worth the false positives</title><link>https://readme.news/static-analysis-is-finally-worth-the-false-positives/</link><guid isPermaLink="true">https://readme.news/static-analysis-is-finally-worth-the-false-positives/</guid><pubDate>Fri, 05 Jun 2026 09:00:00 +0000</pubDate><description>The tools got dramatically better while everyone was ignoring them because of a bad experience in 2015.</description><content:encoded><![CDATA[<p>A lot of engineers formed their opinion of static analysis from a tool that produced four thousand warnings on first run, 95% of which were noise, and got disabled within a month.</p>
<p>That was an accurate assessment of the tools at the time. The tools are substantially different now and the assessment has not updated.</p>
<h2 id="what-changed">what changed<a class="anchor" href="#what-changed" aria-label="link to this section">#</a></h2>
<p><strong>Flow-sensitive analysis became standard.</strong> Older linters matched patterns in the syntax tree. Modern analyzers track values through the control flow graph, which means they can tell that a variable was checked for null on line 12 and therefore is not null on line 40. That single capability eliminates the largest source of false positives.</p>
<p><strong>Language servers made it interactive.</strong> A warning in your editor while you type is a different product from a report generated in CI. You fix it in context, in seconds, instead of triaging a list a week later.</p>
<p><strong><a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">Type systems</a> absorbed much of it.</strong> A lot of what static analysis used to catch is now caught by the compiler in a typed language, for free, with no false positives at all.</p>
<p><strong>The defaults got sane.</strong> Modern tools ship with a curated recommended set rather than everything enabled. <code>clippy</code>, <code>ruff</code>, <code>biome</code>, <code>staticcheck</code> and their peers are opinionated about what is worth reporting.</p>
<p><strong>They got fast.</strong> Analyzers written in compiled languages run over a large codebase in seconds. Speed matters more than people credit — a check that takes two minutes gets run in CI, and a check that takes two seconds gets run on every save, which is where it actually changes behavior.</p>
<h2 id="what-to-actually-run">what to actually run<a class="anchor" href="#what-to-actually-run" aria-label="link to this section">#</a></h2>
<p><strong>A fast linter with a good default set</strong>, on save, in the editor. <code>ruff</code> for Python, <code>clippy</code> for Rust, <code>biome</code> or <code>eslint</code> for JavaScript, <code>staticcheck</code> for Go.</p>
<p><strong>A typechecker in strict mode</strong>, in CI. This is the highest-value item on the list for a gradually-typed language, and the non-strict modes permit exactly the holes that make the guarantees unreliable.</p>
<p><strong>A security-focused analyzer</strong> if you handle untrusted input. Taint tracking — does data from a request reach a SQL query, a shell command, or a template without sanitization — is the specific capability worth having, and it is a genuinely different analysis from ordinary linting.</p>
<p><strong>A dependency scanner</strong>, ranked by reachability if your tooling supports it. A critical CVE in code you never call is lower priority than a medium in your request path, and a scanner that cannot tell you which is which produces a queue nobody reads.</p>
<h2 id="the-adoption-sequence">the adoption sequence<a class="anchor" href="#the-adoption-sequence" aria-label="link to this section">#</a></h2>
<p>Turning on a full rule set against an existing codebase produces thousands of warnings and gets the tool disabled. The sequence that works:</p>
<p><strong>1. Run it in report-only mode.</strong> Get the number. Do not fix anything yet.</p>
<p><strong>2. Enable a small subset that has near-zero false positives</strong> and fix those. Usually: unused variables, unreachable code, obviously wrong comparisons, missing awaits. Twenty rules, not four hundred.</p>
<p><strong>3. Make it blocking for new and changed code only.</strong> Most tools support this, either natively or through a diff-aware wrapper. This is the key move — the existing violations do not block anyone, and the codebase stops getting worse immediately.</p>
<p><strong>4. Burn down the backlog opportunistically.</strong> When you touch a file, fix its warnings. No cleanup sprint, no dedicated project.</p>
<p><strong>5. Add rules gradually</strong>, one at a time, each with a burn-down.</p>
<p>Steps three and four are where most adoptions succeed or fail. A tool that blocks the whole team on a pre-existing backlog gets turned off; one that only blocks new violations is uncontroversial.</p>
<h2 id="the-rules-worth-arguing-about">the rules worth arguing about<a class="anchor" href="#the-rules-worth-arguing-about" aria-label="link to this section">#</a></h2>
<p>Some checks are genuinely contested and you should decide deliberately rather than accepting the default:</p>
<p><strong>Cyclomatic complexity limits.</strong> Sometimes a function is legitimately complex because the domain is. A hard limit produces artificially split functions that are harder to read, not easier.</p>
<p><strong>Line length.</strong> Real disagreement, formatter should handle it, not worth a rule.</p>
<p><strong>Naming conventions.</strong> Worth enforcing, and pick your convention rather than the tool's default if they differ.</p>
<p><strong>Anything with more than a few percent false positives.</strong> A rule that is wrong one time in ten trains people to ignore it, and that habit generalizes to the rules that are right.</p>
<h2 id="the-honest-limits">the honest limits<a class="anchor" href="#the-honest-limits" aria-label="link to this section">#</a></h2>
<p>Static analysis finds a specific class of bug: local, syntactic, pattern-matchable. It does not find logic errors, wrong business rules, race conditions in most cases, or performance problems.</p>
<p>It is not a substitute for tests, review, or thought. It is a way to spend zero human attention on the errors that do not require human attention, which frees attention for the ones that do.</p>
<p>That framing — attention allocation rather than bug finding — is the one that makes it worth the setup.</p>
<h2 id="the-new-reason-it-matters">the new reason it matters<a class="anchor" href="#the-new-reason-it-matters" aria-label="link to this section">#</a></h2>
<p>Machine-generated code has a characteristic error profile: plausible, syntactically valid, and wrong in specific recurring ways. Unchecked errors, missing awaits, resource leaks, off-by-one in boundary conditions.</p>
<p>Those are exactly the errors static analysis is good at. Running a strict analyzer over generated code is verification you get for free, and it is one of the few places where the verification bottleneck has an automated answer.</p>
<p>Turn it on.</p>]]></content:encoded></item><item><title>Platform teams that don't get resented</title><link>https://readme.news/platform-teams-that-dont-get-resented/</link><guid isPermaLink="true">https://readme.news/platform-teams-that-dont-get-resented/</guid><pubDate>Fri, 22 May 2026 09:00:00 +0000</pubDate><description>Internal platforms fail for predictable reasons. The successful ones share four properties.</description><content:encoded><![CDATA[<p>Most internal platform teams end up resented by the engineers they serve. The pattern is consistent enough that the causes are identifiable.</p>
<h2 id="the-failure-pattern">the failure pattern<a class="anchor" href="#the-failure-pattern" aria-label="link to this section">#</a></h2>
<ol><li>Platform team forms to reduce duplicated infrastructure work.</li><li>They build an abstraction over the cloud provider.</li><li>The abstraction covers 80% of cases well.</li><li>The remaining 20% is impossible, and the escape hatch is either absent or punished.</li><li>Product teams work around the platform.</li><li>Platform team responds by mandating the platform.</li><li>Everyone is unhappy and the platform is now a tax.</li></ol>
<p>Every step follows from the previous one. The root is step four.</p>
<h2 id="the-four-properties-of-platforms-that-work">the four properties of platforms that work<a class="anchor" href="#the-four-properties-of-platforms-that-work" aria-label="link to this section">#</a></h2>
<p><strong>1. An escape hatch that is not punished.</strong></p>
<p>The platform covers the common case. It cannot cover every case, and pretending otherwise is what breaks trust.</p>
<p>There must be a supported path for "I need something the platform does not do," and taking that path must not require an exception process, a meeting, or an apologetic Slack message.</p>
<p>The best platforms make the escape hatch cheap and then compete on being better than it. The worst make it forbidden, which does not eliminate the need — it drives it underground.</p>
<p><strong>2. Adoption is voluntary, at least at first.</strong></p>
<p>A platform that teams choose is a platform that is good. A platform teams are required to use never gets the feedback that would make it good, because the feedback mechanism — people leaving — has been disabled.</p>
<p>If you cannot get voluntary adoption, that is information. Mandating it does not fix the underlying problem; it hides it and converts a product problem into a political one.</p>
<p>Mandate later, when it is genuinely better, and the mandate will be uncontroversial because everyone already uses it.</p>
<p><strong>3. The abstraction leaks deliberately, not accidentally.</strong></p>
<p>Every abstraction leaks. The question is whether you planned for it.</p>
<p>A good platform lets you drop a level when you need to: use the paved path for the deployment, and reach the underlying resource directly when you need something specific. A bad one hides the underlying system entirely, so that when it fails you cannot debug it and neither can the platform team, because now there are two systems to understand.</p>
<p><strong>Concretely:</strong> if your platform generates infrastructure configuration, let people see it. If it wraps a cloud API, let people access the underlying resource. If it runs their container, give them the logs from the actual runtime, not a filtered view.</p>
<p><strong>4. The platform team is measured on adoption and satisfaction, not on compliance.</strong></p>
<p>If the platform team's metric is "percentage of services on the platform," they will optimize for mandating it.</p>
<p>If the metric is "would you use this if you had a choice," they will optimize for making it good.</p>
<p>Ask that question quarterly, anonymously, and publish the answer.</p>
<h2 id="the-specific-things-that-generate-resentment">the specific things that generate resentment<a class="anchor" href="#the-specific-things-that-generate-resentment" aria-label="link to this section">#</a></h2>
<p><strong>Slow escape.</strong> A team needs something the platform does not support. The answer is "file a request, we will look at it next quarter." Their deadline is Friday.</p>
<p><strong>Breaking changes without migration paths.</strong> The platform is infrastructure. Break it and every team stops. Platform teams frequently hold themselves to a lower compatibility standard than they would accept from a vendor.</p>
<p><strong>Opaque failures.</strong> The deploy failed. The error is a platform-internal message. The product engineer cannot debug it and must escalate, which means waiting.</p>
<p><strong>Being a gate rather than a service.</strong> A platform that must approve things is a bureaucracy. A platform that makes the right thing easy is infrastructure.</p>
<p><strong>Solving the platform team's problems.</strong> Standardization is valuable to the platform team and is not automatically valuable to product teams. If the pitch for a migration is "this makes our lives easier," expect a cool reception.</p>
<h2 id="the-framing-that-works">the framing that works<a class="anchor" href="#the-framing-that-works" aria-label="link to this section">#</a></h2>
<p><strong>You are building a product. Your users are engineers. They have alternatives.</strong></p>
<p>That framing produces the right behaviors automatically: user research before building, documentation that assumes nothing, onboarding that works, support that responds, and a roadmap driven by what users need rather than by architectural preference.</p>
<p>The platform teams I have seen work best behave exactly like a startup selling to a skeptical market, and they say so out loud.</p>
<p>The ones that fail behave like an internal standards body, and they are usually correct about the standards and wrong about how to get them adopted.</p>
<h2 id="the-measurement-that-matters">the measurement that matters<a class="anchor" href="#the-measurement-that-matters" aria-label="link to this section">#</a></h2>
<p>Time from "a new engineer joins" to "their code is running in production."</p>
<p>That single number captures most of what a platform is for, it is measurable, and it is the thing product teams actually care about. If it is going down, the platform is working, regardless of what the adoption <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> says.</p>]]></content:encoded></item><item><title>I/O 2026 and the assistant that lives in everything</title><link>https://readme.news/io-2026-and-the-assistant-that-lives-in-everything/</link><guid isPermaLink="true">https://readme.news/io-2026-and-the-assistant-that-lives-in-everything/</guid><pubDate>Fri, 08 May 2026 09:00:00 +0000</pubDate><description>More model, more surfaces, and a search product that keeps changing what the web is for.</description><content:encoded><![CDATA[<p>Google's developer conference happened this week and the shape is consistent with where the company has been heading since 2024: a capable model, deployed everywhere they already have users, priced aggressively because they own the silicon.</p>
<h2 id="the-distribution-advantage-compounding">the distribution advantage, compounding<a class="anchor" href="#the-distribution-advantage-compounding" aria-label="link to this section">#</a></h2>
<p>The thing no competitor can replicate is that Google can ship a capability into products that billions of people already open daily, on launch day.</p>
<p>That is worth more than a benchmark lead and it is becoming more visible each year. A model that is marginally better but reaches users through a signup flow loses to a model that is marginally worse and is already in the search box.</p>
<p>The strategic implication for everyone else — including the other frontier labs — is that raw capability is not the competition anymore. Distribution, price, and integration are.</p>
<h2 id="the-search-question-again">the search question, again<a class="anchor" href="#the-search-question-again" aria-label="link to this section">#</a></h2>
<p>Every year this conference makes the same thing more true: informational queries are increasingly answered on the results page rather than by sending someone to a site.</p>
<p>For anyone who publishes on the web, the consequences are now well past theoretical:</p>
<ul><li><strong>Referral traffic to informational content keeps falling.</strong> This is measurable and it is not recovering.</li><li><strong>Your documentation is being summarized by a system you do not control</strong>, and users are acting on the summary.</li><li><strong>The <a class="xref" href="/why-your-tests-are-slow/" title="Why your tests are slow">feedback loop</a> is broken.</strong> You cannot see what people asked, what answer they got, or whether it was right.</li></ul>
<p>I do not have a satisfying answer. The mitigations available to an individual project are marginal: write documentation that is hard to summarize badly, keep a machine-readable <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a>, make <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a> self-explanatory so they do not require a search at all.</p>
<p>The structural problem — that the economic model funding web content is being removed without a replacement — is not solvable by any individual publisher, and the people who could solve it have no incentive to.</p>
<h2 id="the-developer-surface">the developer surface<a class="anchor" href="#the-developer-surface" aria-label="link to this section">#</a></h2>
<p>The genuinely useful announcements, as always, are the boring ones:</p>
<p><strong>Model pricing and the cheap tier.</strong> The cost-per-capability at the small end continues falling, and Google's TPU position means they can price below what competitors renting accelerators can match. If you run high-volume inference, the arithmetic is worth redoing quarterly.</p>
<p><strong>Longer context, better retention.</strong> Incremental and real. The practical question is always the degradation curve, not the maximum, and that continues to improve.</p>
<p><strong>Agent tooling in the cloud.</strong> Same category everyone is building: runtimes, memory, identity, observability. Evaluate against a real problem rather than a demo.</p>
<h2 id="the-thing-to-actually-do">the thing to actually do<a class="anchor" href="#the-thing-to-actually-do" aria-label="link to this section">#</a></h2>
<p>The recommendation has not changed in two years and I will keep repeating it because it keeps being right:</p>
<p><strong>Have an eval set in your repository.</strong> Fifty examples from your real domain, with expected outputs, run against every candidate model.</p>
<p>Every conference like this produces a new model that is claimed to be better. With an eval harness, evaluating that claim for your workload takes an hour. Without one, it takes a week of impressions and you will get it wrong.</p>
<p>This is a day of setup that pays back on every model release, forever, and the number of teams that have done it remains surprisingly small.</p>
<h2 id="the-honest-summary">the honest summary<a class="anchor" href="#the-honest-summary" aria-label="link to this section">#</a></h2>
<p>A very good model, deployed extremely well, in a company with structural advantages that are getting stronger.</p>
<p>Whether that is good for the web is a separate question, and I keep arriving at the same uncomfortable answer.</p>]]></content:encoded></item><item><title>Build 2026 and the platform that keeps absorbing</title><link>https://readme.news/build-2026-and-the-platform-that-keeps-absorbing/</link><guid isPermaLink="true">https://readme.news/build-2026-and-the-platform-that-keeps-absorbing/</guid><pubDate>Mon, 04 May 2026 09:00:00 +0000</pubDate><description>Agent infrastructure at the OS layer, more of the developer stack open-sourced, and a strategy that has not changed in a decade.</description><content:encoded><![CDATA[<p>Microsoft's developer conference ran this week. The individual announcements matter less than the consistency of the strategy, which has been unchanged for about ten years and keeps working.</p>
<h2 id="the-strategy">the strategy<a class="anchor" href="#the-strategy" aria-label="link to this section">#</a></h2>
<ol><li>Meet developers where they are, including on other people's platforms.</li><li>Adopt other people's standards rather than inventing competing ones.</li><li>Open-source the layers where control is not worth the friction.</li><li>Monetize the cloud underneath.</li></ol>
<p>Every Build for a decade has been an execution of that, and the cumulative result is that a company which was actively hostile to open source in 2005 is now the largest corporate contributor to it and owns the default editor, the default code host, and a large share of the developer toolchain.</p>
<h2 id="the-agent-infrastructure">the agent infrastructure<a class="anchor" href="#the-agent-infrastructure" aria-label="link to this section">#</a></h2>
<p>The substantive announcements this year continue the theme of putting agent capabilities at the operating system layer rather than in an application: a <a class="xref" href="/nodejs-24-and-the-slow-reinvention-of-the-runtime/" title="Node.js 24 and the slow reinvention of the runtime">permission model</a> for what agents may reach, an identity model for agents acting on a user's behalf, and audit surfaces for what they did.</p>
<p>That is the right layer for it. The alternative — every application implementing its own agent permission model — produces exactly the inconsistency that made mobile permissions a mess for a decade before the platforms standardized.</p>
<p>The questions that matter and that a keynote cannot answer:</p>
<p><strong>Granularity.</strong> "Filesystem access" is not a permission, it is a surrender. Does the model support "read from this directory for this task"?</p>
<p><strong>Consent fatigue.</strong> If the prompts are frequent, users click through them, and the control is theater. The design problem is asking rarely and meaningfully.</p>
<p><strong>Revocation and audit.</strong> Can a user see what an agent did and undo it? This is the part that is hardest and gets the least attention.</p>
<p>I will believe the <a class="xref" href="/the-component-model-and-the-plugin-problem/" title="The component model and the plugin problem">security model</a> when someone publishes an analysis of it, not when it is demonstrated on a stage.</p>
<h2 id="the-enterprise-angle">the enterprise angle<a class="anchor" href="#the-enterprise-angle" aria-label="link to this section">#</a></h2>
<p>The genuinely differentiating position Microsoft has is that they can offer agent capabilities inside an enterprise's existing identity, compliance, and audit infrastructure.</p>
<p>That is worth more to a large organization than raw capability. An agent that works within the existing access control model, logs to the existing audit system, and is governed by the existing data policies clears procurement. An agent that requires a new trust boundary does not, regardless of how good it is.</p>
<p>This is the same advantage that won enterprise cloud and it is being applied identically.</p>
<h2 id="what-a-developer-should-actually-do-with-this">what a developer should actually do with this<a class="anchor" href="#what-a-developer-should-actually-do-with-this" aria-label="link to this section">#</a></h2>
<p><strong>If you build on Windows:</strong> the tooling story is genuinely good now — WSL, the terminal, winget, PowerShell 7 — and if your Windows support has been a grudging afterthought since 2018, it is worth revisiting.</p>
<p><strong>If you build agent-adjacent products:</strong> design against the OS permission model rather than around it. Products that require users to disable platform protections do not get enterprise adoption.</p>
<p><strong>If you are evaluating anything announced here:</strong> wait for the second version. Microsoft's first releases in a new category are consistently rough and consistently improved within a year. That is a reasonable pattern and it means the launch-day evaluation is not the useful one.</p>
<h2 id="the-pattern-to-watch">the pattern to watch<a class="anchor" href="#the-pattern-to-watch" aria-label="link to this section">#</a></h2>
<p>The layer where the industry is currently fighting is not the model. It is the control plane for agents: who they are, what they may do, on whose behalf, with what audit trail.</p>
<p>Every platform vendor is building this. The one that becomes standard will have the same kind of position that identity providers have today, and it will be very durable.</p>
<p>That is the strategic story of the next three years and it is being fought in permission dialogs rather than benchmarks.</p>]]></content:encoded></item><item><title>Search is hard and you should probably not build it</title><link>https://readme.news/search-is-hard-and-you-should-probably-not-build-it/</link><guid isPermaLink="true">https://readme.news/search-is-hard-and-you-should-probably-not-build-it/</guid><pubDate>Wed, 29 Apr 2026 09:00:00 +0000</pubDate><description>Relevance ranking is a specialist discipline. Here&#x27;s the decision tree, and what to do at each level.</description><content:encoded><![CDATA[<p>Every product eventually adds a search box. The distance between "a search box that works" and "a search box users trust" is much larger than it looks, and most teams discover this after committing.</p>
<h2 id="the-levels">the levels<a class="anchor" href="#the-levels" aria-label="link to this section">#</a></h2>
<p><strong>Level 0: <code>LIKE '%query%'</code>.</strong></p>
<p>Works for tiny datasets. No ranking, no stemming, no typo tolerance, and a full table scan on every query.</p>
<p>Fine for an admin tool with a thousand rows. Not fine for anything a customer touches.</p>
<p><strong>Level 1: your database's full-text search.</strong></p>
<p>Postgres <code>tsvector</code>, MySQL full-text, SQLite FTS5. You get stemming, stop words, ranking, and index support.</p>
<div class="code"><span class="code-lang">sql</span><pre><code class="lang-sql">ALTER TABLE articles ADD COLUMN search tsvector
  GENERATED ALWAYS AS (
    setweight(to_tsvector('english', coalesce(title,'')), 'A') ||
    setweight(to_tsvector('english', coalesce(body,'')), 'B')
  ) STORED;

CREATE INDEX ON articles USING GIN (search);

SELECT id, title, ts_rank(search, q) AS rank
FROM articles, websearch_to_tsquery('english', $1) q
WHERE search @@ q
ORDER BY rank DESC LIMIT 20;</code></pre></div>
<p>The <code>setweight</code> calls are the part people miss: a match in the title should outrank a match in the body, and without weighting it does not.</p>
<p><strong>This handles most applications.</strong> If you have under a few million documents and your users search for terms that appear in them, stop here. One system, no synchronization problem, joins to your relational data.</p>
<p><strong>Level 2: a dedicated search engine.</strong></p>
<p>Elasticsearch, OpenSearch, Typesense, Meilisearch. You get: typo tolerance, faceting, synonyms, custom analyzers, distributed scaling, and much better relevance tuning.</p>
<p>The cost is a second system with a synchronization problem. Your search index is now eventually consistent with your database, and every write path must update both. That inconsistency will produce bugs — a deleted item still appearing in results is the classic — and handling it correctly is real work.</p>
<p><strong>Level 3: hybrid semantic search.</strong></p>
<p>Vector embeddings alongside keyword search, combined with reciprocal rank fusion or a learned reranker.</p>
<p>This handles the case where the user's words are not the document's words. "How do I cancel" should find "Terminating your subscription."</p>
<p><strong>Important:</strong> hybrid, not pure vector. Pure semantic search is bad at exact matches — product codes, error numbers, names, function names — precisely the queries where users are most certain about what they want and least tolerant of a wrong answer.</p>
<h2 id="what-makes-search-actually-good">what makes search actually good<a class="anchor" href="#what-makes-search-actually-good" aria-label="link to this section">#</a></h2>
<p>The engine is the easy part. Relevance is the hard part, and it is mostly not about the algorithm.</p>
<p><strong>Weight your fields.</strong> Title beats body. Exact phrase beats individual terms. Recent beats old, for content where recency matters.</p>
<p><strong>Use behavioral signals.</strong> What users clicked on for similar queries is the strongest relevance signal available, and it requires logging queries and clicks from day one. Retrofitting this means starting your data collection from zero.</p>
<p><strong>Handle the empty result.</strong> "No results for X" is a failure. Show something: did-you-mean, related content, popular items, a way to browse. An empty page is where users leave.</p>
<p><strong>Handle the head queries manually.</strong> A small number of queries make up a large share of volume. Look at your top hundred, check what they return, and pin the correct answer where it is wrong. This is unglamorous, takes an afternoon, and improves perceived quality more than any algorithmic change.</p>
<p><strong>Log everything.</strong> Query, result count, position clicked, whether anything was clicked. The searches that return nothing and the searches where nobody clicks are your improvement backlog, delivered for free.</p>
<h2 id="the-instrumentation-that-matters">the instrumentation that matters<a class="anchor" href="#the-instrumentation-that-matters" aria-label="link to this section">#</a></h2>
<p>Three metrics:</p>
<ol><li><strong>Zero-result rate.</strong> Should be low. Every zero-result query is a user who did not find what they wanted.</li><li><strong>Click-through rate</strong>, and the position clicked. If people consistently click the fifth result, your ranking is wrong.</li><li><strong>Query refinement rate.</strong> Users who search, then immediately search again with different words, did not find it the first time.</li></ol>
<p>Most teams have none of these and are tuning relevance by intuition.</p>
<h2 id="the-recommendation">the recommendation<a class="anchor" href="#the-recommendation" aria-label="link to this section">#</a></h2>
<p>Start at level 1. Postgres full-text search with weighted fields covers more applications than people expect, and it does not introduce a synchronization problem.</p>
<p>Move to level 2 when you have a specific complaint you cannot fix — typo tolerance, faceting at scale, or performance. Move to level 3 when you have evidence that users search for concepts rather than terms.</p>
<p>And whatever level you are at: log the queries. That data is the input to every future improvement and you cannot get it retroactively.</p>]]></content:encoded></item><item><title>Why your build is slow</title><link>https://readme.news/why-your-build-is-slow/</link><guid isPermaLink="true">https://readme.news/why-your-build-is-slow/</guid><pubDate>Wed, 25 Mar 2026 09:00:00 +0000</pubDate><description>Six causes, in order of how often they&#x27;re the actual problem, with the fix for each.</description><content:encoded><![CDATA[<p>Build time is the tax you pay on every single change, and most teams have never measured where it goes. Here are the causes, roughly ordered by how often I find each one to be the actual bottleneck.</p>
<h2 id="1-you-are-not-caching-or-your-cache-never-hits">1. You are not caching, or your cache never hits<a class="anchor" href="#1-you-are-not-caching-or-your-cache-never-hits" aria-label="link to this section">#</a></h2>
<p>By far the most common. Not "we have no cache" — teams have caches. The caches do not hit.</p>
<p>Cache keys that include a timestamp, a branch name, a commit SHA, or anything else that changes every run will produce a 0% hit rate while looking like a working cache. Worse than nothing, because you also pay the upload.</p>
<p><strong>Measure the hit rate first.</strong> Most CI systems report it and almost nobody looks. If it is not above 80% on a typical branch build, fix the key before you do anything else.</p>
<p>The key should be a hash of the inputs that actually determine the output: the lockfile for dependencies, the source files for a compilation unit. Nothing else.</p>
<h2 id="2-you-are-rebuilding-things-that-did-not-change">2. You are rebuilding things that did not change<a class="anchor" href="#2-you-are-rebuilding-things-that-did-not-change" aria-label="link to this section">#</a></h2>
<p>If your build runs everything on every commit, you are paying for the whole repository to verify a one-line change in one module.</p>
<p>The fix is a build graph that knows what depends on what and only rebuilds the affected subtree. This is what Bazel, Buck, Nx, Turborepo, and Gradle's configuration cache exist for.</p>
<p>The cost is real — adopting a graph-aware build system is a project, and for a small repository it is not worth it. The threshold is roughly when your full build exceeds five minutes and most changes touch one small part.</p>
<h2 id="3-your-tests-are-serial-or-parallel-badly">3. Your tests are serial, or parallel badly<a class="anchor" href="#3-your-tests-are-serial-or-parallel-badly" aria-label="link to this section">#</a></h2>
<p>Two separate failures.</p>
<p><strong>Serial:</strong> you have 4,000 tests running one after another on one core. Almost every test framework has a parallel mode and it is frequently off by default because parallel tests expose shared state.</p>
<p>If your tests fail when run in parallel, that is a bug in the tests — shared fixtures, a common database, an order dependency — and fixing it is worth doing regardless of speed.</p>
<p><strong>Parallel badly:</strong> you split tests into eight shards by file count, and one shard takes six minutes while the others take ninety seconds. Your suite takes six minutes.</p>
<p>Split by <em>historical duration</em>, not by count. Most CI systems support this and it frequently halves the wall clock for free.</p>
<h2 id="4-you-are-doing-full-clones">4. You are doing full clones<a class="anchor" href="#4-you-are-doing-full-clones" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">yaml</span><pre><code class="lang-yaml">- uses: actions/checkout@v4
  with:
    fetch-depth: 1</code></pre></div>
<p>A repository with a long history and large binary files can take minutes to clone. Almost no build step needs history. If one does — a <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a> generator, a version-from-tags scheme — give that one job the full clone and shallow-clone everything else.</p>
<h2 id="5-your-runners-are-too-small">5. Your runners are too small<a class="anchor" href="#5-your-runners-are-too-small" aria-label="link to this section">#</a></h2>
<p>The least satisfying answer and frequently the correct one.</p>
<p>Engineer time is expensive. A larger runner that halves your build time pays for itself immediately at any reasonable team size, and the arithmetic is easy to do:</p>
<div class="code"><pre><code>build minutes/day × engineers × wait fraction × hourly cost
    vs.
extra runner cost</code></pre></div>
<p>The extra runner is almost always cheaper. Teams resist this because compute cost is a visible line item and engineer waiting is not.</p>
<h2 id="6-you-are-compiling-in-the-container-build">6. You are compiling in the container build<a class="anchor" href="#6-you-are-compiling-in-the-container-build" aria-label="link to this section">#</a></h2>
<p>If your Dockerfile runs <code>npm install</code> and then a compile, and you are not using BuildKit cache mounts and multi-stage builds properly, you are redoing work on every image build.</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile"># cache mount survives across builds
RUN --mount=type=cache,target=/root/.npm \
    npm ci</code></pre></div>
<p>Layer ordering matters too: copy the lockfile and install before copying source, so a source change does not invalidate the dependency layer. This is well known and consistently gotten wrong.</p>
<h2 id="the-measurement">the measurement<a class="anchor" href="#the-measurement" aria-label="link to this section">#</a></h2>
<p>Before fixing anything, get the breakdown:</p>
<div class="code"><pre><code>queue wait     ← runner capacity or concurrency limits
checkout       ← clone depth, LFS
dependency install ← caching
build          ← incrementality
test           ← parallelism
publish        ← usually fine</code></pre></div>
<p>Most teams find that one stage is 60% of the time and they had assumed it was a different one.</p>
<h2 id="the-target">the target<a class="anchor" href="#the-target" aria-label="link to this section">#</a></h2>
<p>Under ten minutes for the check that gates merge. Under two minutes if you can get there.</p>
<p>The reason for the ten-minute line is behavioral: past it, people stop waiting, start batching, and the whole delivery process degrades in ways that are hard to attribute back to the build.</p>
<p>Not everything has to be in that ten minutes. Split the blocking check from the exhaustive one. Lint, typecheck, and unit tests gate the merge. Integration, e2e, and the full platform matrix run after and page someone on failure.</p>
<p>That single split is usually the largest available improvement and it requires no new tooling.</p>]]></content:encoded></item><item><title>AGENTS.md and the repository that explains itself</title><link>https://readme.news/agentsmd-and-the-repository-that-explains-itself/</link><guid isPermaLink="true">https://readme.news/agentsmd-and-the-repository-that-explains-itself/</guid><pubDate>Sun, 04 Jan 2026 09:00:00 +0000</pubDate><description>A convention nobody standardized became standard anyway. Here&#x27;s what belongs in it and what doesn&#x27;t.</description><content:encoded><![CDATA[<p>Over the last year, essentially every coding agent converged on the same mechanism: a markdown file in the repository root telling the agent how to work in this codebase.</p>
<p>The names varied — <code>AGENTS.md</code>, <code>CLAUDE.md</code>, <code>.cursorrules</code>, <code>.github/copilot-instructions.md</code> — and the format did not. It is prose. The model reads it. That is the whole protocol.</p>
<p>A convention that emerges independently in five products in one year is telling you something about the shape of the problem.</p>
<h2 id="what-actually-belongs-in-it">what actually belongs in it<a class="anchor" href="#what-actually-belongs-in-it" aria-label="link to this section">#</a></h2>
<p>I have written and rewritten these a dozen times. The version that works is shorter than you think and more specific than you want.</p>
<p><strong>Commands.</strong> The exact invocations. Not "run the tests" — the command, including the flags, including how to run one test file.</p>
<div class="code"><span class="code-lang">markdown</span><pre><code class="lang-markdown">## commands
- install: `pnpm install --frozen-lockfile`
- test: `pnpm vitest run`
- test one file: `pnpm vitest run src/foo.test.ts`
- typecheck: `pnpm tsc --noEmit`
- lint: `pnpm biome check --write .`
- dev server: `pnpm dev` (port 5173)</code></pre></div>
<p>This section alone eliminates most wasted agent turns. Without it, the agent guesses, guesses wrong, and spends four tool calls discovering your test runner.</p>
<p><strong>Non-obvious structure.</strong> Where things live, when it is not inferable. "Database migrations are in <code>db/migrations</code> and must be created with <code>pnpm db:new</code>, never by hand." "The <code>legacy/</code> directory is frozen — do not modify it."</p>
<p><strong>Conventions that a linter does not enforce.</strong> If your linter catches it, do not write it down; the linter will tell the agent. Write down the things that are policy rather than syntax: "prefer composition over inheritance in <code>src/domain</code>", "all public functions in <code>api/</code> need a docstring", "we do not use default exports".</p>
<p><strong>Things that will break.</strong> "Do not run <code>pnpm build</code> — it takes 12 minutes and is not needed for tests." "The integration tests require Docker; skip them if it is not running." "Never modify <code>schema.sql</code> directly."</p>
<p><strong>Boundaries.</strong> What the agent may not touch without asking. Production configs, migration files, anything with a security review requirement.</p>
<h2 id="what-does-not-belong">what does not belong<a class="anchor" href="#what-does-not-belong" aria-label="link to this section">#</a></h2>
<p><strong>Your entire architecture document.</strong> The agent reads this every session. A three-thousand-word essay costs tokens on every single request and dilutes attention across a lot of text that is irrelevant to most tasks.</p>
<p>Keep it under about 200 lines. If you need more, put it in a separate document and reference it: "for the event pipeline design, read <code>docs/events.md</code> before changing anything in <code>src/events/</code>."</p>
<p><strong>Anything the code already says.</strong> Do not restate the type signatures. The agent can read.</p>
<p><strong>Aspirations.</strong> "We value clean code" is not an instruction. "Functions over 40 lines get split" is.</p>
<p><strong>Rules you do not actually follow.</strong> If the codebase contradicts the file, the codebase wins in the model's attention, and now your instructions are noise.</p>
<h2 id="the-part-i-did-not-expect">the part I did not expect<a class="anchor" href="#the-part-i-did-not-expect" aria-label="link to this section">#</a></h2>
<p>Writing these has improved my documentation for humans.</p>
<p>The discipline of writing "here is exactly how to run the tests, here are the things that will surprise you, here is what not to touch" is precisely what a new engineer needs on day one, and it is precisely what most onboarding documents fail to contain because they were written by someone who already knew.</p>
<p>An agent is an infinitely patient new hire who will follow instructions literally and never ask a clarifying question out of politeness. That turns out to be an excellent test of whether your instructions are any good.</p>
<p>Several teams I know have merged their onboarding doc and their agent file into one. That is the correct end state.</p>
<h2 id="the-standardization-question">the standardization question<a class="anchor" href="#the-standardization-question" aria-label="link to this section">#</a></h2>
<p>There is an ongoing effort to consolidate on <code>AGENTS.md</code> as the common name, with tools reading it as a fallback. That would be good and it is a coordination problem, which means it will take longer than it should.</p>
<p>In the meantime: write one file, symlink the rest. It costs nothing.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">ln -s AGENTS.md CLAUDE.md</code></pre></div>]]></content:encoded></item><item><title>Gemini 3 arrives with an IDE attached</title><link>https://readme.news/gemini-3-arrives-with-an-ide-attached/</link><guid isPermaLink="true">https://readme.news/gemini-3-arrives-with-an-ide-attached/</guid><pubDate>Wed, 19 Nov 2025 09:00:00 +0000</pubDate><description>Google ships a frontier model and Antigravity, an agent-first development environment. The bundling is the strategy.</description><content:encoded><![CDATA[<p>Google released Gemini 3 Pro yesterday along with Antigravity, an agent-first development environment, and integration of the model directly into Search's AI Mode on launch day.</p>
<h2 id="the-model">the model<a class="anchor" href="#the-model" aria-label="link to this section">#</a></h2>
<p>Strong across reasoning, multimodal understanding, and coding benchmarks. A "Deep Think" mode for the hardest problems. The million-token context window carries over.</p>
<p>The benchmark numbers are competitive at the frontier. At this point that sentence describes every major release, which is the actual news — the frontier is a cluster, not a leader.</p>
<p>What differentiates a release now is not the top-line capability. It is:</p>
<ul><li><strong>Price per unit of capability</strong>, where Google's TPU position is a real structural advantage.</li><li><strong>Context handling at length</strong>, where Google has led for a while.</li><li><strong>Multimodal</strong>, where native training rather than adapters keeps paying off.</li><li><strong>Distribution</strong>, where shipping into Search on day one is something no competitor can do.</li></ul>
<p>That last one deserves emphasis. Google put a new <a class="xref" href="/small-models-ate-the-middle/" title="Small models ate the middle">frontier model</a> into the search product used by billions of people on launch day. The previous norm was a staged rollout over months. That is a capability nobody else has and it is the reason Google's position looks different than it did in 2023.</p>
<h2 id="antigravity">Antigravity<a class="anchor" href="#antigravity" aria-label="link to this section">#</a></h2>
<p>An agent-first IDE — a VS Code derivative where the primary interaction is directing agents rather than editing text, with a manager surface for orchestrating multiple agents in parallel across editor, terminal, and browser.</p>
<p>The interesting design decision is <strong>artifacts</strong>: agents produce task lists, plans, screenshots, and browser recordings as reviewable outputs, rather than requiring you to read a raw transcript to figure out what happened.</p>
<p>That addresses the actual problem with delegated agents, which I have written about before: review is the bottleneck. A transcript of four hundred tool calls is not reviewable. A plan, a diff, and a recording of the browser test passing is.</p>
<p>Whether this specific implementation is good, I do not know yet — first releases of IDEs rarely are. The direction is right, and it is the first serious attempt I have seen at designing for review rather than for generation.</p>
<h2 id="the-bundling">the bundling<a class="anchor" href="#the-bundling" aria-label="link to this section">#</a></h2>
<p>Model, IDE, CLI, cloud, and search distribution, from one vendor, priced aggressively.</p>
<p>This is the classic platform playbook and Google is executing it more coherently than they have on anything in a decade. The pieces reinforce each other: the IDE drives model usage, the model drives cloud usage, the cloud subsidizes the free tiers, and the search distribution provides the consumer volume that funds all of it.</p>
<p>The competitive question for everyone else is whether best-of-breed beats integrated. Historically it has, in developer tools, because developers choose their own tools and choose the best one. It has not, in enterprise procurement, where bundles win.</p>
<p>Both markets exist. The bundle is going to do well in one of them.</p>
<h2 id="the-practical-advice">the practical advice<a class="anchor" href="#the-practical-advice" aria-label="link to this section">#</a></h2>
<p>Same as every model release, and I will keep repeating it because it keeps being the right answer:</p>
<p>Run your evals. Gemini 3 is likely better than what you are using on some dimensions and different on all of them. The migration cost is a day if you have an eval harness and a week of guessing if you do not.</p>
<p>Try Antigravity on a real task, not a demo task. Agent IDEs differ enormously in how they handle a twenty-minute task versus a two-minute one, and the demos are all two-minute tasks.</p>]]></content:encoded></item><item><title>Rust 1.90 and the linker nobody talks about</title><link>https://readme.news/rust-190-and-the-linker-nobody-talks-about/</link><guid isPermaLink="true">https://readme.news/rust-190-and-the-linker-nobody-talks-about/</guid><pubDate>Fri, 19 Sep 2025 09:00:00 +0000</pubDate><description>LLD becomes the default linker on x86-64 Linux. Link times drop and nobody notices, which is the point.</description><content:encoded><![CDATA[<p>Rust 1.90 makes LLD the default linker for <code>x86_64-unknown-linux-gnu</code>. This is a build-time performance change with no user-facing API and it is one of the more useful things to ship this year.</p>
<h2 id="why-linking-matters">why linking matters<a class="anchor" href="#why-linking-matters" aria-label="link to this section">#</a></h2>
<p>For a large Rust binary, linking is frequently the dominant cost of an incremental build. You change one line, the compiler recompiles one crate quickly, and then the linker spends several seconds stitching together a hundred megabytes of object files and debug information.</p>
<p>GNU <code>ld</code> is old, single-threaded in the parts that matter, and was designed for a different era of binary sizes. <code>lld</code> is parallel, substantially faster, and has been production-ready for years — it is the default on several other platforms already.</p>
<p>Reported improvements vary by project, with the largest gains on debug builds of large dependency trees. Two to five times faster linking is a common range. If your edit-compile-run loop is four seconds and linking was two of them, you feel this every single time.</p>
<h2 id="how-to-check">how to check<a class="anchor" href="#how-to-check" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">cargo build --timings</code></pre></div>
<p>Opens an HTML report showing where time went, including link time as a distinct phase. Most people have never looked at this and are surprised by what it says.</p>
<p>If you were already opting into <code>lld</code> or <code>mold</code> via <code>.cargo/config.toml</code>, nothing changes for you. If you were not — which is most people, because the configuration was obscure and required knowing it existed — you get the improvement for free on upgrade.</p>
<h2 id="the-general-lesson">the general lesson<a class="anchor" href="#the-general-lesson" aria-label="link to this section">#</a></h2>
<p>Defaults are the most impactful thing a toolchain ships.</p>
<p>A faster linker has been available for years. Anyone could have configured it. The instructions were in a blog post. Almost nobody did, because it required knowing the option existed, knowing it was safe, and caring enough to edit a config file.</p>
<p>Changing the default delivers the improvement to everyone at once. That is worth more than a decade of documentation.</p>
<p>This generalizes. If you maintain a tool and you find yourself writing documentation explaining how to enable a better behavior, ask whether the better behavior should be the default. The answer is usually yes, and the reason it is not is usually caution about breaking a small number of unusual setups — which is a real concern that should be handled with a flag to opt <em>out</em>, not a flag to opt in.</p>
<h2 id="the-rest-of-190">the rest of 1.90<a class="anchor" href="#the-rest-of-190" aria-label="link to this section">#</a></h2>
<ul><li><strong>Cargo workspace publishing</strong> — <code>cargo publish --workspace</code> publishes multiple interdependent crates in the correct order. Anyone maintaining a multi-crate project has written a shell script for this. Now you can delete it.</li><li><strong>Demoted target tiers</strong> for several less-used platforms.</li><li><strong><code>x86_64-apple-darwin</code> demoted to tier 2</strong> with host tools still provided, which reflects where Apple hardware has actually gone.</li><li>The usual const stabilizations and API additions.</li></ul>
<h2 id="the-trend-line">the trend line<a class="anchor" href="#the-trend-line" aria-label="link to this section">#</a></h2>
<p>Rust's reputation for slow compilation is the single most cited objection to adopting it, and it has been steadily addressed for years: incremental compilation, parallel front-end work, the new <a class="xref" href="/rust-184-and-the-quiet-rewrite-underneath/" title="Rust 1.84 and the quiet rewrite underneath">trait solver</a>, better <a class="xref" href="/caching-is-the-only-optimization-that-reliably-works/" title="Caching is the only optimization that reliably works">caching</a>, and now the linker.</p>
<p>Compile times are meaningfully better than they were three years ago and still slower than Go. That gap is partly fundamental — monomorphization and the amount of optimization Rust does are not free — and partly still addressable.</p>
<p>Anyone who tried Rust in 2021 and bounced off the build times should try again. The numbers are different now.</p>]]></content:encoded></item><item><title>The monorepo question has a boring answer</title><link>https://readme.news/the-monorepo-question-has-a-boring-answer/</link><guid isPermaLink="true">https://readme.news/the-monorepo-question-has-a-boring-answer/</guid><pubDate>Wed, 03 Sep 2025 09:00:00 +0000</pubDate><description>It&#x27;s not about repository count. It&#x27;s about whether you can make an atomic change across your system.</description><content:encoded><![CDATA[<p>The monorepo debate is one of the few technical arguments that has been running for fifteen years without either side updating. That usually means both sides are arguing about the wrong thing.</p>
<h2 id="the-actual-question">the actual question<a class="anchor" href="#the-actual-question" aria-label="link to this section">#</a></h2>
<p>Not "one repository or many." The question is:</p>
<p><strong>Can you make a change that spans multiple components, atomically, with a single review and a single test run?</strong></p>
<p>If yes, you have the property that matters. If no, you do not, regardless of how many repositories you have.</p>
<p>You can have one repository where every service has its own pipeline, its own version, and its own release train — and get none of the benefit, because a cross-cutting change still requires coordination. That is a polyrepo wearing a monorepo's directory structure.</p>
<p>You can have separate repositories with a workspace tool that checks them all out and builds them together, and get most of the benefit.</p>
<h2 id="what-atomicity-buys">what atomicity buys<a class="anchor" href="#what-atomicity-buys" aria-label="link to this section">#</a></h2>
<p><strong>Refactoring at scale becomes possible.</strong> Renaming a function used by twelve services is one commit. In a polyrepo it is twelve pull requests, a deprecation period, a compatibility shim, and a cleanup task nobody does. This is the single biggest difference and it compounds over years — polyrepo codebases accumulate <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> debt because changing an interface is expensive.</p>
<p><strong>One version of every dependency.</strong> No "service A is on library v2, service B is on v4, and the shared types are subtly incompatible." Upgrading is a single change that either passes CI everywhere or does not.</p>
<p><strong>Discoverability.</strong> Grep works. Find-references works. A new engineer can read the whole system.</p>
<h2 id="what-it-costs">what it costs<a class="anchor" href="#what-it-costs" aria-label="link to this section">#</a></h2>
<p><strong>Tooling.</strong> Above a certain size, standard tools break. <code>git status</code> gets slow. Your IDE indexes forever. CI running everything on every commit becomes infeasible. You need build graph analysis to test only what changed. That is real infrastructure investment and it never ends.</p>
<p><strong>Coupling by default.</strong> Shared code is easy, so shared code happens, and suddenly a change to a "utility" function affects nine teams. Polyrepo makes coupling expensive, which is a crude but effective forcing function.</p>
<p><strong>Access control.</strong> One repository means one permission boundary, mostly. Fine for a company, not fine for open source with contractors.</p>
<p><strong>Blast radius on the build system.</strong> Everyone depends on it. When it breaks, everyone stops.</p>
<h2 id="the-honest-heuristic">the honest heuristic<a class="anchor" href="#the-honest-heuristic" aria-label="link to this section">#</a></h2>
<ul><li><strong>Under ~50 engineers</strong>: monorepo, almost always. The tooling cost has not kicked in and the coordination benefit is immediate.</li><li><strong>50 to 500</strong>: monorepo if you are willing to staff a build/<a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">platform team</a>. If you are not, the tooling debt will eat you.</li><li><strong>Over 500</strong>: whatever you already have, improved. Migrating a large organization between these models is a multi-year project with a poor track record, and the energy is better spent on making the current arrangement work.</li><li><strong>Open source with external contributors</strong>: polyrepo, usually. The access control and contribution-scope arguments dominate.</li></ul>
<h2 id="the-thing-both-sides-get-wrong">the thing both sides get wrong<a class="anchor" href="#the-thing-both-sides-get-wrong" aria-label="link to this section">#</a></h2>
<p>Monorepo advocates point at Google and Meta. Those companies have hundreds of engineers working full-time on build infrastructure. You do not. Their approach is not transferable without their investment, and citing them is like citing an F1 team's pit strategy for your commute.</p>
<p>Polyrepo advocates point at microservice independence. Then they build a shared library, and a shared types package, and a shared client SDK, and now every service is coupled through three packages with independent release cycles and version skew — which is strictly worse than the coupling they were avoiding.</p>
<h2 id="what-to-actually-do">what to actually do<a class="anchor" href="#what-to-actually-do" aria-label="link to this section">#</a></h2>
<p>Whichever you have, invest in the property you are missing.</p>
<p><strong>In a monorepo</strong>: enforce boundaries. Ownership files, dependency rules that fail the build, and a rule that a module's public API is explicit. Physical colocation must not mean logical coupling.</p>
<p><strong>In a polyrepo</strong>: invest in cross-repo change tooling. A script that opens coordinated pull requests. A dependency <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>. A way to run integration tests across repositories before merge. This is unglamorous and it is the difference between a polyrepo that works and one that slowly ossifies.</p>
<p>The repository count is a detail. The workflow is the thing.</p>]]></content:encoded></item><item><title>Gemini CLI puts a free agent in your terminal</title><link>https://readme.news/gemini-cli-puts-a-free-agent-in-your-terminal/</link><guid isPermaLink="true">https://readme.news/gemini-cli-puts-a-free-agent-in-your-terminal/</guid><pubDate>Thu, 26 Jun 2025 09:00:00 +0000</pubDate><description>Apache 2.0, generous free limits, and a very direct shot at the terminal-agent category.</description><content:encoded><![CDATA[<p>Google released Gemini CLI: an open-source terminal agent under Apache 2.0, with a free tier that includes a large daily request allowance and access to <a class="xref" href="/gemini-25-pro-is-googles-best-model-and-it-shows/" title="Gemini 2.5 Pro is Google&#x27;s best model and it shows">Gemini 2.5 Pro</a> with its million-token context.</p>
<p>The free tier is the story. This is a competitive move priced at zero.</p>
<h2 id="what-it-does">what it does<a class="anchor" href="#what-it-does" aria-label="link to this section">#</a></h2>
<p>The familiar shape: a terminal agent with filesystem access, shell execution, git awareness, and web search. It reads a <code>GEMINI.md</code> for project-specific instructions. It supports MCP servers for extending its tool set.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">npx https://github.com/google-gemini/gemini-cli</code></pre></div>
<p>Sign in with a Google account and you are running. No API key, no billing setup, no credit card.</p>
<h2 id="why-free">why free<a class="anchor" href="#why-free" aria-label="link to this section">#</a></h2>
<p>Two reasons, both strategic.</p>
<p><strong>TPU economics.</strong> Google serves its own models on its own silicon in its own datacenters. The marginal cost of inference is genuinely lower for them than for anyone renting Nvidia capacity, and they can afford a free tier that competitors cannot match without eating a loss.</p>
<p><strong>Developer mindshare is a leading indicator.</strong> The developers who adopt a tool today choose the platform their company standardizes on in two years. Google has lost that fight repeatedly — to AWS in cloud, to OpenAI in AI APIs — and this is a deliberate attempt to not lose it again.</p>
<h2 id="the-open-source-part">the open source part<a class="anchor" href="#the-open-source-part" aria-label="link to this section">#</a></h2>
<p>Apache 2.0 on the whole client. That means you can fork it, audit it, run it against a different model endpoint if you rewire it, and inspect exactly what it sends where.</p>
<p>That last one matters more than people acknowledge. A terminal agent has filesystem and shell access. "What does this thing actually transmit" is a question a security team will ask, and "here is the source" is a much better answer than a data processing addendum.</p>
<p>I expect the open-source-ness to be adopted as table stakes. It is very hard to argue for a closed terminal agent when a competitive one is Apache 2.0.</p>
<h2 id="the-current-state-of-the-category">the current state of the category<a class="anchor" href="#the-current-state-of-the-category" aria-label="link to this section">#</a></h2>
<p>By my count there are now four credible terminal-based coding agents from major vendors, plus several from startups, plus a healthy open-source contingent. All shipped within about six months.</p>
<p>They are converging fast on the same feature set: filesystem tools, shell, project instruction files, MCP support, permission prompts, git integration. The differences that remain:</p>
<ul><li><strong>Model quality on long-horizon tasks</strong>, which is the real one.</li><li><strong>Context handling</strong> — how they decide what to read and when to compact.</li><li><strong>Permission ergonomics</strong> — how annoying the safety prompts are, which sounds trivial and determines whether people turn them off.</li><li><strong>Price</strong>, where Google just set an aggressive anchor.</li></ul>
<h2 id="the-practical-advice">the practical advice<a class="anchor" href="#the-practical-advice" aria-label="link to this section">#</a></h2>
<p>Try more than one on the same task. They differ more in practice than the feature lists suggest, and the differences show up on tasks that take twenty minutes, not on tasks that take two.</p>
<p>And whichever you use: sandbox it. Container, VM, or at minimum a dedicated user account without your production credentials in its environment. The tooling is good. The <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> are still real, and "the agent ran a command I did not read carefully" is a bad way to learn that.</p>]]></content:encoded></item><item><title>Cursor at $9 billion and the coding-tool land grab</title><link>https://readme.news/cursor-at-9-billion-and-the-coding-tool-land-grab/</link><guid isPermaLink="true">https://readme.news/cursor-at-9-billion-and-the-coding-tool-land-grab/</guid><pubDate>Mon, 23 Jun 2025 09:00:00 +0000</pubDate><description>An editor fork raises at a valuation that only makes sense if you believe the IDE is the control point.</description><content:encoded><![CDATA[<p>Anysphere, which makes Cursor, raised at a reported $9 billion valuation this month. The product is a fork of VS Code with AI features. That sentence undersells it, and it also explains why the valuation is contested.</p>
<h2 id="what-they-actually-built">what they actually built<a class="anchor" href="#what-they-actually-built" aria-label="link to this section">#</a></h2>
<p>Cursor's technical differentiation is real and it is mostly not the model.</p>
<p><strong>Codebase indexing that works.</strong> Semantic search over the whole repository, kept fresh, so the model gets relevant context without you selecting files. This is harder than it sounds at repository scale and it is where a lot of competitors are visibly worse.</p>
<p><strong>Fast apply.</strong> A specialized model that takes a proposed edit and applies it to the existing file correctly. Sounds trivial. Is not. Getting a large model to reproduce an entire file with one changed function is slow and error-prone; a <a class="xref" href="/haiku-45-and-the-collapsing-cost-of-good-enough/" title="Haiku 4.5 and the collapsing cost of good-enough">small model</a> trained specifically on edit application is fast and reliable.</p>
<p><strong>Tab completion that predicts your next action</strong>, including where you are about to move the cursor, not just what you are about to type.</p>
<p>All three are inference-layer engineering rather than model training, and all three are the kind of thing that is unglamorous to build and immediately obvious in daily use.</p>
<h2 id="the-bear-case">the bear case<a class="anchor" href="#the-bear-case" aria-label="link to this section">#</a></h2>
<p>Cursor is a fork of VS Code, and VS Code is made by the company that owns GitHub Copilot, GitHub, and a large share of the developer's toolchain. Microsoft has been shipping the same features, more slowly, with better distribution and a bundled price.</p>
<p>Cursor also pays for inference on models it does not own, from vendors who are themselves building competing products. That is a structurally uncomfortable position: your primary cost is paid to your competitor, and your gross margin depends on their pricing decisions.</p>
<p>And the switching cost is approximately zero. It is an editor. Your settings sync in five minutes.</p>
<h2 id="the-bull-case">the bull case<a class="anchor" href="#the-bull-case" aria-label="link to this section">#</a></h2>
<p>Distribution in developer tools is earned by being better, not by being bundled — which is why VS Code beat Atom, why Git beat SVN, and why every attempt to push a mandated IDE on a team has failed. Cursor is currently better at the specific thing developers do all day, and developers are unusually willing to switch tools and unusually vocal when they do.</p>
<p>The valuation implies the IDE becomes the control point for AI-assisted development — the place where context lives, where policy is enforced, where every model call routes through. If that is true, whoever owns it has a durable position regardless of which model wins.</p>
<h2 id="my-read">my read<a class="anchor" href="#my-read" aria-label="link to this section">#</a></h2>
<p>The category is real and enormous. The specific defensibility is unclear. And the whole category has a structural problem that nobody has solved: as models get better at long-horizon delegated work, the <em>editor</em> becomes less central, because you are not editing. You are reviewing.</p>
<p>The product that wins the next phase might not be an editor at all. It might be whatever tool makes reviewing twelve agent-generated pull requests tolerable, and that looks a lot more like a code review <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> than an IDE.</p>
<p>Nobody has built a good version of that yet. That is where I would be looking.</p>]]></content:encoded></item><item><title>I/O 2025: Google puts AI in the search box and means it</title><link>https://readme.news/io-2025-google-puts-ai-in-the-search-box-and-means-it/</link><guid isPermaLink="true">https://readme.news/io-2025-google-puts-ai-in-the-search-box-and-means-it/</guid><pubDate>Wed, 21 May 2025 09:00:00 +0000</pubDate><description>AI Mode, Veo 3 with audio, Jules, and an Android XR headset. The distribution advantage, deployed.</description><content:encoded><![CDATA[<p>Google I/O was a two-hour demonstration of what it looks like when a company with a competitive model also owns the surfaces people already use.</p>
<h2 id="ai-mode-in-search">AI Mode in Search<a class="anchor" href="#ai-mode-in-search" aria-label="link to this section">#</a></h2>
<p>A tab in Google Search that runs a conversational, multi-step research flow instead of returning ten blue links. Rolling out broadly in the US.</p>
<p>This is the announcement with the largest downstream consequences and almost none of them are technical.</p>
<p>If a meaningful share of informational queries get answered in the results page, the traffic that funded the open web's content layer for twenty-five years goes away. Publishers have been shouting about this since AI Overviews launched and the shouting is going to get louder, because the data is going to get worse.</p>
<p>For developers specifically: your documentation site's traffic is going to fall, and the answers people get about your project will be synthesized from your docs by a model you do not control. The mitigations are unsatisfying — write docs that are hard to summarize badly, keep a <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a> that is machine-readable, make sure your <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a> contain enough context to be searchable — but the structural shift is not something an individual project can opt out of.</p>
<h2 id="veo-3">Veo 3<a class="anchor" href="#veo-3" aria-label="link to this section">#</a></h2>
<p>Video generation with synchronized audio — dialogue, sound effects, ambience — generated together rather than dubbed on afterward. The quality jump is substantial and the audio integration is the part that makes it feel different.</p>
<p>Ethically this is the loudest thing at the conference and it got the least discussion time. SynthID watermarking is included, which is genuinely better than nothing and is not a solution, because watermarks survive exactly as long as nobody is motivated to remove them.</p>
<h2 id="jules">Jules<a class="anchor" href="#jules" aria-label="link to this section">#</a></h2>
<p>An asynchronous coding agent. Clones your repo into a VM, works on a task, opens a pull request. Same shape as Codex and Copilot's agent mode. Everybody arrived at the same product within a month of each other, which tells you the capability threshold was crossed at roughly the same time for everyone.</p>
<h2 id="gemini-25-deep-think">Gemini 2.5 Deep Think<a class="anchor" href="#gemini-25-deep-think" aria-label="link to this section">#</a></h2>
<p>An enhanced reasoning mode for 2.5 Pro that explores multiple hypotheses in parallel before answering. Aimed at the hardest math and coding problems.</p>
<p>The technique — parallel sampling with selection, rather than one longer chain — is a different axis of test-time compute than "think longer," and I suspect it generalizes better. A single long chain compounds its own errors. Multiple independent attempts do not.</p>
<h2 id="android-xr">Android XR<a class="anchor" href="#android-xr" aria-label="link to this section">#</a></h2>
<p>A headset with Gemini integrated, plus glasses in development with Warby Parker and Gentle Monster as partners. Google has attempted this category twice and failed twice. The bet is that a genuinely useful assistant is the app that makes the form factor worth wearing.</p>
<p>Possible. The failure mode for the last decade has been that the hardware was uncomfortable and the software was a solution looking for a problem. This addresses the second one.</p>
<h2 id="the-throughline">the throughline<a class="anchor" href="#the-throughline" aria-label="link to this section">#</a></h2>
<p>Google's problem for the last two years was never capability. It was that OpenAI had the mindshare and Google had the users but was afraid to touch them.</p>
<p>This was the conference where they stopped being afraid. Whether that is good for the web is a separate question, and the answer is probably no.</p>]]></content:encoded></item><item><title>Build 2025: Microsoft open-sources WSL and puts MCP in Windows</title><link>https://readme.news/build-2025-microsoft-open-sources-wsl-and-puts-mcp-in-windows/</link><guid isPermaLink="true">https://readme.news/build-2025-microsoft-open-sources-wsl-and-puts-mcp-in-windows/</guid><pubDate>Tue, 20 May 2025 09:00:00 +0000</pubDate><description>A decade-old request granted, a coding agent in GitHub, and an agent protocol shipped at the OS layer.</description><content:encoded><![CDATA[<p>Microsoft Build ran this week. Three announcements matter to developers and one of them has been requested since 2016.</p>
<h2 id="wsl-is-open-source">WSL is open source<a class="anchor" href="#wsl-is-open-source" aria-label="link to this section">#</a></h2>
<p>The Windows Subsystem for Linux is now open source at <code>github.com/microsoft/WSL</code>. Not all of it — some Windows-side components remain closed — but the substantial majority, including the init system, the networking daemons, and the file sharing layer.</p>
<p>This was the single most requested item on the WSL issue tracker for years. The practical effect is that people can finally debug the parts of WSL that go wrong, which historically meant filing an issue and waiting.</p>
<p>The strategic read: WSL won. It is how a very large number of developers run Linux tooling on a Windows machine, it removed most of the reason to dual-boot, and open-sourcing it costs Microsoft nothing now that the position is secure. It is a mature move from a company that spent the nineties doing the opposite.</p>
<h2 id="github-copilot-coding-agent">GitHub Copilot coding agent<a class="anchor" href="#github-copilot-coding-agent" aria-label="link to this section">#</a></h2>
<p>Copilot gets a delegated mode: assign an issue to it, and it opens a pull request. It runs in GitHub Actions, follows repository instructions in <code>.github/copilot-instructions.md</code>, and its pull requests go through normal review and CI.</p>
<p>Notably it cannot approve its own pull requests and its Actions runs require approval. Those are the correct guardrails and it is good that they shipped with them rather than after an incident.</p>
<p>The integration point is the interesting bit. An agent that operates on issues and pull requests slots into a workflow every team already has, rather than requiring a new one. That is a much lower adoption barrier than "install this IDE."</p>
<h2 id="mcp-in-windows">MCP in Windows<a class="anchor" href="#mcp-in-windows" aria-label="link to this section">#</a></h2>
<p>Windows gets native support for the <a class="xref" href="/openai-adopts-mcp-and-a-protocol-becomes-a-standard/" title="OpenAI adopts MCP, and a protocol becomes a standard">Model Context</a> Protocol: an MCP registry, built-in servers exposing filesystem, windowing, and shell capabilities, and a <a class="xref" href="/nodejs-24-and-the-slow-reinvention-of-the-runtime/" title="Node.js 24 and the slow reinvention of the runtime">permission model</a> gating what an agent can reach.</p>
<p>An operating system vendor adopting a protocol that a model company published six months ago is fast by any standard. It also tells you Microsoft has concluded that the agent-to-system <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> is going to be standardized and would rather be the platform for it than fight it.</p>
<p>The <a class="xref" href="/the-component-model-and-the-plugin-problem/" title="The component model and the plugin problem">security model</a> is where this gets interesting and where I would like to see more detail. An OS-level MCP surface means any agent with the right permission can drive the machine. The consent flow, the scoping granularity, and the audit trail are the whole ballgame, and "there is a permission model" is not yet enough information to evaluate it.</p>
<h2 id="also-announced">also announced<a class="anchor" href="#also-announced" aria-label="link to this section">#</a></h2>
<p><strong>NLWeb</strong>, a project to let websites expose a natural-language query interface with MCP underneath. Positioned as "HTML for the agentic web." Ambitious; adoption is everything and adoption is unknowable.</p>
<p><strong>Edit</strong>, a new command-line text editor for Windows, open source, small. A genuine gap — Windows has had no default CLI editor since <code>edit.com</code> — and filling it is a nice piece of housekeeping.</p>
<p><strong>Windows AI Foundry</strong>, the rebranded local model runtime story, with Foundry Local for running models <a class="xref" href="/pixel-10-and-the-on-device-model-as-a-platform-feature/" title="Pixel 10 and the on-device model as a platform feature">on-device</a>.</p>
<h2 id="the-pattern">the pattern<a class="anchor" href="#the-pattern" aria-label="link to this section">#</a></h2>
<p>Microsoft's developer strategy for a decade has been: meet developers where they are, adopt other people's standards, open-source the layers where control is not worth the friction, and monetize the cloud underneath.</p>
<p>It keeps working. Build 2025 was more of it, executed well.</p>]]></content:encoded></item><item><title>Codex, and the agent that opens pull requests</title><link>https://readme.news/codex-and-the-agent-that-opens-pull-requests/</link><guid isPermaLink="true">https://readme.news/codex-and-the-agent-that-opens-pull-requests/</guid><pubDate>Mon, 19 May 2025 09:00:00 +0000</pubDate><description>OpenAI ships a cloud software engineering agent that works in a sandbox and hands you a diff.</description><content:encoded><![CDATA[<p>OpenAI released Codex as a <a class="xref" href="/operator-and-the-long-road-to-an-agent-that-can-click/" title="Operator, and the long road to an agent that can click">research preview</a>: a cloud-based software engineering agent that runs in an isolated container with your repository, works on a task for up to some tens of minutes, and produces a diff with a log of what it did.</p>
<p>This is a different product shape from the coding assistants we have been using, and the difference is worth being precise about.</p>
<h2 id="the-three-shapes">the three shapes<a class="anchor" href="#the-three-shapes" aria-label="link to this section">#</a></h2>
<p><strong>Completion.</strong> The model suggests the next few lines as you type. Copilot's original form. Latency budget: milliseconds. You stay in control of everything.</p>
<p><strong>Conversational.</strong> You describe a change, the model proposes an edit, you accept or reject. Latency budget: seconds. You review each step.</p>
<p><strong>Delegated.</strong> You describe a task, close the tab, and come back to a pull request. Latency budget: minutes to tens of minutes. You review the result, not the process.</p>
<p>Codex is firmly the third. So is Claude Code in its non-interactive mode, so is Devin, and so is what GitHub is building into Copilot. This is where the category is going.</p>
<h2 id="why-the-shape-matters">why the shape matters<a class="anchor" href="#why-the-shape-matters" aria-label="link to this section">#</a></h2>
<p>Delegated agents change the unit of work from "edit" to "task," and that changes everything downstream:</p>
<ul><li><strong>You cannot course-correct mid-flight.</strong> If the agent misunderstood the task, you find out after twenty minutes of work. That makes the task description vastly more important than a prompt in a chat.</li><li><strong>Parallelism becomes free.</strong> Five tasks running at once is the same wall-clock as one. That is a genuine multiplier and it is the actual value proposition.</li><li><strong>Review becomes the bottleneck.</strong> If an agent produces five pull requests an hour and you can meaningfully review two, you have not multiplied throughput. You have created a queue.</li></ul>
<p>That last point is where I think most teams are going to struggle. The constraint on software delivery in most organizations was never typing speed. It was understanding, coordination, and review. Delegated agents attack the part that was not the bottleneck.</p>
<h2 id="the-sandbox-is-the-good-part">the sandbox is the good part<a class="anchor" href="#the-sandbox-is-the-good-part" aria-label="link to this section">#</a></h2>
<p>Codex runs in a container with no network access during execution — dependencies are preloaded, then the network is cut. That is a meaningful security design and more products should copy it.</p>
<p>The threat model: an agent that can read your repository and reach the internet can exfiltrate your repository. It does not have to be malicious; a <a class="xref" href="/prompt-injection-is-sql-injection-without-the-fix/" title="Prompt injection is SQL injection without the fix">prompt injection</a> in a dependency's README is enough. Cutting network access after setup eliminates a whole class of attack for the cost of some inconvenience.</p>
<h2 id="the-practical-guidance">the practical guidance<a class="anchor" href="#the-practical-guidance" aria-label="link to this section">#</a></h2>
<p>For delegated agents to be worth it, you need:</p>
<ol><li><strong>Tasks with clear acceptance criteria.</strong> "Fix the flaky test in <code>test_payments.py::test_retry</code>" works. "Improve the payment system" does not.</li><li><strong>A test suite that actually gates.</strong> The agent's self-verification is only as good as your tests. If your tests pass on broken code, you will get broken code that passes tests.</li><li><strong>A review culture that has not been eroded.</strong> See above. This is the hard part and it is organizational, not technical.</li></ol>
<p>Start with the tasks you have been putting off: dependency upgrades, test coverage on a neglected module, migrating a deprecated API call across a hundred files. Mechanical, verifiable, tedious. That is where this technology is already clearly worth it, today, without any argument about whether it will replace anyone.</p>]]></content:encoded></item><item><title>OpenAI adopts MCP, and a protocol becomes a standard</title><link>https://readme.news/openai-adopts-mcp-and-a-protocol-becomes-a-standard/</link><guid isPermaLink="true">https://readme.news/openai-adopts-mcp-and-a-protocol-becomes-a-standard/</guid><pubDate>Wed, 26 Mar 2025 09:00:00 +0000</pubDate><description>Anthropic&#x27;s Model Context Protocol gets its most important endorsement four months after release.</description><content:encoded><![CDATA[<p>OpenAI announced support for the Model Context Protocol across its products today. Sam Altman posted about it. Four months after Anthropic open-sourced MCP, its primary competitor adopted it.</p>
<p>That is the moment a protocol stops being a vendor's format and starts being infrastructure.</p>
<h2 id="what-mcp-is-briefly">what MCP is, briefly<a class="anchor" href="#what-mcp-is-briefly" aria-label="link to this section">#</a></h2>
<p>A protocol for connecting language models to external context and tools. A server exposes three primitive types:</p>
<ul><li><strong>Resources</strong> — data the model can read (files, database rows, API responses).</li><li><strong>Tools</strong> — functions the model can call, with JSON Schema parameters.</li><li><strong>Prompts</strong> — reusable templates the user can invoke.</li></ul>
<p>The transport is JSON-RPC 2.0 over stdio for local servers or HTTP with server-sent events for remote ones. That is deliberately unexciting; the value is in the shape of the <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>, not the wire format.</p>
<div class="code"><span class="code-lang">json</span><pre><code class="lang-json">{
  "jsonrpc": "2.0", "id": 3, "method": "tools/call",
  "params": {
    "name": "query_db",
    "arguments": { "sql": "select count(*) from orders where day = '2025-03-25'" }
  }
}</code></pre></div>
<h2 id="why-this-mattered-enough-to-win">why this mattered enough to win<a class="anchor" href="#why-this-mattered-enough-to-win" aria-label="link to this section">#</a></h2>
<p>The problem MCP solves is combinatorial. Before it, connecting <code>M</code> AI clients to <code>N</code> data sources meant <code>M × N</code> bespoke integrations, each one a slightly different function-calling schema with slightly different auth. Every new assistant had to rebuild every connector.</p>
<p>MCP makes it <code>M + N</code>. Write one server for your internal ticketing system and every MCP-speaking client can use it. That is the same argument that won for LSP in editors, and it won for the same reason: the integration burden was crushing the ecosystem and everyone knew it.</p>
<p>The LSP comparison is not incidental — MCP's designers cite it explicitly, and the JSON-RPC choice is a direct inheritance.</p>
<h2 id="the-security-part-which-is-underdiscussed">the security part, which is underdiscussed<a class="anchor" href="#the-security-part-which-is-underdiscussed" aria-label="link to this section">#</a></h2>
<p>An MCP server is a program you run that a model can invoke. That is a substantially larger attack surface than it appears.</p>
<ul><li><strong><a class="xref" href="/prompt-injection-is-sql-injection-without-the-fix/" title="Prompt injection is SQL injection without the fix">Prompt injection</a> through resources.</strong> If a server returns content from an untrusted source, that content is now in the model's context and can attempt to influence tool calls. This is the central unsolved problem of the entire agent category and MCP does not fix it.</li><li><strong>Tool description poisoning.</strong> The tool descriptions themselves go into the model's context. A malicious server can write a description that manipulates behavior.</li><li><strong>Over-broad servers.</strong> A filesystem server with root access is a filesystem server with root access. The convenience of "just point it at my home directory" is how this goes wrong.</li></ul>
<p>Practical guidance while the ecosystem matures: run servers with the narrowest possible scope, treat any server you did not write as untrusted code, and do not combine a server that reads untrusted content with a server that can take destructive action in the same session. That last one is the composition hazard and it is easy to walk into.</p>
<h2 id="what-happens-next">what happens next<a class="anchor" href="#what-happens-next" aria-label="link to this section">#</a></h2>
<p>Expect: an official governance structure, a registry, and a fight about authorization semantics. Expect every developer tool company to ship a server within a quarter. Expect at least one significant security incident involving a popular community server, because that is what happens to every successful plugin ecosystem, without exception.</p>
<p>Standards win when the alternative is worse for everyone including the people who would prefer to own the standard. That happened here, unusually fast.</p>]]></content:encoded></item><item><title>The eight git commands worth actually learning</title><link>https://readme.news/the-eight-git-commands-worth-actually-learning/</link><guid isPermaLink="true">https://readme.news/the-eight-git-commands-worth-actually-learning/</guid><pubDate>Fri, 07 Mar 2025 09:00:00 +0000</pubDate><description>Not a tutorial. A short list of the operations that separate people who fight git from people who use it.</description><content:encoded><![CDATA[<p>Most developers use about six git commands and route around everything else by deleting the repository and cloning it again. That works. It is also leaving an enormous amount of leverage on the table for maybe two hours of investment.</p>
<p>Here is the short list. Not the complete list — the list where each item changes how you work.</p>
<h2 id="1-git-rebase-i">1. <code>git rebase -i</code><a class="anchor" href="#1-git-rebase-i" aria-label="link to this section">#</a></h2>
<p>The one that matters most. Interactive rebase lets you rewrite your local history before anyone sees it: squash the four "fix typo" commits, reorder, split a commit that did two things, reword a message you wrote badly at 6 p.m.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">git rebase -i main   # everything since you branched</code></pre></div>
<p>The rule that makes this safe: rewrite history that only exists on your machine or your own branch. Never rewrite a shared branch. That is the entire safety model and it is not complicated.</p>
<h2 id="2-git-add-p">2. <code>git add -p</code><a class="anchor" href="#2-git-add-p" aria-label="link to this section">#</a></h2>
<p>Stage hunks instead of files. This changes how you commit, because it removes the excuse for the giant mixed commit. You made an unrelated fix while working? Stage it separately, commit it separately.</p>
<p><code>y</code> to stage, <code>n</code> to skip, <code>s</code> to split the hunk smaller, <code>e</code> to edit it by hand.</p>
<h2 id="3-git-log-s">3. <code>git log -S</code><a class="anchor" href="#3-git-log-s" aria-label="link to this section">#</a></h2>
<p>Search history by <em>code</em>, not by message. "When did this string get introduced?" "Who deleted this function?"</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">git log -S "retryPolicy" --oneline
git log -S "retryPolicy" -p -- src/http.ts    # with diffs</code></pre></div>
<p>This is the single fastest way to answer "why is this here," and it beats <code>git blame</code> because blame only shows you the last change to a line.</p>
<h2 id="4-git-bisect">4. <code>git bisect</code><a class="anchor" href="#4-git-bisect" aria-label="link to this section">#</a></h2>
<p>Binary search for the commit that broke something. It feels like a niche tool until the first time it saves you a day.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">git bisect start
git bisect bad                 # current is broken
git bisect good v2.3.0         # this was fine
# git checks out a midpoint; test it; say good or bad
git bisect run ./test.sh       # or automate it entirely</code></pre></div>
<p>That last form is the good one. Give it a script that exits nonzero on failure and it finds the culprit unattended.</p>
<h2 id="5-git-reflog">5. <code>git reflog</code><a class="anchor" href="#5-git-reflog" aria-label="link to this section">#</a></h2>
<p>The undo button for everything. Every position HEAD has occupied for the last ninety days, including states unreachable from any branch.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">git reflog
git reset --hard HEAD@{4}</code></pre></div>
<p>Deleted a branch, botched a rebase, hard-reset over uncommitted work you meant to keep? Reflog. The number of people who have re-cloned a repository over something reflog would have fixed in ten seconds is very large.</p>
<h2 id="6-git-worktree">6. <code>git worktree</code><a class="anchor" href="#6-git-worktree" aria-label="link to this section">#</a></h2>
<p>Multiple branches checked out simultaneously in different directories, sharing one object store.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">git worktree add ../myrepo-hotfix hotfix/urgent</code></pre></div>
<p>No stashing, no switching, no rebuilding your entire dependency tree because you had to jump to a different branch for one review. This has become dramatically more useful in the age of coding agents, where you may want two working copies being edited at once.</p>
<h2 id="7-git-rerere">7. <code>git rerere</code><a class="anchor" href="#7-git-rerere" aria-label="link to this section">#</a></h2>
<p>"Reuse recorded resolution." Turn it on once and git remembers how you resolved a given conflict, then replays that resolution automatically the next time the same conflict appears.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">git config --global rerere.enabled true</code></pre></div>
<p>If you maintain a long-lived branch that rebases repeatedly, this saves you from resolving the same three conflicts every single time.</p>
<h2 id="8-git-commit-fixup-and-autosquash">8. <code>git commit --fixup</code> and <code>--autosquash</code><a class="anchor" href="#8-git-commit-fixup-and-autosquash" aria-label="link to this section">#</a></h2>
<p>Found a problem in an earlier commit on your branch? Do not amend it out of order and do not write "fix review comment."</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">git commit --fixup=a1b2c3d
git rebase -i --autosquash main   # squashes it into a1b2c3d automatically</code></pre></div>
<h2 id="the-meta-lesson">the meta-lesson<a class="anchor" href="#the-meta-lesson" aria-label="link to this section">#</a></h2>
<p>Git's model is small: commits are immutable snapshots, branches are pointers, everything is content-addressed. The porcelain is inconsistent and the <a class="xref" href="/error-messages-are-a-user-interface/" title="Error messages are a user interface">error messages</a> are famously unhelpful, but the underlying model is simple enough to hold in your head completely.</p>
<p>Spend an afternoon learning the model rather than memorizing commands, and every one of the above becomes obvious rather than magic.</p>]]></content:encoded></item><item><title>Claude 3.7 Sonnet ships, and so does a terminal</title><link>https://readme.news/claude-37-sonnet-ships-and-so-does-a-terminal/</link><guid isPermaLink="true">https://readme.news/claude-37-sonnet-ships-and-so-does-a-terminal/</guid><pubDate>Mon, 24 Feb 2025 09:00:00 +0000</pubDate><description>A hybrid reasoning model plus a command-line coding agent in research preview. The CLI is the more interesting release.</description><content:encoded><![CDATA[<p>Anthropic released Claude 3.7 Sonnet today alongside Claude Code, a coding agent that runs in your terminal, as a <a class="xref" href="/operator-and-the-long-road-to-an-agent-that-can-click/" title="Operator, and the long road to an agent that can click">research preview</a>.</p>
<h2 id="the-model">the model<a class="anchor" href="#the-model" aria-label="link to this section">#</a></h2>
<p>3.7 Sonnet is a hybrid reasoning model: one model that can answer immediately or think first, with the <a class="xref" href="/gemini-25-goes-generally-available-with-a-thinking-dial/" title="Gemini 2.5 goes generally available with a thinking dial">thinking budget</a> controllable via the API. That is a different product shape from having a separate reasoning model, and it is the right one — the routing decision belongs to the caller, who knows whether this particular request is worth the latency.</p>
<p>The API exposes a token budget for extended thinking. You set it per request:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">message = client.messages.create(
    model="claude-3-7-sonnet-20250219",
    max_tokens=8000,
    thinking={"type": "enabled", "budget_tokens": 4000},
    messages=[{"role": "user", "content": "..."}],
)</code></pre></div>
<p>The coding numbers are strong, particularly on agentic software engineering benchmarks where a model has to navigate a real repository rather than complete a function in isolation. That distinction is the one that matters for actual work and it is the one most benchmark discourse ignores.</p>
<h2 id="the-terminal-thing">the terminal thing<a class="anchor" href="#the-terminal-thing" aria-label="link to this section">#</a></h2>
<p>Claude Code is more interesting than the model, because it is a bet on a particular shape of tool that runs against the industry's current instinct.</p>
<p>Everyone else is putting the agent in the IDE. Anthropic put it in the terminal, with direct filesystem access, the ability to run commands, and git integration. No editor plugin, no separate UI, no sidebar.</p>
<p>The argument for the terminal is that it is the lowest common denominator that already has everything: your shell history, your credentials, your build tools, your test runner, your deployment scripts. An agent in the terminal inherits your entire working environment for free. An agent in an IDE inherits the IDE's model of your project, which is always partial.</p>
<p>The argument against is that terminals are a bad UI for reviewing a multi-file diff, and a bad UI for anything requiring a mental model of parallel state.</p>
<p>Both are correct. My guess is the terminal wins for the "do this task" workflow and the IDE wins for the "help me while I work" workflow, and in two years both exist and nobody thinks this was ever a debate.</p>
<h2 id="the-part-to-be-careful-about">the part to be careful about<a class="anchor" href="#the-part-to-be-careful-about" aria-label="link to this section">#</a></h2>
<p>An agent with shell access on your development machine is a genuinely different security posture than an autocomplete. It can <code>rm</code>. It can <code>curl | sh</code>. It can commit and push.</p>
<p>The mitigations that matter, in order:</p>
<ol><li>Run it in a container or VM for anything you did not write.</li><li>Do not give it credentials it does not need. Especially not production ones.</li><li>Read the diff before you commit. Every time. The moment you stop reading diffs is the moment the tool becomes a liability.</li></ol>
<p>That third one is the hard one, because the whole value proposition is not having to. The discipline that keeps this useful is treating agent output exactly like a pull request from a fast, capable, slightly overconfident junior engineer: worth reviewing, usually right, occasionally catastrophically wrong in a way that looks fine.</p>]]></content:encoded></item><item><title>Your CI is the slowest developer on the team</title><link>https://readme.news/your-ci-is-the-slowest-developer-on-the-team/</link><guid isPermaLink="true">https://readme.news/your-ci-is-the-slowest-developer-on-the-team/</guid><pubDate>Tue, 04 Feb 2025 09:00:00 +0000</pubDate><description>A twenty-minute pipeline doesn&#x27;t cost twenty minutes. It costs the context switch, the batching, and the review you skipped.</description><content:encoded><![CDATA[<p>There is a number in your organization that nobody owns and everybody pays: the wall-clock time between pushing a commit and knowing whether it is good.</p>
<p>Call it T. If T is two minutes, your team behaves one way. If T is thirty minutes, it behaves completely differently, and the difference is not "things take twenty-eight minutes longer."</p>
<h2 id="what-actually-happens-as-t-grows">what actually happens as T grows<a class="anchor" href="#what-actually-happens-as-t-grows" aria-label="link to this section">#</a></h2>
<p><strong>Below ~2 minutes</strong>, you wait. You keep the change in your head, you see the result, you fix it. The loop is tight enough that debugging is interactive.</p>
<p><strong>Between 2 and 10 minutes</strong>, you context switch. You go read something else. When the result comes back you have to page the change back in. The reload cost is real and it is roughly proportional to how complex the change was — which means it is worst exactly when you can least afford it.</p>
<p><strong>Past 10 minutes</strong>, behavior changes qualitatively. People start batching. They stop pushing small commits because the feedback is too expensive per commit, so they push bigger ones, which are harder to review and more likely to fail, which makes each failure more expensive to diagnose. It is a doom loop and it is entirely emergent from one number.</p>
<p><strong>Past 30 minutes</strong>, people route around CI. They test locally in ways that diverge from CI, they merge on green-ish, they add <code>[skip ci]</code>, and eventually somebody proposes a nightly build as the "real" signal, at which point you have given up.</p>
<h2 id="the-diagnostic">the diagnostic<a class="anchor" href="#the-diagnostic" aria-label="link to this section">#</a></h2>
<p>Before optimizing, measure the right thing. Not average pipeline duration — <strong>p90 time-to-signal for the check that actually blocks merge</strong>. Most teams measure the wrong thing here and optimize a stage nobody was waiting on.</p>
<p>Then break it down:</p>
<div class="code"><pre><code>queue wait      ← runners are saturated or your concurrency limit is wrong
checkout        ← you're doing a full clone of a repo with 200k commits
dependency install ← almost always the biggest fixable chunk
build           ← is anything cached? really?
test            ← is it parallel? is it parallel *well*?</code></pre></div>
<h2 id="the-fixes-roughly-in-order-of-return">the fixes, roughly in order of return<a class="anchor" href="#the-fixes-roughly-in-order-of-return" aria-label="link to this section">#</a></h2>
<ol><li><strong>Cache dependencies properly.</strong> Not "we have a cache step." Verify the hit rate. A cache that misses on every branch because the key includes the branch name is worse than no cache, because it also uploads.</li><li><strong>Shallow clone.</strong> <code>fetch-depth: 1</code> unless you genuinely need history. If you need history for one job, do it in one job.</li><li><strong>Split the blocking check from the exhaustive check.</strong> Lint, typecheck, and unit tests block merge. Integration, e2e, and the full matrix run after, and page someone if they fail. Not everything needs to gate.</li><li><strong>Parallelize by timing, not by file count.</strong> Splitting tests evenly by count gives you one shard that takes four times as long as the others. Split by historical duration.</li><li><strong>Kill the flaky tests.</strong> A test that fails 2% of the time in a suite of 50 parallel jobs means your pipeline fails constantly. Quarantine them the day they are identified. A quarantined test is a bug ticket; a flaky test in the blocking path is a tax on everyone forever.</li><li><strong>Bigger runners.</strong> This is the least intellectually satisfying fix and often the best one. Engineer time costs more than compute. Do the arithmetic before you spend a sprint on a clever solution.</li></ol>
<h2 id="the-cultural-part">the cultural part<a class="anchor" href="#the-cultural-part" aria-label="link to this section">#</a></h2>
<p>Someone has to own T. Not "the <a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">platform team</a> should look at CI sometime." A named person, a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>, and a number that goes in the same review as uptime.</p>
<p>Because the alternative is that it degrades a little every sprint — one more test, one more dependency, one more required check — and nobody notices until it is thirty minutes and everybody has quietly stopped trusting it.</p>]]></content:encoded></item><item><title>Bun 1.2 stops being a runtime and starts being a platform</title><link>https://readme.news/bun-12-stops-being-a-runtime-and-starts-being-a-platform/</link><guid isPermaLink="true">https://readme.news/bun-12-stops-being-a-runtime-and-starts-being-a-platform/</guid><pubDate>Fri, 17 Jan 2025 09:00:00 +0000</pubDate><description>Native S3 and Postgres clients, a real Node compatibility push, and a bundled package manager that keeps getting faster.</description><content:encoded><![CDATA[<p>Bun 1.2 landed this week and the release is a good moment to notice what the project has actually become. It started as "a fast JavaScript runtime." It is now attempting to be the entire server-side JavaScript toolchain in one binary, and the 1.2 feature list is unsubtle about it.</p>
<h2 id="whats-new">what's new<a class="anchor" href="#whats-new" aria-label="link to this section">#</a></h2>
<p><strong>A built-in S3 client.</strong> <code>Bun.s3</code> gives you presigned URLs, streaming reads and writes, and it works against any S3-compatible endpoint — R2, MinIO, Backblaze, the actual thing.</p>
<div class="code"><span class="code-lang">js</span><pre><code class="lang-js">import { s3 } from "bun";

const file = s3.file("uploads/report.pdf");
await file.write(pdfBytes, { type: "application/pdf" });
const url = file.presign({ expiresIn: 3600 });</code></pre></div>
<p><strong>A built-in Postgres client.</strong> <code>Bun.sql</code> is a tagged-template SQL <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> with <a class="xref" href="/connection-pooling-explained-properly/" title="Connection pooling, explained properly">connection pooling</a>, prepared statements, and parameter binding that does not require you to think about <code>$1</code> ordering.</p>
<div class="code"><span class="code-lang">js</span><pre><code class="lang-js">import { sql } from "bun";
const users = await sql`select * from users where team = ${teamId} limit 20`;</code></pre></div>
<p><strong>Node compatibility as a first-class metric.</strong> Bun now publishes its pass rate against Node's own test suite and treats regressions as bugs. That is a much more honest signal than a feature checklist, and the number went up substantially in this cycle.</p>
<p><strong>Text-based lockfile.</strong> <code>bun.lock</code> replaces the binary <code>bun.lockb</code> as the default. Reviewable in a pull request, diffable, mergeable. This was the single most common complaint about adopting Bun in a team setting and it is now gone.</p>
<h2 id="the-strategic-read">the strategic read<a class="anchor" href="#the-strategic-read" aria-label="link to this section">#</a></h2>
<p>Node's answer to the same problem has been to add capabilities to the runtime incrementally and carefully: a built-in test runner, <code>--watch</code>, a <a class="xref" href="/nodejs-24-and-the-slow-reinvention-of-the-runtime/" title="Node.js 24 and the slow reinvention of the runtime">permission model</a>, and now stable-ish TypeScript stripping. Deno's answer was to bundle everything from the start and then spend years walking back the parts of its opinionated design that made adoption hard.</p>
<p>Bun's answer is to bundle everything <em>and</em> be compatible with Node's ecosystem rather than replacing it. That is a harder engineering problem and a much easier sales problem. Nobody has to rewrite anything. You point <code>bun</code> at an existing Express app and it mostly runs.</p>
<h2 id="the-honest-caveats">the honest caveats<a class="anchor" href="#the-honest-caveats" aria-label="link to this section">#</a></h2>
<p>Bun's ecosystem edge cases still bite. Native modules that assume V8 internals will not work, because Bun is JavaScriptCore. Some observability agents assume Node. Long-tail packages that reach into <code>process.binding</code> or undocumented internals will surprise you.</p>
<p>The pattern for adoption in 2025 looks like: use Bun as the package manager and test runner immediately, because those are drop-in and dramatically faster; use the runtime in production when your dependency tree is one you actually understand.</p>
<p>That is not a hedge. That is just what shipping looks like.</p>]]></content:encoded></item>
</channel>
</rss>
