<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — devops</title>
<link>https://readme.news/tags/devops/</link>
<atom:link href="https://readme.news/tags/devops/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged devops.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>Shell scripts that outlive you</title><link>https://readme.news/shell-scripts-that-outlive-you/</link><guid isPermaLink="true">https://readme.news/shell-scripts-that-outlive-you/</guid><pubDate>Sat, 22 Aug 2026 09:00:00 +0000</pubDate><description>Six lines at the top of a bash script are the difference between a tool and a trap.</description><content:encoded><![CDATA[<p>Every codebase has a <code>scripts/</code> directory. Most of the files in it were written in ten minutes, work correctly on exactly one machine, and fail in ways that produce no error and no output.</p>
<p>Shell is a fine language for gluing programs together. It is a terrible language for doing it <em>safely</em> unless you tell it to be, and telling it to be takes six lines.</p>
<h2 id="the-preamble">the preamble<a class="anchor" href="#the-preamble" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">#!/usr/bin/env bash
set -euo pipefail
IFS=$'\n\t'</code></pre></div>
<p>What each one prevents:</p>
<p><strong><code>set -e</code></strong> — exit on any command that fails. Without it, a script continues merrily after <code>cd /nonexistent</code> and then runs the rest of its commands in the wrong directory. This is how a cleanup script deletes the wrong thing.</p>
<p><strong><code>set -u</code></strong> — error on an undefined variable. Without it, <code>rm -rf "$BUILD_DIR/"</code> with an unset <code>BUILD_DIR</code> expands to <code>rm -rf /</code>. This has happened to real people, at real companies, more than once.</p>
<p><strong><code>set -o pipefail</code></strong> — a pipeline fails if <em>any</em> stage fails, not just the last one. Without it, <code>curl bad-url | jq .</code> succeeds, because <code>jq</code> was happy with the empty input.</p>
<p><strong><code>IFS=$'\n\t'</code></strong> — stop splitting on spaces. This is what makes filenames with spaces stop being a source of bugs.</p>
<p>Four lines. They convert an entire category of silent wrong behaviour into loud failure.</p>
<h2 id="the-next-four-things">the next four things<a class="anchor" href="#the-next-four-things" aria-label="link to this section">#</a></h2>
<p><strong>Quote every expansion.</strong> <code>"$var"</code>, not <code>$var</code>. Always, including inside <code>[[ ]]</code> where it usually does not matter, because "usually" is not a rule anyone remembers correctly.</p>
<p><strong>Use <code>"${var:?message}"</code> for required inputs.</strong></p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">: "${DATABASE_URL:?DATABASE_URL is required}"</code></pre></div>
<p>One line, fails immediately with a useful message rather than three steps later with a confusing one.</p>
<p><strong>Make it idempotent.</strong> A script that is safe to run twice is a script that is safe to run at all. <code>mkdir -p</code>, <code>rm -f</code>, check-before-create. The second run is the one that happens during an incident when nobody is sure whether the first one worked.</p>
<p><strong>Clean up with <code>trap</code>.</strong></p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT</code></pre></div>
<p>Now the temporary directory is removed whether the script succeeds, fails, or is interrupted.</p>
<h2 id="the-usability-part">the usability part<a class="anchor" href="#the-usability-part" aria-label="link to this section">#</a></h2>
<p>A script that other people run needs the same courtesy as any other <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">usage() {
  cat &lt;&lt;'EOF'
Usage: deploy.sh ENVIRONMENT [--dry-run]

  ENVIRONMENT   staging | production
  --dry-run     print what would happen, change nothing

Requires: awscli &gt;= 2, jq. Reads DEPLOY_ROLE from the environment.
EOF
}
[[ $# -eq 0 || "${1:-}" == "-h" ]] &amp;&amp; { usage; exit 0; }</code></pre></div>
<p><strong>Add a dry-run mode to anything destructive.</strong> It costs one conditional and it is the difference between a script people trust and a script people read three times before running.</p>
<p><strong>Echo what you are about to do.</strong> <code>set -x</code> is the crude version and it is better than silence. A script that prints "deleting 4 objects from s3://bucket/prefix/" before doing it lets a human catch the mistake.</p>
<h2 id="when-to-stop-using-shell">when to stop using shell<a class="anchor" href="#when-to-stop-using-shell" aria-label="link to this section">#</a></h2>
<p>Shell is right for: calling other programs in sequence, moving files, gluing a pipeline together. Under about a hundred lines.</p>
<p>Switch to a real language when you need:</p>
<ul><li><strong>Data structures.</strong> Bash arrays are a trap and associative arrays are worse.</li><li><strong>Any arithmetic beyond counting.</strong></li><li><strong>Error handling with recovery</strong>, rather than exit-on-failure.</li><li><strong>Parsing anything structured.</strong> If you are pulling JSON apart with <code>sed</code>, stop.</li><li><strong>Tests.</strong> You can test shell, and almost nobody does, which tells you something.</li></ul>
<p>Python or Go for anything past that line. The rewrite is an hour and it pays back the first time somebody has to change it.</p>
<h2 id="the-check-that-costs-nothing">the check that costs nothing<a class="anchor" href="#the-check-that-costs-nothing" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">shellcheck scripts/*.sh</code></pre></div>
<p>Run it in CI. It catches unquoted expansions, useless <code>cat</code>, subshell variable scoping, and roughly a dozen other things that produce silent wrong behaviour.</p>
<p>Every shell script I have ever run it against had at least one real finding. It takes five minutes to add and it is the highest-value linting available for the least-linted language in most repositories.</p>]]></content:encoded></item><item><title>Vendor lock-in: an honest cost model</title><link>https://readme.news/vendor-lock-in-an-honest-cost-model/</link><guid isPermaLink="true">https://readme.news/vendor-lock-in-an-honest-cost-model/</guid><pubDate>Fri, 17 Jul 2026 09:00:00 +0000</pubDate><description>Portability is not free and neither is dependence. A framework for deciding how much abstraction to buy.</description><content:encoded><![CDATA[<p>"Avoid vendor lock-in" is treated as self-evidently good advice. It is not advice, it is a preference, and following it uncritically produces systems that are worse in exchange for optionality nobody will ever exercise.</p>
<p>Here is a way to actually decide.</p>
<h2 id="the-cost-of-lock-in">the cost of lock-in<a class="anchor" href="#the-cost-of-lock-in" aria-label="link to this section">#</a></h2>
<p><strong>Switching cost.</strong> How much engineering time to move off, if you had to.</p>
<p><strong>Pricing power.</strong> A vendor who knows you cannot leave prices accordingly. This is real and it is usually the largest ongoing cost.</p>
<p><strong>Capability ceiling.</strong> You are limited to what they support, on their timeline.</p>
<p><strong>Correlated risk.</strong> They have an outage, you have an outage. They change their terms, you comply. They get acquired and sunset the product, you migrate on their schedule.</p>
<h2 id="the-cost-of-avoiding-lock-in">the cost of avoiding lock-in<a class="anchor" href="#the-cost-of-avoiding-lock-in" aria-label="link to this section">#</a></h2>
<p>This is the half that gets ignored, and it is frequently larger.</p>
<p><strong>The abstraction layer itself.</strong> Code to write, maintain, test, and debug through. It is a permanent tax and it makes stack traces longer.</p>
<p><strong>Lowest common denominator.</strong> Your abstraction can only expose what all candidate providers support. You give up the features that made the good option good.</p>
<p><strong>The abstraction is usually wrong anyway.</strong> It was designed against one provider's model. When you actually try to swap, you discover the abstraction encoded assumptions that do not hold, and you rewrite it.</p>
<p><strong>Delayed value.</strong> Time spent on portability is time not spent on the product.</p>
<h2 id="the-framework">the framework<a class="anchor" href="#the-framework" aria-label="link to this section">#</a></h2>
<p>For each dependency, estimate:</p>
<ol><li><strong>Switching cost</strong> — engineer-weeks to move.</li><li><strong>Probability you switch</strong> in the next three years.</li><li><strong>Cost of the abstraction</strong> — engineer-weeks now, plus ongoing drag.</li></ol>
<p>Then: if <code>switching_cost × probability &lt; abstraction_cost</code>, do not abstract.</p>
<p>The numbers are rough. The exercise still clarifies, because it forces you to state the probability out loud, and stated probabilities are usually much lower than the implied ones people are acting on.</p>
<h2 id="the-categories-worked-through">the categories, worked through<a class="anchor" href="#the-categories-worked-through" aria-label="link to this section">#</a></h2>
<p><strong>Object storage.</strong> Switching cost: low. The S3 API is a de facto standard and every provider implements it. Probability: moderate — people do move for pricing.</p>
<p><strong>Verdict: use the S3 API, do not abstract further.</strong> The API is already the abstraction.</p>
<p><strong>Compute.</strong> Switching cost: moderate if containerized, high if you use provider-specific serverless. Probability: low.</p>
<p><strong>Verdict: containerize</strong> — which is good practice anyway — <strong>and use whatever managed service you want.</strong> Do not build a compute abstraction layer.</p>
<p><strong>Relational database.</strong> Switching cost: high. Probability: low.</p>
<p><strong>Verdict: use the database's features.</strong> Teams that avoid stored procedures, database-specific types, and advanced indexing to stay portable are giving up real capability for an event that will not happen. Postgres-specific SQL is fine. You are not going to migrate to a different engine, and if you do, the SQL dialect will be the smallest part of the pain.</p>
<p><strong>Managed <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a>, streams, and similar.</strong> Switching cost: moderate. The semantics differ enough between providers that a thin abstraction genuinely helps.</p>
<p><strong>Verdict: a thin <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> — publish, subscribe, ack — is worth it.</strong> Not a full abstraction; a boundary.</p>
<p><strong>Authentication.</strong> Switching cost: very high — you have to migrate user credentials, sessions, and integrations. Probability: low, but the consequences of being stuck are severe.</p>
<p><strong>Verdict: use standard protocols.</strong> OIDC and SAML are the abstraction. A provider that supports them is replaceable in principle; one with a proprietary SDK is not.</p>
<p><strong>AI model providers.</strong> Switching cost: low if you kept the interface thin. Probability: <strong>high</strong> — the market is moving fast and the right choice changes quarterly.</p>
<p><strong>Verdict: definitely abstract.</strong> This is the clearest case on the list. A thin interface plus an eval harness in your own repository makes model changes an afternoon. Teams that did this in 2024 have been switching providers casually ever since.</p>
<p><strong>Observability.</strong> Switching cost: moderate to high — instrumentation is everywhere in your code. Probability: moderate, usually driven by cost.</p>
<p><strong>Verdict: OpenTelemetry.</strong> The instrumentation is vendor-neutral, the collector handles routing, and swapping backends is a configuration change.</p>
<h2 id="the-pattern">the pattern<a class="anchor" href="#the-pattern" aria-label="link to this section">#</a></h2>
<p><strong>Abstract where the switching probability is high and the abstraction is cheap.</strong> AI providers, observability backends, object storage.</p>
<p><strong>Do not abstract where switching is unlikely and the abstraction costs you real capability.</strong> Databases, compute platforms, managed services you chose for their specific features.</p>
<p><strong>Use standard protocols wherever they exist.</strong> OIDC, S3, OpenTelemetry, SQL. A standard is an abstraction someone else maintains, and it is always cheaper than yours.</p>
<h2 id="the-thing-that-actually-protects-you">the thing that actually protects you<a class="anchor" href="#the-thing-that-actually-protects-you" aria-label="link to this section">#</a></h2>
<p>Not an abstraction layer. <strong>Your data, in a format you can export, and a documented process for leaving.</strong></p>
<p>Ask, before adopting anything: can I get all my data out, in a usable format, without their cooperation? If yes, you have real optionality regardless of how coupled your code is, because the expensive part of a migration is never the code — it is the data.</p>
<p>If the answer is no, that is a much bigger red flag than any API coupling, and it is the question almost nobody asks during procurement.</p>]]></content:encoded></item><item><title>Why your container image is 1.4 gigabytes</title><link>https://readme.news/why-your-container-image-is-14-gigabytes/</link><guid isPermaLink="true">https://readme.news/why-your-container-image-is-14-gigabytes/</guid><pubDate>Fri, 10 Jul 2026 09:00:00 +0000</pubDate><description>It should be forty megabytes. Here is where the rest of it came from and how to get it back.</description><content:encoded><![CDATA[<p>A container image for a compiled service should be tens of megabytes. For an interpreted one, low hundreds. If yours is over a gigabyte, something specific went wrong and it is usually one of six things.</p>
<p>Size matters for real reasons: pull time on cold start, registry cost, deployment speed when you are scaling out under load, and attack surface — every package in the image is something that can have a CVE you have to answer for.</p>
<h2 id="the-six-causes">the six causes<a class="anchor" href="#the-six-causes" aria-label="link to this section">#</a></h2>
<p><strong>1. You shipped the build toolchain.</strong></p>
<p>The compiler, the headers, the package manager cache, the source tree, the test fixtures. All of it needed to build, none of it needed to run.</p>
<p>Multi-stage builds fix this completely:</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile">FROM golang:1.24 AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 go build -ldflags="-s -w" -o /app ./cmd/server

FROM gcr.io/distroless/static-debian12
COPY --from=build /app /app
ENTRYPOINT ["/app"]</code></pre></div>
<p>Final image: the binary, plus CA certificates and timezone data. Tens of megabytes.</p>
<p><strong>2. You started from a full distribution image.</strong></p>
<p><code>FROM ubuntu</code> is roughly 80 MB before you install anything, and it includes a package manager, a shell, and a hundred utilities you will never invoke.</p>
<p>The ladder, from largest to smallest:</p>
<ul><li>Full distribution — 80 MB+</li><li><code>-slim</code> variants — 30–80 MB</li><li>Alpine — 5–10 MB, with musl libc, which will occasionally surprise you</li><li>Distroless — just the runtime, no shell, no package manager</li><li><code>scratch</code> — nothing at all, for static binaries</li></ul>
<p><strong>The Alpine caveat</strong>, since it bites people: musl's allocator and DNS resolver behave differently from glibc's. Python performance in particular can be significantly worse, and some binary wheels do not exist for musl. Test rather than assume.</p>
<p><strong>3. Your layers are ordered wrong.</strong></p>
<p>Each instruction creates a layer. A layer is invalidated when it or anything before it changes.</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile"># bad — any source change reinstalls every dependency
COPY . .
RUN npm ci

# good — dependencies are cached until the lockfile changes
COPY package.json package-lock.json ./
RUN npm ci
COPY . .</code></pre></div>
<p>This does not shrink the final image but it dramatically speeds up builds, which is usually what people actually care about.</p>
<p><strong>4. You deleted things in a later layer.</strong></p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile">RUN apt-get install -y build-essential   # layer 1: +400 MB
RUN apt-get remove -y build-essential    # layer 2: marks deleted, image unchanged</code></pre></div>
<p>Layers are additive. Deleting a file in a later layer hides it and does not remove it. The bytes are still in the image and still transferred on pull.</p>
<p>Everything must happen in one <code>RUN</code>:</p>
<div class="code"><span class="code-lang">dockerfile</span><pre><code class="lang-dockerfile">RUN apt-get update \
 &amp;&amp; apt-get install -y --no-install-recommends build-essential \
 &amp;&amp; make \
 &amp;&amp; apt-get purge -y build-essential \
 &amp;&amp; apt-get autoremove -y \
 &amp;&amp; rm -rf /var/lib/apt/lists/*</code></pre></div>
<p>Better: use a multi-stage build and do not install the toolchain in the final image at all.</p>
<p><strong>5. You have no <code>.dockerignore</code>.</strong></p>
<p><code>COPY . .</code> copies <code>.git</code>, <code>node_modules</code>, build artifacts, test fixtures, and your local <code>.env</code>.</p>
<div class="code"><pre><code>.git
node_modules
dist
*.log
.env*
**/__pycache__
coverage</code></pre></div>
<p>The <code>.git</code> directory alone is frequently hundreds of megabytes on a mature repository, and it is in a lot of images.</p>
<p><strong>6. Your dependencies are enormous.</strong></p>
<p>Sometimes it is genuinely the dependencies — machine learning stacks with CUDA libraries are legitimately multiple gigabytes.</p>
<p>Check whether you need the GPU variant. <code>torch</code> with CUDA is roughly 2.5 GB; the CPU build is a fraction of that. If you are serving on CPU, you are shipping GPU libraries for nothing.</p>
<h2 id="finding-out-where-it-went">finding out where it went<a class="anchor" href="#finding-out-where-it-went" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">docker history --no-trunc &lt;image&gt;       # size per layer</code></pre></div>
<p>Or use a layer inspection tool that shows you which files are in which layer and how much space is wasted. Ten minutes with one of those tells you exactly what to fix.</p>
<h2 id="the-security-dimension">the security dimension<a class="anchor" href="#the-security-dimension" aria-label="link to this section">#</a></h2>
<p>Every package in the image is potential CVE surface, and your scanner will report all of them regardless of whether the code is reachable.</p>
<p>A distroless image has almost nothing to report, which means the reports you do get are signal rather than noise. That is worth more than the size reduction — a vulnerability report with three entries gets read; one with four hundred does not.</p>
<p>The trade-off: no shell means you cannot <code>docker exec</code> in to debug. Use ephemeral debug containers that attach to the running pod's namespaces instead, which is a better practice anyway because it means your production image is not a debugging toolkit.</p>
<h2 id="the-target">the target<a class="anchor" href="#the-target" aria-label="link to this section">#</a></h2>
<ul><li>Compiled language, static binary: <strong>under 30 MB.</strong></li><li>Interpreted with dependencies: <strong>under 200 MB.</strong></li><li>Anything over a gigabyte without a machine learning stack: something is wrong and it is one of the six above.</li></ul>]]></content:encoded></item><item><title>Migrations that don't wake anyone up</title><link>https://readme.news/migrations-that-dont-wake-anyone-up/</link><guid isPermaLink="true">https://readme.news/migrations-that-dont-wake-anyone-up/</guid><pubDate>Mon, 27 Apr 2026 09:00:00 +0000</pubDate><description>Data migrations at scale, done in a way that is reversible at every step and boring in the middle.</description><content:encoded><![CDATA[<p>A migration on a table with a thousand rows is a command. A migration on a table with two billion rows is a project, and the difference in approach is total.</p>
<p>Here is the pattern that works, and the specific things that go wrong.</p>
<h2 id="the-locks-that-kill-you">the locks that kill you<a class="anchor" href="#the-locks-that-kill-you" aria-label="link to this section">#</a></h2>
<p>The failure mode is almost always a lock held longer than expected, blocking every query behind it, until connections exhaust and the application falls over.</p>
<p><strong>In Postgres</strong>, these are safe (metadata-only, fast):</p>
<ul><li><code>ADD COLUMN</code> with no default, or with a non-volatile default (since PG 11).</li><li><code>DROP COLUMN</code> (marks it dead, does not rewrite).</li><li><code>ADD CONSTRAINT ... NOT VALID</code>, then <code>VALIDATE CONSTRAINT</code> separately.</li><li><code>CREATE INDEX CONCURRENTLY</code>.</li><li>Renaming things.</li></ul>
<p>These rewrite the table and hold an exclusive lock for the duration:</p>
<ul><li><code>ALTER COLUMN TYPE</code> (most of the time).</li><li><code>ADD COLUMN</code> with a volatile default.</li><li><code>SET NOT NULL</code> directly (use a <code>CHECK</code> constraint validated separately, then convert).</li></ul>
<p><strong>The one everyone hits:</strong> <code>CREATE INDEX</code> without <code>CONCURRENTLY</code> blocks writes for the entire build. On a large table that is minutes to hours. <code>CONCURRENTLY</code> takes longer and does not block, and it can fail — leaving an invalid index you must drop and retry.</p>
<p><strong>The one nobody expects:</strong> even a fast <code>ALTER TABLE</code> must acquire an exclusive lock, and it will queue behind any long-running transaction. Then everything <a class="xref" href="/the-queues-you-did-not-know-you-had/" title="The queues you did not know you had">queues</a> behind <em>it</em>. A migration that takes 5 ms can cause a 10-minute outage if a reporting query is holding a lock.</p>
<p>Set <code>lock_timeout</code> before every DDL statement:</p>
<div class="code"><span class="code-lang">sql</span><pre><code class="lang-sql">SET lock_timeout = '3s';
ALTER TABLE orders ADD COLUMN region text;</code></pre></div>
<p>If it cannot get the lock in three seconds, it fails and you retry, rather than queueing behind something and taking the site down. This one line prevents a large share of migration incidents.</p>
<h2 id="expand-and-contract-always">expand and contract, always<a class="anchor" href="#expand-and-contract-always" aria-label="link to this section">#</a></h2>
<p>Never change a column in place. Four deploys:</p>
<p><strong>1. Expand.</strong> Add the new column, nullable. Deploy. Old code ignores it.</p>
<p><strong>2. Backfill and dual-write.</strong> Application writes both columns. Backfill existing rows in batches. Deploy.</p>
<div class="code"><span class="code-lang">sql</span><pre><code class="lang-sql">-- in a loop, with a pause between batches
UPDATE orders SET region = derive_region(country)
WHERE region IS NULL AND id BETWEEN $1 AND $2;</code></pre></div>
<p>Batch size in the thousands. Sleep between batches. Monitor replication lag and back off if it grows — a backfill that outruns replication is how you take down your read replicas.</p>
<p><strong>3. Switch reads.</strong> Read from the new column. Deploy. Still fully reversible, the old column is intact and still being written.</p>
<p><strong>4. Contract.</strong> Stop writing the old column, deploy, and drop it days later.</p>
<p>Four deploys instead of one. Every intermediate state works with both the old and new code, so any deploy can be rolled back independently.</p>
<h2 id="the-backfill-rules">the backfill rules<a class="anchor" href="#the-backfill-rules" aria-label="link to this section">#</a></h2>
<p><strong>Idempotent.</strong> It will be interrupted. It must be safe to re-run.</p>
<p><strong>Resumable.</strong> Track progress. A backfill that starts over from the beginning after a failure will never finish on a large table.</p>
<p><strong>Rate limited.</strong> Not "as fast as possible." A backfill competing with production traffic for I/O is a self-inflicted incident. Watch replication lag, watch p99 latency, and throttle.</p>
<p><strong>Observable.</strong> Log progress. "Backfilled 4.2M of 180M rows, ETA 6h" lets someone make a decision. Silence does not.</p>
<p><strong>Kill-switchable.</strong> You need to be able to stop it immediately without deploying. A flag it checks between batches.</p>
<h2 id="testing-it">testing it<a class="anchor" href="#testing-it" aria-label="link to this section">#</a></h2>
<p><strong>On a copy of production, at production size.</strong> A migration tested on a 10,000-row development database tells you the syntax is right and nothing about the runtime.</p>
<p><strong>With concurrent load.</strong> Locks only matter under contention. Run the migration against a restored copy while replaying production traffic.</p>
<p><strong>With the rollback.</strong> Actually run it. A rollback plan that has never been executed is a hypothesis.</p>
<h2 id="the-checklist">the checklist<a class="anchor" href="#the-checklist" aria-label="link to this section">#</a></h2>
<p>Before any production migration on a large table:</p>
<ul><li>[ ] Tested on production-sized data with concurrent load</li><li>[ ] <code>lock_timeout</code> set on every DDL statement</li><li>[ ] Expand-contract, not in-place</li><li>[ ] Backfill is batched, rate-limited, resumable, and killable</li><li>[ ] Rollback tested</li><li>[ ] Replication lag monitored with a threshold to pause at</li><li>[ ] Someone is watching, with the <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">kill switch</a>, for the duration</li><li>[ ] Not on a Friday</li></ul>
<p>That last one is not superstition. It is that the people who understand the change should be available for the two days after it, and on Friday they are not.</p>]]></content:encoded></item><item><title>You still do not need Kubernetes</title><link>https://readme.news/you-still-do-not-need-kubernetes/</link><guid isPermaLink="true">https://readme.news/you-still-do-not-need-kubernetes/</guid><pubDate>Mon, 13 Apr 2026 09:00:00 +0000</pubDate><description>It&#x27;s excellent software solving a real problem that most teams do not have. The honest threshold, and what to do below it.</description><content:encoded><![CDATA[<p>Kubernetes is genuinely good software. It solves a real problem well. It has an enormous ecosystem and a large pool of people who know it.</p>
<p>It is also, for a majority of the teams running it, a substantial amount of complexity in exchange for benefits they do not receive, and saying so is still mildly heretical.</p>
<h2 id="the-problem-it-actually-solves">the problem it actually solves<a class="anchor" href="#the-problem-it-actually-solves" aria-label="link to this section">#</a></h2>
<p>Kubernetes was built for: many services, many teams, heterogeneous workloads, on a fleet of machines, where you want bin-packing efficiency and declarative self-healing, and where the platform is operated by people whose job that is.</p>
<p>If you have all of those, it is the right answer and there is no close second.</p>
<h2 id="what-it-costs">what it costs<a class="anchor" href="#what-it-costs" aria-label="link to this section">#</a></h2>
<p><strong>A permanent learning tax.</strong> Pods, deployments, services, ingresses, configmaps, secrets, persistent volume claims, storage classes, service accounts, roles, network policies, resource quotas, and a YAML dialect for each. Every engineer who deploys anything must learn a meaningful fraction of it.</p>
<p><strong>Operational surface.</strong> Control plane upgrades, node upgrades, CNI plugin, CSI driver, ingress controller, cert manager, metrics server, log shipper. Each is a component that can break and that must be upgraded on someone else's schedule.</p>
<p><strong>Debugging distance.</strong> "Why is my service not reachable" has a dozen possible answers across five layers, and diagnosing it requires understanding all of them.</p>
<p><strong>Cost, frequently.</strong> A managed control plane plus nodes sized for the platform's own overhead plus the observability stack it needs is often more than the equivalent capacity on simpler infrastructure.</p>
<p><strong>Resume-driven adoption.</strong> This is real and worth naming. Kubernetes on your CV is worth money. That is a genuine incentive pointed away from the simplest solution that works.</p>
<h2 id="the-honest-threshold">the honest threshold<a class="anchor" href="#the-honest-threshold" aria-label="link to this section">#</a></h2>
<p>You probably want Kubernetes if:</p>
<ul><li>More than roughly fifteen to twenty distinct services, deployed independently.</li><li>More than a handful of teams that need to deploy without coordinating.</li><li>You have someone whose job includes operating the platform, not as a side task.</li><li>Genuinely heterogeneous workloads with different scaling characteristics.</li><li>Multi-tenancy requirements with real isolation needs.</li></ul>
<p>You probably do not if:</p>
<ul><li>Under ten services.</li><li>One or two teams.</li><li>Nobody owns the platform.</li><li>Traffic is predictable.</li><li>You are running one application with a database.</li></ul>
<h2 id="what-to-do-below-the-threshold">what to do below the threshold<a class="anchor" href="#what-to-do-below-the-threshold" aria-label="link to this section">#</a></h2>
<p>The options are better than they were, and all of them are boring:</p>
<p><strong>A platform-as-a-service.</strong> Push code, it runs. This is the correct answer for a very large number of applications and the reason people avoid it is usually aesthetic.</p>
<p><strong>Containers on a managed container service</strong> without the orchestrator — the various "run this container, scale it, load balance it" products every cloud offers. You get containers, autoscaling, and rolling deploys without the platform.</p>
<p><strong>A couple of servers and a process manager.</strong> Systemd units, a reverse proxy with automatic certificates, and a deploy script. This runs an enormous amount of traffic, is trivially debuggable, and every engineer already understands it.</p>
<p><strong>Docker Compose on one machine.</strong> For staging, for internal tools, for anything where a single host is enough. Unfashionable, works.</p>
<h2 id="the-migration-path-argument">the migration-path argument<a class="anchor" href="#the-migration-path-argument" aria-label="link to this section">#</a></h2>
<p>"We will need Kubernetes eventually, so we should start now."</p>
<p>This is the most common justification and it is usually wrong, for two reasons.</p>
<p>First, the complexity cost is paid every day from now until then, and "eventually" frequently never arrives.</p>
<p>Second, containerizing your application is the actual hard part of a future migration, and you can do that without an orchestrator. A containerized application running under a simple process manager can move to Kubernetes later in a couple of weeks.</p>
<p>Build the container. Skip the platform until you have the problem it solves.</p>
<h2 id="the-position-i-will-defend">the position I will defend<a class="anchor" href="#the-position-i-will-defend" aria-label="link to this section">#</a></h2>
<p>The teams I have seen most successfully run Kubernetes are large organizations with dedicated <a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">platform teams</a>, where it is genuinely the right tool.</p>
<p>The teams I have seen most damaged by it are small ones where a single engineer set it up, that engineer left, and the remaining team is operating a system nobody understands and is afraid to touch.</p>
<p>That second failure is common, expensive, and entirely predictable from the staffing at adoption time. If nobody's job is going to be operating the platform, you should not have a platform.</p>]]></content:encoded></item><item><title>Infrastructure as code, ten years of lessons</title><link>https://readme.news/infrastructure-as-code-ten-years-of-lessons/</link><guid isPermaLink="true">https://readme.news/infrastructure-as-code-ten-years-of-lessons/</guid><pubDate>Wed, 08 Apr 2026 09:00:00 +0000</pubDate><description>State files, drift, modules that became frameworks, and the one practice that separates teams that like their IaC from teams that don&#x27;t.</description><content:encoded><![CDATA[<p>Infrastructure as code won the argument. Nobody clicks through a console to provision production anymore, or admits to it.</p>
<p>What did not get settled is how to do it without producing something everyone hates. Here is what a decade of watching this has produced.</p>
<h2 id="the-state-file-is-the-whole-problem">the state file is the whole problem<a class="anchor" href="#the-state-file-is-the-whole-problem" aria-label="link to this section">#</a></h2>
<p>Declarative infrastructure tools work by comparing three things: your configuration, a recorded state, and reality. Every hard problem comes from those three disagreeing.</p>
<p><strong>State drift.</strong> Someone changed something by hand. Now state says one thing and reality says another, and the next apply will either revert their change or fail confusingly.</p>
<p>The fix is process, not tooling: <strong>nobody has write access to production infrastructure except the pipeline.</strong> Break-glass access exists, is audited, and triggers a drift check afterward. Without this, drift is continuous and IaC becomes theater.</p>
<p><strong>State locking.</strong> Two applies at once corrupt state. Every backend supports locking. Verify yours is actually enabled — this is a default that people leave off and discover during an incident.</p>
<p><strong>State as a blast radius.</strong> One giant state file means every apply touches everything and a corrupted state loses everything. Split by lifecycle and blast radius: networking, data stores, and applications should not share state.</p>
<p>Rule of thumb: if a plan takes more than a couple of minutes, your state is too big.</p>
<h2 id="modules-that-became-frameworks">modules that became frameworks<a class="anchor" href="#modules-that-became-frameworks" aria-label="link to this section">#</a></h2>
<p>The pattern is universal. A team writes a module to wrap a resource with sensible defaults. It grows parameters. It grows conditionals. Three years later it has forty inputs, nested conditional logic, and nobody can tell what it creates without running a plan.</p>
<p>At that point the abstraction is costing more than the duplication it prevented.</p>
<p><strong>The guidance that holds:</strong></p>
<ul><li><strong>Wrap only when there is real shared policy</strong> — tagging, naming, security baselines. Not to save typing.</li><li><strong>Modules should be shallow.</strong> A module that takes one resource and adds organizational defaults is good. A module that composes twelve resources with conditional branches is a framework, and frameworks need maintainers.</li><li><strong>Prefer duplication over the wrong abstraction</strong>, more so here than in application code, because infrastructure changes less and the cost of a bad abstraction is higher.</li><li><strong>No conditionals that change what resources exist.</strong> <code>count = var.enabled ? 1 : 0</code> is fine once. Nested versions of it produce configurations nobody can reason about.</li></ul>
<h2 id="the-practice-that-separates-teams">the practice that separates teams<a class="anchor" href="#the-practice-that-separates-teams" aria-label="link to this section">#</a></h2>
<p><strong>Plan output is reviewed as part of the pull request.</strong></p>
<p>Not "run apply and see." The plan — the exact set of creates, changes, and destroys — posted to the pull request, read by a human, before merge.</p>
<p>Teams that do this catch the accidental destroy, the security group opened to the world, the database replacement that would have caused an outage. Teams that do not find out during apply.</p>
<p>This is a one-day CI setup and it is the single highest-value practice in the whole category.</p>
<p>The corollary: <strong>pay attention to every destroy in a plan.</strong> Most infrastructure incidents caused by IaC are a resource being replaced when the author expected it to be updated in place. Some attribute changes force replacement. The plan tells you. Nobody reads it.</p>
<h2 id="the-things-to-keep-out">the things to keep out<a class="anchor" href="#the-things-to-keep-out" aria-label="link to this section">#</a></h2>
<p><strong>Secrets.</strong> Never in the configuration, never in state. State files contain resource attributes in plaintext, including some you would not expect. Use a secret manager and reference it, and encrypt the state backend regardless.</p>
<p><strong>Application deployment.</strong> Infrastructure tools are bad at deploying applications — they are declarative and convergent, and deployment is imperative and sequenced. Use them to provision the platform; use something else to deploy onto it.</p>
<p><strong>Anything with a fast change cadence.</strong> If a value changes weekly, it should be configuration read at runtime, not infrastructure applied through a pipeline.</p>
<h2 id="the-fork-question">the fork question<a class="anchor" href="#the-fork-question" aria-label="link to this section">#</a></h2>
<p>The OpenTofu fork means there is now a genuinely open-source implementation with a foundation behind it, alongside the commercial one. Both work. The configuration language is largely compatible.</p>
<p>The practical guidance: this is a lower-stakes decision than it feels. The <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a> is in your modules and your provider usage, not in the binary. Pick based on your licensing requirements and your appetite for the respective governance models, and know that migrating between them is a smaller project than it sounds.</p>
<h2 id="the-thing-i-would-tell-a-team-starting-today">the thing I would tell a team starting today<a class="anchor" href="#the-thing-i-would-tell-a-team-starting-today" aria-label="link to this section">#</a></h2>
<p>Start smaller than you think. One state file per environment per major system. Plain resources before modules. Plan review in CI from day one. No manual changes, ever, enforced by permissions.</p>
<p>The teams that hate their infrastructure code all made the same mistake: they built an abstraction layer before they understood the domain, and then they were maintaining two things.</p>]]></content:encoded></item><item><title>The state of self-hosting</title><link>https://readme.news/the-state-of-self-hosting/</link><guid isPermaLink="true">https://readme.news/the-state-of-self-hosting/</guid><pubDate>Fri, 20 Mar 2026 09:00:00 +0000</pubDate><description>Running your own infrastructure got dramatically easier while the industry was arguing about the cloud. A practical assessment.</description><content:encoded><![CDATA[<p>The default answer to "where should this run" has been "the cloud" for fifteen years, and for most of that time it was correct.</p>
<p>Several things changed and the answer is now more nuanced than the reflex suggests.</p>
<h2 id="what-changed-in-favor-of-self-hosting">what changed in favor of self-hosting<a class="anchor" href="#what-changed-in-favor-of-self-hosting" aria-label="link to this section">#</a></h2>
<p><strong>Machines got enormous.</strong> A single server you can rent for a few hundred dollars a month has more cores, more memory, and dramatically more I/O than a rack of hardware from 2012. A very large number of applications fit on one machine with room to spare.</p>
<p><strong>The tooling got good.</strong> Configuration management, container runtimes, reverse proxies with automatic certificates, backup tooling. What required a team a decade ago requires a competent person and a weekend.</p>
<p><strong>Cloud egress pricing did not fall.</strong> Compute prices came down. Bandwidth pricing at the major clouds is still a large multiple of what it costs, and for bandwidth-heavy applications it dominates the bill.</p>
<p><strong>Managed service prices are high relative to the alternative.</strong> A managed database costs several times what the equivalent instance costs, for operational convenience that is real and is not always worth the multiple.</p>
<h2 id="what-changed-against-it">what changed against it<a class="anchor" href="#what-changed-against-it" aria-label="link to this section">#</a></h2>
<p><strong>Security expectations rose.</strong> Patching, hardening, monitoring, incident response. Self-hosting means you own all of it, and the threat environment is worse than it was.</p>
<p><strong>Compliance frameworks assume cloud controls.</strong> SOC 2, ISO 27001, and their relatives are achievable self-hosted and the evidence collection is more work.</p>
<p><strong>The talent assumption inverted.</strong> A decade ago every team had someone who knew Linux systems administration. Now a lot of teams do not, and hiring for it is harder than hiring for cloud skills.</p>
<h2 id="the-honest-decision-framework">the honest decision framework<a class="anchor" href="#the-honest-decision-framework" aria-label="link to this section">#</a></h2>
<p><strong>Self-host when:</strong></p>
<ul><li>Your workload is steady rather than spiky. Cloud's core value proposition is elasticity, and you are paying for elasticity you do not use.</li><li>Bandwidth is a large share of your bill.</li><li>You have or can hire operational competence.</li><li>Data locality or sovereignty is a requirement.</li><li>You are at a scale where the cloud premium is a meaningful number — which starts lower than most people assume.</li></ul>
<p><strong>Use the cloud when:</strong></p>
<ul><li>Traffic is spiky or unpredictable.</li><li>You are early and optimizing for speed of iteration over unit economics.</li><li>You need global presence and do not want to operate it.</li><li>Your team's time is better spent on the product, which for most early-stage companies it is.</li><li>Compliance requirements are easier to satisfy with a provider's attestations.</li></ul>
<p><strong>The hybrid that most people should consider:</strong> run the steady baseline on owned or rented hardware, burst to cloud for peaks, keep object storage and CDN with a provider. This captures most of the cost advantage without giving up elasticity where it matters.</p>
<h2 id="the-middle-option-nobody-talks-about">the middle option nobody talks about<a class="anchor" href="#the-middle-option-nobody-talks-about" aria-label="link to this section">#</a></h2>
<p>Between "hyperscaler" and "rack in a colo" there is a large market of dedicated server providers and mid-size clouds: a real machine, in a real datacenter, with network and power handled, for a monthly fee.</p>
<p>You get root, predictable performance without noisy neighbors, and bandwidth allowances that are not priced as a profit center. You do not get managed databases, autoscaling, or a hundred adjacent services.</p>
<p>For a very large number of applications this is the correct answer and it is under-considered because the discourse is binary.</p>
<h2 id="the-operational-minimum">the operational minimum<a class="anchor" href="#the-operational-minimum" aria-label="link to this section">#</a></h2>
<p>If you self-host, these are non-negotiable:</p>
<ul><li><strong>Automated, tested restores.</strong> Not backups — <em>restores</em>. A backup you have never restored is a hypothesis. Test it quarterly, on a schedule, with a timer running.</li><li><strong>Unattended security updates</strong>, at least for the OS.</li><li><strong>Monitoring with alerting that reaches a human.</strong> Disk full is the most common self-hosted outage and it is entirely preventable.</li><li><strong><a class="xref" href="/infrastructure-as-code-ten-years-of-lessons/" title="Infrastructure as code, ten years of lessons">Infrastructure as code</a>.</strong> The machine must be reproducible. If rebuilding it requires someone's memory, you have a single point of failure that is a person.</li><li><strong>A documented runbook</strong> for the failures you expect: disk, certificate expiration, service crash, host failure.</li></ul>
<p>That is a weekend of setup and a few hours a month. If nobody on the team will own those hours, use the cloud — that is a legitimate reason and it is the actual deciding factor more often than cost is.</p>
<h2 id="the-thing-that-changed-my-mind">the thing that changed my mind<a class="anchor" href="#the-thing-that-changed-my-mind" aria-label="link to this section">#</a></h2>
<p>I used to treat "we run our own servers" as a red flag. I now treat "we are on the cloud and have never modeled the alternative" as an equal one.</p>
<p>Both are defaults applied without analysis. The analysis takes an afternoon and the answer is frequently not what the reflex says.</p>]]></content:encoded></item><item><title>Feature flags and the state space nobody tests</title><link>https://readme.news/feature-flags-and-the-state-space-nobody-tests/</link><guid isPermaLink="true">https://readme.news/feature-flags-and-the-state-space-nobody-tests/</guid><pubDate>Mon, 02 Mar 2026 09:00:00 +0000</pubDate><description>Twenty flags is a million configurations. You are testing one of them. Here&#x27;s how to keep that from being a problem.</description><content:encoded><![CDATA[<p>Feature flags decouple deploy from release, enable gradual rollout, and give you a kill switch. All of that is genuinely valuable.</p>
<p>They also multiply your state space by two for every flag you add, and almost nobody accounts for that.</p>
<h2 id="the-arithmetic">the arithmetic<a class="anchor" href="#the-arithmetic" aria-label="link to this section">#</a></h2>
<p>Twenty boolean flags is 2^20 possible configurations — over a million. Your test suite exercises the default configuration. Your staging environment exercises one other. Production is running dozens simultaneously, because different flags are on for different user segments.</p>
<p>Which means: <strong>most of the configurations your users are running have never been executed anywhere before.</strong></p>
<p>Most of the time this is fine, because most flags are independent. The failures come from the ones that are not, and you find out about those from a support ticket.</p>
<h2 id="the-failure-modes">the failure modes<a class="anchor" href="#the-failure-modes" aria-label="link to this section">#</a></h2>
<p><strong>Interaction bugs.</strong> Flag A changes the data format. Flag B reads that data. Both work alone. Together, one of them is reading a shape it does not expect. This is the classic and it is very hard to catch because neither change is wrong.</p>
<p><strong>Flag debt.</strong> A flag that has been at 100% for a year is still in the code, with its dead branch, and nobody remembers whether it can be removed. The dead branch does not compile-error and does not test-fail; it just sits there being wrong.</p>
<p><strong>The stale branch.</strong> Once a flag is at 100%, the disabled path stops being exercised. Six months later someone flips it as a rollback and discovers the old path no longer works with the current schema. Your kill switch is broken and you find out during an incident.</p>
<p><strong>Configuration drift.</strong> Flags set differently in staging and production means staging tests a configuration nobody runs.</p>
<h2 id="the-rules-that-keep-this-manageable">the rules that keep this manageable<a class="anchor" href="#the-rules-that-keep-this-manageable" aria-label="link to this section">#</a></h2>
<p><strong>Every flag has an owner and an expiry date at creation.</strong> Not a suggestion — a required field. A flag that has passed its expiry shows up in a report and somebody has to either extend it with a reason or delete it.</p>
<p><strong>Flags are deleted, not left at 100%.</strong> The cleanup is part of the work, not a follow-up ticket. A flag rollout is not done when it reaches 100%; it is done when the flag and the dead branch are gone.</p>
<p><strong>Flags never nest.</strong> If flag A only means something when flag B is on, you have created a configuration nobody can reason about. Combine them into one flag with three states, or restructure.</p>
<p><strong>Kill switches are a separate category</strong> with separate rules. They live forever, they are documented, and they are <em>tested on a schedule</em> — a quarterly exercise where you flip each one in staging and verify it works. An untested kill switch is not a kill switch.</p>
<p><strong>Test both branches.</strong> Your test suite should run the critical path with each significant flag both on and off. Not the full combinatorial space — that is impossible — but each flag independently against the default configuration. That catches the "old path rotted" failure, which is the expensive one.</p>
<h2 id="the-taxonomy-that-helps">the taxonomy that helps<a class="anchor" href="#the-taxonomy-that-helps" aria-label="link to this section">#</a></h2>
<p>Four kinds of flag, with different lifecycles:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">kind</th><th style="text-align:left">lifetime</th><th style="text-align:left">example</th></tr></thead><tbody><tr><td style="text-align:left">release</td><td style="text-align:left">days to weeks</td><td style="text-align:left">shipping a new checkout</td></tr><tr><td style="text-align:left">experiment</td><td style="text-align:left">weeks</td><td style="text-align:left">A/B test</td></tr><tr><td style="text-align:left">ops / kill switch</td><td style="text-align:left">permanent</td><td style="text-align:left">disable recommendations</td></tr><tr><td style="text-align:left">permission</td><td style="text-align:left">permanent</td><td style="text-align:left">enterprise-tier feature</td></tr></tbody></table></div>
<p>The first two must expire. The last two must be documented and tested. Conflating them is how you get a thousand flags and no idea which matter.</p>
<h2 id="the-observability-part">the observability part<a class="anchor" href="#the-observability-part" aria-label="link to this section">#</a></h2>
<p>Every event you log should carry the flag configuration that produced it.</p>
<div class="code"><span class="code-lang">json</span><pre><code class="lang-json">{ "event": "checkout", "status": 500, "flags": ["new_pricing", "fast_path"] }</code></pre></div>
<p>Without this, a bug that only affects one flag combination is undiagnosable — you see errors, you cannot correlate them to anything, and you spend a day guessing. With it, the correlation is a single query.</p>
<p>This is a fifteen-minute change and it is the single highest-value thing you can do if you use flags at all.</p>
<h2 id="the-counterargument">the counterargument<a class="anchor" href="#the-counterargument" aria-label="link to this section">#</a></h2>
<p>Some teams respond to all of this by using fewer flags and doing more trunk-based deployment with fast rollback.</p>
<p>That is a legitimate position and it works when your rollback is genuinely fast and your changes are genuinely reversible. Flags are a tool for when they are not — when a change is expensive to revert, when you need per-segment control, or when release timing must be decoupled from deployment for business reasons.</p>
<p>Use them for those. Do not use them because they feel safer, because a flag you do not test is not safety, it is the appearance of it.</p>]]></content:encoded></item><item><title>re:Invent and the year of the boring cloud announcement</title><link>https://readme.news/reinvent-and-the-year-of-the-boring-cloud-announcement/</link><guid isPermaLink="true">https://readme.news/reinvent-and-the-year-of-the-boring-cloud-announcement/</guid><pubDate>Tue, 02 Dec 2025 09:00:00 +0000</pubDate><description>Silicon, agents, and a keynote that mostly described things that already existed. That&#x27;s a healthy sign.</description><content:encoded><![CDATA[<p>AWS re:Invent is underway and the announcement pace is, as always, absurd. Most of it will not matter to you. Here is the part that will.</p>
<h2 id="silicon">silicon<a class="anchor" href="#silicon" aria-label="link to this section">#</a></h2>
<p>Amazon continues pushing Trainium and Inferentia as the alternative to buying Nvidia. The pitch is price-performance for customers who can tolerate a different software stack.</p>
<p>The honest state of it: the hardware is competitive on paper, the software ecosystem is meaningfully behind CUDA, and the gap is closing slowly. If your workload runs through PyTorch with standard operations, the port is manageable. If you have custom kernels, it is a project.</p>
<p>The strategic point is the same for every hyperscaler: reducing dependency on a single supplier with enormous pricing power. Whether the customer benefit materializes depends entirely on whether the savings get passed through.</p>
<h2 id="the-agent-announcements">the agent announcements<a class="anchor" href="#the-agent-announcements" aria-label="link to this section">#</a></h2>
<p>Every cloud vendor is now shipping agent infrastructure: runtimes, memory services, gateways for tool access, identity for agents, observability for agent traces.</p>
<p>This category is real. Running agents in production has genuine infrastructure requirements that are different from running services:</p>
<ul><li><strong>Long-lived sessions</strong> with state that outlives a request.</li><li><strong>Non-deterministic execution paths</strong> that make traditional tracing awkward.</li><li><strong>Cost per invocation that varies by orders of magnitude.</strong></li><li><strong>Identity and permission scoping</strong> for a thing acting on a user's behalf.</li><li><strong>Human approval gates</strong> in the middle of automated flows.</li></ul>
<p>Those are real problems and the tooling is early everywhere. Evaluate on whether it solves a problem you actually have, not on whether the demo was good.</p>
<h2 id="the-pattern-i-would-push-back-on">the pattern I would push back on<a class="anchor" href="#the-pattern-i-would-push-back-on" aria-label="link to this section">#</a></h2>
<p>Every vendor's agent framework wants to be the place your orchestration lives. That is a <a class="xref" href="/vendor-lock-in-an-honest-cost-model/" title="Vendor lock-in: an honest cost model">lock-in</a> position, and orchestration is the layer most likely to be absorbed by the models themselves — it has been happening steadily for two years.</p>
<p>Keep your orchestration portable. Use the managed pieces for the genuinely hard infrastructure — identity, secure tool access, session storage — and keep the logic in your own code.</p>
<h2 id="the-underrated-announcements">the underrated announcements<a class="anchor" href="#the-underrated-announcements" aria-label="link to this section">#</a></h2>
<p>The ones nobody writes about and everybody uses:</p>
<ul><li>Incremental improvements to S3 consistency and performance.</li><li>Networking latency reductions.</li><li>Cost management tooling that is slightly less bad.</li><li>Database engine version updates.</li></ul>
<p>These are worth more to most organizations than any AI announcement, and they get four slides between two hours of agent demos.</p>
<h2 id="the-meta-observation">the meta-observation<a class="anchor" href="#the-meta-observation" aria-label="link to this section">#</a></h2>
<p>Cloud conferences have gotten less interesting, and that is good. It means the platform is mature. The exciting years of a platform are the years when fundamental things are missing.</p>
<p>The interesting question for AWS is not what they announced. It is whether the operational excellence that justified the premium is still there after a year that included a major regional outage. That is an execution question and it does not get answered at a conference.</p>
<h2 id="what-to-actually-do-with-this">what to actually do with this<a class="anchor" href="#what-to-actually-do-with-this" aria-label="link to this section">#</a></h2>
<p>Skip the keynote. Read the "what's new" feed filtered to the services you actually use. Look for the deprecations, which are the announcements that will cost you time and which are never on stage.</p>
<p>And check your bill. The single highest-value hour available to most engineering organizations is someone competent looking at the AWS bill line by line, and almost nobody does it.</p>]]></content:encoded></item><item><title>us-east-1 goes down and takes a large chunk of the internet with it</title><link>https://readme.news/us-east-1-goes-down-and-takes-a-large-chunk-of-the-internet-with-it/</link><guid isPermaLink="true">https://readme.news/us-east-1-goes-down-and-takes-a-large-chunk-of-the-internet-with-it/</guid><pubDate>Tue, 21 Oct 2025 09:00:00 +0000</pubDate><description>A DNS race condition in DynamoDB&#x27;s automation cascades across dozens of AWS services. The lesson is about coupling, not DNS.</description><content:encoded><![CDATA[<p>AWS's us-east-1 region suffered a multi-hour outage yesterday that affected a very large number of services and, through them, a very large fraction of consumer internet applications.</p>
<p>AWS's public summary attributes the trigger to a latent race condition in the automation that manages DynamoDB's DNS records, which resulted in an empty record set for a regional endpoint and no automatic recovery path.</p>
<h2 id="the-cascade">the cascade<a class="anchor" href="#the-cascade" aria-label="link to this section">#</a></h2>
<p>The failure sequence is the interesting part.</p>
<p>DynamoDB's endpoint became unresolvable. That alone would be bad. What made it a regional event is that <strong>an enormous number of AWS's own services use DynamoDB internally.</strong> The EC2 instance launch path, IAM's control plane, Lambda's invocation machinery, and dozens of others depend on it.</p>
<p>So the failure propagated: DynamoDB down means new EC2 instances cannot launch, which means autoscaling cannot replace failing capacity, which means load shedding, which means more failures. Network Load Balancer health checks destabilized. The recovery itself was slowed by the backlog of queued work that had accumulated.</p>
<p>This is textbook <strong>metastable failure</strong>: a system that is stable under normal load and stable under no load, but which, once pushed past a threshold, sustains its own failure through retry amplification and queue buildup even after the original trigger is fixed.</p>
<h2 id="why-us-east-1">why us-east-1<a class="anchor" href="#why-us-east-1" aria-label="link to this section">#</a></h2>
<p>It is the oldest region, the largest, and the default in approximately every tutorial ever written. Several global AWS control planes are homed there — IAM, CloudFront configuration, Route 53's control plane, and others. That means a us-east-1 event has global blast radius even for customers with no resources in the region.</p>
<p>That architecture is a historical artifact. It is also extremely hard to change now, which is a lesson about early decisions in systems that grow.</p>
<h2 id="the-honest-customer-takeaway">the honest customer takeaway<a class="anchor" href="#the-honest-customer-takeaway" aria-label="link to this section">#</a></h2>
<p>The reflexive response is "multi-region." Before you spend a year on that, do the arithmetic.</p>
<p><strong>Multi-region active-active is genuinely hard.</strong> Data consistency across regions, failover testing that actually works, doubled infrastructure cost, and a substantially more complex system that fails in new ways. Many organizations that attempt it end up with a system that is <em>less</em> reliable overall because the complexity introduces more <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> than the regional risk it removes.</p>
<p><strong>The dependency you cannot escape.</strong> If your multi-region architecture depends on a global control plane that lives in us-east-1, you did not achieve independence. Check this specifically. A lot of people discovered it yesterday.</p>
<p><strong>What is actually worth doing, in order:</strong></p>
<ol><li><strong>Know your dependencies.</strong> Most teams cannot enumerate what their service requires to start. Write it down. The exercise is revealing.</li><li><strong>Static stability.</strong> Design so existing capacity keeps serving when the control plane is unavailable. If your service needs to call an API to keep running, it will stop when that API stops. Cache aggressively, fail open where safe, and do not require a control plane call on the request path.</li><li><strong>Graceful degradation.</strong> Decide in advance which features can be turned off. A checkout that works without recommendations is much better than a site that is down.</li><li><strong>Exponential backoff with jitter, and circuit breakers.</strong> Retry storms are what turns an incident into an outage. This is the single highest-leverage code change available.</li><li><strong>Multi-region for the tier that genuinely warrants it.</strong> Which is usually not everything.</li></ol>
<h2 id="the-industry-level-observation">the industry-level observation<a class="anchor" href="#the-industry-level-observation" aria-label="link to this section">#</a></h2>
<p>A meaningful fraction of the world's software depends on a small number of regions operated by a small number of companies. That concentration produces excellent reliability most of the time and correlated failure occasionally.</p>
<p>There is no individual fix. Every company independently choosing the most reliable provider produces exactly this concentration. It is a collective action problem, and the only actors who can address it are regulators thinking about systemic risk, who are — belatedly — starting to.</p>]]></content:encoded></item><item><title>In defense of the boring deploy</title><link>https://readme.news/in-defense-of-the-boring-deploy/</link><guid isPermaLink="true">https://readme.news/in-defense-of-the-boring-deploy/</guid><pubDate>Mon, 10 Mar 2025 09:00:00 +0000</pubDate><description>Blue-green, canaries, feature flags, and the deeply unfashionable practice of shipping the same way every time.</description><content:encoded><![CDATA[<p>The best deployment I ever worked on took eleven minutes, ran the same way every time, and had a rollback that was a single command anyone on the team could run from their phone. Nobody wrote a conference talk about it.</p>
<p>Deployment is a solved problem that organizations keep re-opening because the solution is boring and boring does not get anyone promoted.</p>
<h2 id="the-properties-that-matter">the properties that matter<a class="anchor" href="#the-properties-that-matter" aria-label="link to this section">#</a></h2>
<p>Strip away the tooling debates and a good deploy has four properties.</p>
<p><strong>It is the same every time.</strong> Not "mostly the same, except for the database migration, and except on Fridays, and except when Sarah does it." The same. If there is a manual step, it is in the pipeline as an approval gate, not in someone's head.</p>
<p><strong>It is reversible in under a minute.</strong> This is the one people get wrong. They build elaborate progressive rollouts and then discover that rolling back requires reverting a migration that dropped a column. Reversibility is a design constraint on your <em>changes</em>, not a feature of your deploy tool.</p>
<p><strong>Failure is detected without a human looking.</strong> If the way you find out about a bad deploy is a customer email, you do not have a deployment process, you have a deployment ritual.</p>
<p><strong>It happens often enough to be unremarkable.</strong> The failure rate of a deploy scales superlinearly with the amount of change in it. Ten small deploys are dramatically safer than one deploy containing ten changes, and this is the single most robust finding in the entire delivery-metrics literature.</p>
<h2 id="the-expand-contract-discipline">the expand-contract discipline<a class="anchor" href="#the-expand-contract-discipline" aria-label="link to this section">#</a></h2>
<p>Most rollback pain is schema pain. The fix is a discipline, not a tool:</p>
<ol><li><strong>Expand.</strong> Add the new column, nullable. Deploy. The old code ignores it.</li><li><strong>Migrate.</strong> Backfill. Dual-write from application code. Deploy. Both shapes work.</li><li><strong>Switch.</strong> Read from the new column. Deploy. Still reversible — the old column is intact.</li><li><strong>Contract.</strong> Drop the old column, days or weeks later, once you are certain.</li></ol>
<p>Four deploys instead of one. Every intermediate state is reversible. It feels slow and it is the reason you do not have a 3 a.m. incident where the rollback made things worse.</p>
<h2 id="flags-are-not-free">flags are not free<a class="anchor" href="#flags-are-not-free" aria-label="link to this section">#</a></h2>
<p><a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">Feature flags</a> decouple deploy from release and that is genuinely valuable. They also multiply your state space: <code>n</code> flags means <code>2^n</code> possible configurations, and you are testing approximately one of them.</p>
<p>Rules that keep this manageable:</p>
<ul><li>Every flag gets an owner and an expiry date at creation.</li><li>A flag that has been at 100% for a month is deleted, not left "just in case."</li><li>Flags never nest. If flag A only matters when flag B is on, you have made something unreasonable.</li><li>Kill-switch flags are a separate category with separate rules and they can live forever.</li></ul>
<p>The failure mode is flag debt: a codebase with two hundred flags where nobody knows which combinations are actually exercised in production. That is worse than no flags, because it looks like safety.</p>
<h2 id="the-actually-unfashionable-opinion">the actually unfashionable opinion<a class="anchor" href="#the-actually-unfashionable-opinion" aria-label="link to this section">#</a></h2>
<p>Most teams do not need Kubernetes to deploy well, and many teams running Kubernetes deploy worse than they did before, because the platform's complexity consumed the attention that used to go into the process.</p>
<p>If your deploys are fast, boring, and reversible, your infrastructure is correct regardless of what it is. If they are not, no amount of platform will fix it, because the problem is that nobody has treated the deploy as a product with a user.</p>
<p>The user is your team at 2 a.m. Design for them.</p>]]></content:encoded></item><item><title>Your CI is the slowest developer on the team</title><link>https://readme.news/your-ci-is-the-slowest-developer-on-the-team/</link><guid isPermaLink="true">https://readme.news/your-ci-is-the-slowest-developer-on-the-team/</guid><pubDate>Tue, 04 Feb 2025 09:00:00 +0000</pubDate><description>A twenty-minute pipeline doesn&#x27;t cost twenty minutes. It costs the context switch, the batching, and the review you skipped.</description><content:encoded><![CDATA[<p>There is a number in your organization that nobody owns and everybody pays: the wall-clock time between pushing a commit and knowing whether it is good.</p>
<p>Call it T. If T is two minutes, your team behaves one way. If T is thirty minutes, it behaves completely differently, and the difference is not "things take twenty-eight minutes longer."</p>
<h2 id="what-actually-happens-as-t-grows">what actually happens as T grows<a class="anchor" href="#what-actually-happens-as-t-grows" aria-label="link to this section">#</a></h2>
<p><strong>Below ~2 minutes</strong>, you wait. You keep the change in your head, you see the result, you fix it. The loop is tight enough that debugging is interactive.</p>
<p><strong>Between 2 and 10 minutes</strong>, you context switch. You go read something else. When the result comes back you have to page the change back in. The reload cost is real and it is roughly proportional to how complex the change was — which means it is worst exactly when you can least afford it.</p>
<p><strong>Past 10 minutes</strong>, behavior changes qualitatively. People start batching. They stop pushing small commits because the feedback is too expensive per commit, so they push bigger ones, which are harder to review and more likely to fail, which makes each failure more expensive to diagnose. It is a doom loop and it is entirely emergent from one number.</p>
<p><strong>Past 30 minutes</strong>, people route around CI. They test locally in ways that diverge from CI, they merge on green-ish, they add <code>[skip ci]</code>, and eventually somebody proposes a nightly build as the "real" signal, at which point you have given up.</p>
<h2 id="the-diagnostic">the diagnostic<a class="anchor" href="#the-diagnostic" aria-label="link to this section">#</a></h2>
<p>Before optimizing, measure the right thing. Not average pipeline duration — <strong>p90 time-to-signal for the check that actually blocks merge</strong>. Most teams measure the wrong thing here and optimize a stage nobody was waiting on.</p>
<p>Then break it down:</p>
<div class="code"><pre><code>queue wait      ← runners are saturated or your concurrency limit is wrong
checkout        ← you're doing a full clone of a repo with 200k commits
dependency install ← almost always the biggest fixable chunk
build           ← is anything cached? really?
test            ← is it parallel? is it parallel *well*?</code></pre></div>
<h2 id="the-fixes-roughly-in-order-of-return">the fixes, roughly in order of return<a class="anchor" href="#the-fixes-roughly-in-order-of-return" aria-label="link to this section">#</a></h2>
<ol><li><strong>Cache dependencies properly.</strong> Not "we have a cache step." Verify the hit rate. A cache that misses on every branch because the key includes the branch name is worse than no cache, because it also uploads.</li><li><strong>Shallow clone.</strong> <code>fetch-depth: 1</code> unless you genuinely need history. If you need history for one job, do it in one job.</li><li><strong>Split the blocking check from the exhaustive check.</strong> Lint, typecheck, and unit tests block merge. Integration, e2e, and the full matrix run after, and page someone if they fail. Not everything needs to gate.</li><li><strong>Parallelize by timing, not by file count.</strong> Splitting tests evenly by count gives you one shard that takes four times as long as the others. Split by historical duration.</li><li><strong>Kill the flaky tests.</strong> A test that fails 2% of the time in a suite of 50 parallel jobs means your pipeline fails constantly. Quarantine them the day they are identified. A quarantined test is a bug ticket; a flaky test in the blocking path is a tax on everyone forever.</li><li><strong>Bigger runners.</strong> This is the least intellectually satisfying fix and often the best one. Engineer time costs more than compute. Do the arithmetic before you spend a sprint on a clever solution.</li></ol>
<h2 id="the-cultural-part">the cultural part<a class="anchor" href="#the-cultural-part" aria-label="link to this section">#</a></h2>
<p>Someone has to own T. Not "the <a class="xref" href="/platform-teams-that-dont-get-resented/" title="Platform teams that don&#x27;t get resented">platform team</a> should look at CI sometime." A named person, a <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>, and a number that goes in the same review as uptime.</p>
<p>Because the alternative is that it degrades a little every sprint — one more test, one more dependency, one more required check — and nobody notices until it is thirty minutes and everybody has quietly stopped trusting it.</p>]]></content:encoded></item>
</channel>
</rss>
