<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — incidents</title>
<link>https://readme.news/tags/incidents/</link>
<atom:link href="https://readme.news/tags/incidents/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged incidents.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>What a good incident channel looks like</title><link>https://readme.news/what-a-good-incident-channel-looks-like/</link><guid isPermaLink="true">https://readme.news/what-a-good-incident-channel-looks-like/</guid><pubDate>Sat, 29 Aug 2026 09:00:00 +0000</pubDate><description>The chat log is where the incident is actually run. Four conventions make it useful rather than a wall of noise.</description><content:encoded><![CDATA[<p>Every incident runs in a chat channel. Almost nobody has thought about how that channel should work, so the default is thirty people speculating in parallel while two people try to fix something.</p>
<p>Four conventions fix most of it, and none of them need a tool.</p>
<h2 id="1-one-channel-created-at-declaration">1. one channel, created at declaration<a class="anchor" href="#1-one-channel-created-at-declaration" aria-label="link to this section">#</a></h2>
<p>Not the team channel. Not a thread in the team channel. A dedicated channel per incident, named predictably: <code>inc-2026-08-29-checkout-errors</code>.</p>
<p>Why it matters: the incident becomes searchable as a unit, the timeline is the channel, and people who join late can read from the top instead of asking what happened. It also means the incident does not drown the normal channel, and the normal channel does not drown the incident.</p>
<p>Archive it afterwards rather than deleting it. It is the primary source for the postmortem.</p>
<h2 id="2-named-roles-stated-out-loud">2. named roles, stated out loud<a class="anchor" href="#2-named-roles-stated-out-loud" aria-label="link to this section">#</a></h2>
<p>Three, at minimum, posted in the channel as the first message:</p>
<div class="code"><pre><code>IC: @priya       — decides, delegates, does not debug
Comms: @marcus   — status page, stakeholders, customer support
Ops: @sam        — hands on keyboard</code></pre></div>
<p>The incident commander not debugging is the rule people resist and the one that matters most. The moment the IC opens a terminal, nobody is tracking the whole picture, and the incident gets longer.</p>
<p>For a small incident one person can hold two roles. They should still say which ones, because the alternative is everyone assuming someone else is doing comms.</p>
<h2 id="3-a-pinned-status-message-edited-in-place">3. a pinned status message, edited in place<a class="anchor" href="#3-a-pinned-status-message-edited-in-place" aria-label="link to this section">#</a></h2>
<p>One message, pinned, rewritten as things change:</p>
<div class="code"><pre><code>STATUS 14:22 — Checkout failing for ~12% of users since 13:58.
Cause: unknown. Suspect the payments deploy at 13:55.
Now: rolling back that deploy (@sam), ETA 5 min.
Impact: card payments only, wallet payments unaffected.
Next update: 14:35</code></pre></div>
<p>Five lines: what is broken, since when, what we think, what we are doing, when the next update comes.</p>
<p>This single convention removes most of the noise, because it answers the question that generates the noise — "what's the current state?" — without anyone having to ask. Everyone joining reads the pin instead of scrolling.</p>
<p><strong>Always include the next-update time.</strong> It is what stops people asking for updates.</p>
<h2 id="4-mark-the-speculation">4. mark the speculation<a class="anchor" href="#4-mark-the-speculation" aria-label="link to this section">#</a></h2>
<p>The most common way incidents go wrong is that a guess gets repeated until it becomes the working theory, and then twenty minutes go into the wrong system.</p>
<p>A one-word convention fixes it:</p>
<div class="code"><pre><code>FACT: error rate went from 0.1% to 12% at 13:58:20
FACT: the payments deploy completed at 13:55:41
GUESS: the deploy caused it
ACTION: rolling back to confirm</code></pre></div>
<p>Facts have evidence attached. Guesses are labelled as guesses. Actions say who is doing them.</p>
<p>It looks pedantic for about four minutes and then it saves the incident, because somebody reading the channel can tell the difference between what is known and what somebody said.</p>
<h2 id="what-to-keep-out">what to keep out<a class="anchor" href="#what-to-keep-out" aria-label="link to this section">#</a></h2>
<p><strong>Speculation from people who are not investigating.</strong> Well-meant and it fills the channel that responders are trying to read. If you are not on a role, watch.</p>
<p><strong>"Is it fixed yet?"</strong> The pinned status has the next update time.</p>
<p><strong>Root-cause analysis during the incident.</strong> Restore first. The good question at 14:22 is "what makes this stop", not "why did this happen". Why is a question for Thursday.</p>
<p><strong>Blame, in any form, including jokes.</strong> It is in a permanent record that a person will read afterwards.</p>
<h2 id="the-handoff">the handoff<a class="anchor" href="#the-handoff" aria-label="link to this section">#</a></h2>
<p>Long incidents cross shift boundaries and the handoff is where context dies. It needs to be explicit and in the channel:</p>
<div class="code"><pre><code>HANDOFF 22:00 — @priya → @dan (IC)
Ruled out: deploy (rolled back, no change), CDN (unaffected regions also failing)
Current theory: connection pool exhaustion on the orders DB
In flight: @sam is capturing pg_stat_activity every 30s → thread above
Not yet tried: failover to replica
Customers: status page updated 21:40, support has the template</code></pre></div>
<p>Ruled out, current theory, in flight, not tried, customer state. Five lines, and the incoming IC starts with the accumulated knowledge instead of rediscovering it.</p>
<h2 id="why-bother">why bother<a class="anchor" href="#why-bother" aria-label="link to this section">#</a></h2>
<p>The channel is not a side effect of the incident. During the incident it is the coordination mechanism, and afterwards it is the only complete record of what was known and when.</p>
<p>A channel that follows these four conventions produces a postmortem that half writes itself, and — more importantly — an incident that ends sooner, because the people fixing it spent their attention on the system rather than on the chat.</p>]]></content:encoded></item><item><title>The postmortem that changes something</title><link>https://readme.news/the-postmortem-that-changes-something/</link><guid isPermaLink="true">https://readme.news/the-postmortem-that-changes-something/</guid><pubDate>Tue, 16 Dec 2025 09:00:00 +0000</pubDate><description>Most incident reviews produce a document and a ticket that never gets done. Here&#x27;s the difference.</description><content:encoded><![CDATA[<p>Most organizations run blameless postmortems. Most organizations also have the same class of incident repeatedly. Both of those things are true simultaneously and nobody finds it strange.</p>
<h2 id="the-failure-pattern">the failure pattern<a class="anchor" href="#the-failure-pattern" aria-label="link to this section">#</a></h2>
<p>The typical postmortem:</p>
<ol><li>Timeline of what happened. Accurate, detailed, useful.</li><li>Root cause. Usually a single technical fact.</li><li>Action items. Five to twelve of them.</li><li>Filed. Two action items get done. The rest age out.</li></ol>
<p>Six months later, a similar incident, with a similar document.</p>
<p>Three things are wrong here.</p>
<h2 id="problem-one-root-cause-is-singular">problem one: "root cause" is singular<a class="anchor" href="#problem-one-root-cause-is-singular" aria-label="link to this section">#</a></h2>
<p>Complex systems do not fail because of one thing. They fail because several conditions aligned, each of which was individually survivable.</p>
<p>The Cloudflare outage in November: a permissions change, a query that returned duplicates, a fixed buffer size, an error path that panicked instead of degrading, and a config pipeline that propagated globally in minutes. Remove any one and it does not happen, or it is much smaller.</p>
<p>Picking one and calling it "the root cause" means you fix one and leave four.</p>
<p>Better framing: <strong>contributing factors</strong>, plural, each with its own assessment of whether it is worth addressing. Some will not be — that is a legitimate decision if it is made explicitly.</p>
<h2 id="problem-two-action-items-without-owners-and-dates-are-wishes">problem two: action items without owners and dates are wishes<a class="anchor" href="#problem-two-action-items-without-owners-and-dates-are-wishes" aria-label="link to this section">#</a></h2>
<p>An action item that says "improve monitoring for the config pipeline" with no owner, no date, and no definition of done is a sentence, not a plan.</p>
<p>The fix is unglamorous:</p>
<ul><li><strong>Every action item has one named person.</strong> Not a team. A person.</li><li><strong>Every action item has a date.</strong> If nobody will commit to a date, it is not going to happen and you should delete it and say so.</li><li><strong>Every action item has a definition of done</strong> that someone else could verify.</li><li><strong>They go in the same backlog as feature work</strong>, prioritized against it. An action item in a separate "incident follow-up" list that is never sprint planned is a list of things that will not be done.</li></ul>
<p>And the one that actually forces it: <strong>review the open action items at the start of the next postmortem.</strong> Nothing motivates completion like a room full of people looking at your undone item from last quarter's incident while discussing this quarter's similar one.</p>
<h2 id="problem-three-nobody-asks-about-the-near-misses">problem three: nobody asks about the near misses<a class="anchor" href="#problem-three-nobody-asks-about-the-near-misses" aria-label="link to this section">#</a></h2>
<p>The incidents you review are the ones that broke through. For every one, there were several that did not — a bad deploy caught by a canary, a config error someone noticed in review, a query that would have taken down the database if it had run on Monday instead of Sunday.</p>
<p>Those contain the same information at a fraction of the cost, and almost nobody collects them.</p>
<p>Add a lightweight channel for it. "I nearly broke prod today, here is how." No document, no meeting, no blame. Just a note. The pattern that emerges over a quarter is more valuable than any individual postmortem.</p>
<h2 id="the-questions-that-produce-useful-findings">the questions that produce useful findings<a class="anchor" href="#the-questions-that-produce-useful-findings" aria-label="link to this section">#</a></h2>
<p>Replace "what was the root cause" with:</p>
<ul><li><strong>"What made this hard to detect?"</strong> Detection time is usually the largest component of impact and is the most improvable.</li><li><strong>"What made this hard to diagnose?"</strong> Usually missing observability. This produces the highest-value action items.</li><li><strong>"What made recovery slow?"</strong> Often a missing runbook, a missing <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">kill switch</a>, or a rollback that was not actually tested.</li><li><strong>"Who knew something that would have helped, and why did that not reach the responders?"</strong> This is an organizational question and it is frequently the real finding.</li><li><strong>"What did we do that helped?"</strong> Genuinely important. Practices that worked should be named so they get kept.</li></ul>
<h2 id="the-blameless-part-done-correctly">the blameless part, done correctly<a class="anchor" href="#the-blameless-part-done-correctly" aria-label="link to this section">#</a></h2>
<p>Blameless does not mean nobody made a mistake. It means the analysis focuses on why the mistake was possible and easy, rather than on the person.</p>
<p>"Alice deployed without running the tests" is blame and it is also useless. "The deploy path does not require tests to pass, and the shortcut that skips them is the fastest way to deploy" is the same fact stated in a way you can act on.</p>
<p>If the answer to "why did they do that" is "because the system made it easy and the correct path was hard," you have found something.</p>
<p>If the answer is genuinely "they were careless," you still fix the system, because the next person will also be careless eventually. Humans are the constant; the system is the variable.</p>]]></content:encoded></item><item><title>Cloudflare falls over because of a config file</title><link>https://readme.news/cloudflare-falls-over-because-of-a-config-file/</link><guid isPermaLink="true">https://readme.news/cloudflare-falls-over-because-of-a-config-file/</guid><pubDate>Thu, 20 Nov 2025 09:00:00 +0000</pubDate><description>A permissions change doubles the size of a generated feature file, which overflows a fixed-size buffer, which 500s a fifth of the web.</description><content:encoded><![CDATA[<p>Cloudflare had a significant outage on Tuesday, returning 5xx errors across a large portion of its network for several hours. Their published postmortem is detailed and worth reading in full.</p>
<p>The chain of events is a small masterpiece of the genre.</p>
<h2 id="what-happened">what happened<a class="anchor" href="#what-happened" aria-label="link to this section">#</a></h2>
<p>A database permissions change caused a query that generates a Bot Management feature configuration file to return <strong>duplicate rows</strong>. The query had been returning one row per feature; after the permissions change it returned rows from multiple underlying schemas.</p>
<p>The generated file therefore roughly doubled in size.</p>
<p>The proxy that consumes this file preallocates a fixed-size buffer sized against a limit of 200 features — comfortably above the ~60 actually in use. The doubled file exceeded that limit.</p>
<p>The Rust code handling this hit an unrecoverable error path and the proxy panicked rather than degrading. Because the configuration file propagates network-wide every few minutes, the failure propagated network-wide within minutes.</p>
<p>Recovery was complicated by the fact that the bad file kept regenerating and redeploying.</p>
<h2 id="the-lessons-which-are-old">the lessons, which are old<a class="anchor" href="#the-lessons-which-are-old" aria-label="link to this section">#</a></h2>
<p><strong>Configuration is code and needs the same rigor.</strong> This was not a code deploy. It was a data change that propagated to production automatically with no staging, no canary, and no validation. Config deployment pipelines are consistently held to a lower standard than code deployment pipelines, and config causes a large share of major outages.</p>
<p><strong>Fixed-size limits need to fail soft.</strong> The 200-feature limit was reasonable. Panicking when exceeded was not. The correct behavior for a proxy encountering an oversized config is to log loudly, alert, and continue with the previous known-good version.</p>
<p><strong>Validate generated artifacts before propagation.</strong> A size check, a schema check, a sanity check on row count against the previous version — any of these would have caught it. Generated files should be validated as rigorously as user input, because the generator can be wrong.</p>
<p><strong>Blast radius follows deployment speed.</strong> Config that propagates globally in minutes is a feature until it propagates a bad config globally in minutes. Staged rollout applies to configuration too.</p>
<p><strong>Have a <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">kill switch</a> for automated pipelines.</strong> Much of the recovery time went to stopping the thing that kept redeploying the bad file. Every automated deployment path needs a way to stop it that does not require fixing the underlying problem first.</p>
<h2 id="the-rust-note">the Rust note<a class="anchor" href="#the-rust-note" aria-label="link to this section">#</a></h2>
<p>This will be used as an argument about Rust, and it should not be.</p>
<p>The panic was a deliberate choice at that call site — the code used a construct that terminates on error rather than propagating it. That is a design decision about error handling, available in any language. In C the equivalent code would have written past the buffer, which is worse.</p>
<p>The actual lesson is about where you choose to make errors fatal. In a proxy handling live traffic, almost nothing should be fatal. Degrade, alert, continue. "Fail fast" is good advice for a batch job and bad advice for a load balancer.</p>
<h2 id="the-credit-due">the credit due<a class="anchor" href="#the-credit-due" aria-label="link to this section">#</a></h2>
<p>Cloudflare published a detailed technical postmortem within a day, named the specific code path, and did not hide behind "an issue with a third-party provider."</p>
<p>That is how it should be done and it is rarer than it should be. A company that publishes real postmortems earns more trust than one that never has visible incidents, because the second one is not telling you about them.</p>]]></content:encoded></item><item><title>us-east-1 goes down and takes a large chunk of the internet with it</title><link>https://readme.news/us-east-1-goes-down-and-takes-a-large-chunk-of-the-internet-with-it/</link><guid isPermaLink="true">https://readme.news/us-east-1-goes-down-and-takes-a-large-chunk-of-the-internet-with-it/</guid><pubDate>Tue, 21 Oct 2025 09:00:00 +0000</pubDate><description>A DNS race condition in DynamoDB&#x27;s automation cascades across dozens of AWS services. The lesson is about coupling, not DNS.</description><content:encoded><![CDATA[<p>AWS's us-east-1 region suffered a multi-hour outage yesterday that affected a very large number of services and, through them, a very large fraction of consumer internet applications.</p>
<p>AWS's public summary attributes the trigger to a latent race condition in the automation that manages DynamoDB's DNS records, which resulted in an empty record set for a regional endpoint and no automatic recovery path.</p>
<h2 id="the-cascade">the cascade<a class="anchor" href="#the-cascade" aria-label="link to this section">#</a></h2>
<p>The failure sequence is the interesting part.</p>
<p>DynamoDB's endpoint became unresolvable. That alone would be bad. What made it a regional event is that <strong>an enormous number of AWS's own services use DynamoDB internally.</strong> The EC2 instance launch path, IAM's control plane, Lambda's invocation machinery, and dozens of others depend on it.</p>
<p>So the failure propagated: DynamoDB down means new EC2 instances cannot launch, which means autoscaling cannot replace failing capacity, which means load shedding, which means more failures. Network Load Balancer health checks destabilized. The recovery itself was slowed by the backlog of queued work that had accumulated.</p>
<p>This is textbook <strong>metastable failure</strong>: a system that is stable under normal load and stable under no load, but which, once pushed past a threshold, sustains its own failure through retry amplification and queue buildup even after the original trigger is fixed.</p>
<h2 id="why-us-east-1">why us-east-1<a class="anchor" href="#why-us-east-1" aria-label="link to this section">#</a></h2>
<p>It is the oldest region, the largest, and the default in approximately every tutorial ever written. Several global AWS control planes are homed there — IAM, CloudFront configuration, Route 53's control plane, and others. That means a us-east-1 event has global blast radius even for customers with no resources in the region.</p>
<p>That architecture is a historical artifact. It is also extremely hard to change now, which is a lesson about early decisions in systems that grow.</p>
<h2 id="the-honest-customer-takeaway">the honest customer takeaway<a class="anchor" href="#the-honest-customer-takeaway" aria-label="link to this section">#</a></h2>
<p>The reflexive response is "multi-region." Before you spend a year on that, do the arithmetic.</p>
<p><strong>Multi-region active-active is genuinely hard.</strong> Data consistency across regions, failover testing that actually works, doubled infrastructure cost, and a substantially more complex system that fails in new ways. Many organizations that attempt it end up with a system that is <em>less</em> reliable overall because the complexity introduces more <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> than the regional risk it removes.</p>
<p><strong>The dependency you cannot escape.</strong> If your multi-region architecture depends on a global control plane that lives in us-east-1, you did not achieve independence. Check this specifically. A lot of people discovered it yesterday.</p>
<p><strong>What is actually worth doing, in order:</strong></p>
<ol><li><strong>Know your dependencies.</strong> Most teams cannot enumerate what their service requires to start. Write it down. The exercise is revealing.</li><li><strong>Static stability.</strong> Design so existing capacity keeps serving when the control plane is unavailable. If your service needs to call an API to keep running, it will stop when that API stops. Cache aggressively, fail open where safe, and do not require a control plane call on the request path.</li><li><strong>Graceful degradation.</strong> Decide in advance which features can be turned off. A checkout that works without recommendations is much better than a site that is down.</li><li><strong>Exponential backoff with jitter, and circuit breakers.</strong> Retry storms are what turns an incident into an outage. This is the single highest-leverage code change available.</li><li><strong>Multi-region for the tier that genuinely warrants it.</strong> Which is usually not everything.</li></ol>
<h2 id="the-industry-level-observation">the industry-level observation<a class="anchor" href="#the-industry-level-observation" aria-label="link to this section">#</a></h2>
<p>A meaningful fraction of the world's software depends on a small number of regions operated by a small number of companies. That concentration produces excellent reliability most of the time and correlated failure occasionally.</p>
<p>There is no individual fix. Every company independently choosing the most reliable provider produces exactly this concentration. It is a collective action problem, and the only actors who can address it are regulators thinking about systemic risk, who are — belatedly — starting to.</p>]]></content:encoded></item><item><title>Shai-Hulud: npm gets a self-replicating worm</title><link>https://readme.news/shai-hulud-npm-gets-a-self-replicating-worm/</link><guid isPermaLink="true">https://readme.news/shai-hulud-npm-gets-a-self-replicating-worm/</guid><pubDate>Tue, 16 Sep 2025 09:00:00 +0000</pubDate><description>A payload that steals credentials and then uses them to publish itself into other packages. This is the escalation everyone predicted.</description><content:encoded><![CDATA[<p>A self-propagating worm spread through npm this week, compromising hundreds of packages. The mechanism is the escalation people have been warning about for years, and it arrived exactly as described.</p>
<h2 id="how-it-works">how it works<a class="anchor" href="#how-it-works" aria-label="link to this section">#</a></h2>
<p>The loop:</p>
<ol><li>A malicious package version is installed. Its <code>postinstall</code> runs.</li><li>The payload harvests credentials from the machine — npm tokens, GitHub tokens, cloud provider credentials — using TruffleHog-style secret scanning.</li><li>It exfiltrates them to a public repository created in the victim's GitHub account.</li><li><strong>It uses the stolen npm token to publish trojanized versions of every package that token can publish to.</strong></li><li>Those packages get installed. Go to 1.</li></ol>
<p>Step four is the difference between an incident and an outbreak. Previous npm compromises required an attacker to manually obtain credentials for each package. This one propagates on its own, and its rate of spread is proportional to how many packages the compromised maintainers control.</p>
<p>Additional behavior observed: creating public forks of private repositories in victims' organizations, and installing a GitHub Actions workflow for persistence.</p>
<h2 id="why-this-was-inevitable">why this was inevitable<a class="anchor" href="#why-this-was-inevitable" aria-label="link to this section">#</a></h2>
<p>Every precondition has been in place for years:</p>
<ul><li>Packages execute arbitrary code on install, by default.</li><li>Publishing credentials are commonly stored on developer machines in plaintext.</li><li>One credential frequently controls many packages.</li><li>Nothing in the publishing flow requires human presence.</li></ul>
<p>Given those four facts, a worm is not a clever attack. It is the obvious one. Security researchers have described this exact scenario in talks for the better part of a decade.</p>
<h2 id="the-immediate-response">the immediate response<a class="anchor" href="#the-immediate-response" aria-label="link to this section">#</a></h2>
<p>npm has accelerated changes it had already announced: shorter token lifetimes, mandatory two-factor for high-impact publishing, restrictions on classic tokens, and a strong push toward trusted publishing.</p>
<p>Those are the right changes. They are also the changes that were "coming" for years and are now arriving under emergency conditions, which is how infrastructure security usually improves.</p>
<h2 id="what-to-do-right-now">what to do right now<a class="anchor" href="#what-to-do-right-now" aria-label="link to this section">#</a></h2>
<p><strong>Rotate every npm token you have</strong>, especially any on a developer machine or in a CI secret store. Assume anything that was on a machine that ran an install during the window is exposed.</p>
<p><strong>Enable trusted publishing</strong> on every package you maintain. This removes the long-lived token entirely — publishing is authorized by an OIDC assertion from a specific workflow in a specific repository.</p>
<p><strong>Audit your GitHub account</strong> for repositories and workflows you did not create. The persistence mechanism was a workflow file; it survives credential rotation.</p>
<p><strong>Turn off install scripts</strong> everywhere you can. This remains the single highest value control and remains widely unused.</p>
<p><strong>Adopt a version cooldown.</strong> Every one of these incidents has a window between publication and detection measured in hours. A 72-hour delay before adopting new versions costs you nothing and avoids nearly all of them.</p>
<h2 id="the-structural-conclusion">the structural conclusion<a class="anchor" href="#the-structural-conclusion" aria-label="link to this section">#</a></h2>
<p>The npm ecosystem's install-time code execution is a design decision from 2010 that made sense when the registry had a few thousand packages maintained by people who mostly knew each other.</p>
<p>It does not make sense now, and every incident makes the argument more forcefully. Other ecosystems handle this differently — Go has no install-time execution at all, and it turns out you can build a package ecosystem without it.</p>
<p>The migration cost for npm is enormous and the cost of not migrating is this, repeatedly, with escalating sophistication. At some point the arithmetic flips. It may have just flipped.</p>]]></content:encoded></item><item><title>chalk and debug get compromised, and 2 billion weekly downloads flinch</title><link>https://readme.news/chalk-and-debug-get-compromised-and-2-billion-weekly-downloads-flinch/</link><guid isPermaLink="true">https://readme.news/chalk-and-debug-get-compromised-and-2-billion-weekly-downloads-flinch/</guid><pubDate>Tue, 09 Sep 2025 09:00:00 +0000</pubDate><description>A phishing email to a maintainer, eighteen packages, and a crypto-stealing payload in the browser.</description><content:encoded><![CDATA[<p>A maintainer of several extremely widely-used npm packages — including <code>chalk</code>, <code>debug</code>, <code>ansi-styles</code>, and <code>strip-ansi</code> — was phished, and malicious versions of eighteen packages were published.</p>
<p>Combined weekly download counts for the affected packages are on the order of two billion.</p>
<h2 id="the-phish">the phish<a class="anchor" href="#the-phish" aria-label="link to this section">#</a></h2>
<p>An email from <code>npmjs.help</code> — a lookalike domain — claiming the account required two-factor re-verification, with a link to a credential harvesting page that also captured the TOTP code.</p>
<p>The maintainer has written publicly about it. The email was well-constructed, it arrived at a plausible time, and the domain was close enough to pass a quick glance.</p>
<p>This is worth saying clearly: the target was a competent, security-aware, long-time open source maintainer. Phishing that captures a TOTP in real time defeats the second factor. The control that actually stops this is a hardware security key or passkey, where the authentication is bound to the origin and a lookalike domain simply cannot complete it.</p>
<p><strong>If you publish packages, use a passkey or hardware key. Today. TOTP is not sufficient against this attack and has not been for years.</strong></p>
<h2 id="the-payload">the payload<a class="anchor" href="#the-payload" aria-label="link to this section">#</a></h2>
<p>Browser-targeted, not server-targeted, which is unusual and clever.</p>
<p>The injected code hooked <code>window.ethereum</code> and intercepted <code>fetch</code> and <code>XMLHttpRequest</code>, watching for cryptocurrency transactions and swapping destination addresses for attacker-controlled ones — selecting a visually similar address to survive a casual glance at the confirmation dialog.</p>
<p>Because these packages are ubiquitous transitive dependencies of frontend build tooling, the payload had a plausible path into a very large number of shipped bundles.</p>
<p>Actual losses appear to have been small. Detection was fast — within a couple of hours — and most builds during the window did not pull the affected versions.</p>
<h2 id="why-the-damage-was-limited">why the damage was limited<a class="anchor" href="#why-the-damage-was-limited" aria-label="link to this section">#</a></h2>
<p>Three things, in order:</p>
<ol><li><strong>Lockfiles.</strong> Most production builds resolve to pinned versions. A new malicious release does not enter an existing lockfile without someone running an update.</li><li><strong>Speed of detection.</strong> The community noticed within hours. Someone diffed the published tarball against the repository and the mismatch was obvious.</li><li><strong>Payload specificity.</strong> Targeting crypto wallets is narrow. A payload targeting build-time credential theft would have done far more damage.</li></ol>
<p>That third point should not be comforting. The same access with a better payload would have been much worse.</p>
<h2 id="the-controls-again">the controls, again<a class="anchor" href="#the-controls-again" aria-label="link to this section">#</a></h2>
<p>The same short list as every one of these:</p>
<ul><li><strong>Phishing-resistant auth on publishing accounts.</strong> Non-negotiable.</li><li><strong>Trusted publishing via OIDC</strong> rather than long-lived tokens.</li><li><strong>Lockfiles, committed, with exact versions in production.</strong></li><li><strong><code>--ignore-scripts</code> in CI.</strong></li><li><strong>A cooldown before adopting new versions.</strong></li></ul>
<p>None of these are new. All of them are still not universally adopted, which is the actual story.</p>
<h2 id="the-structural-observation">the structural observation<a class="anchor" href="#the-structural-observation" aria-label="link to this section">#</a></h2>
<p><code>chalk</code> adds color to terminal output. It is about a thousand lines. It has two billion weekly downloads because it is a transitive dependency of nearly everything in the JavaScript ecosystem.</p>
<p>A thousand-line utility maintained by a volunteer sits in the trusted computing base of a substantial fraction of the world's software, and the ecosystem's security posture depends on that person's email hygiene.</p>
<p>That is not a criticism of the maintainer, who did nothing unreasonable. It is a description of a system that has grown a dependency structure nobody designed and nobody can now change.</p>
<p>The fix is not "fewer dependencies." The fix is publishing infrastructure where a single compromised human credential cannot ship code to two billion machines. The registry has to make that impossible, because asking every maintainer to be unphishable has failed for a decade.</p>]]></content:encoded></item><item><title>The Nx compromise and the credential-stealing postinstall</title><link>https://readme.news/the-nx-compromise-and-the-credential-stealing-postinstall/</link><guid isPermaLink="true">https://readme.news/the-nx-compromise-and-the-credential-stealing-postinstall/</guid><pubDate>Fri, 29 Aug 2025 09:00:00 +0000</pubDate><description>A popular build tool&#x27;s npm packages ship a payload that harvests tokens and pushes them to public repos.</description><content:encoded><![CDATA[<p>Malicious versions of the <code>nx</code> build tool and several related packages were published to npm with a <code>postinstall</code> script that scanned the developer's machine for credentials and published them to a public GitHub repository created under the victim's own account.</p>
<h2 id="the-payload">the payload<a class="anchor" href="#the-payload" aria-label="link to this section">#</a></h2>
<p>The script searched for:</p>
<ul><li>npm and GitHub tokens</li><li>SSH private keys</li><li>Cloud provider credentials in the usual locations</li><li>Environment variables matching credential-shaped patterns</li><li>Cryptocurrency wallet files</li></ul>
<p>It then created a public repository in the victim's GitHub account named with a recognizable prefix and pushed the harvested data to it.</p>
<p>Some variants additionally invoked locally-installed AI coding CLI tools with a prompt asking them to enumerate sensitive files — using the developer's own agent as a discovery mechanism. That detail is new and it is going to be studied.</p>
<h2 id="the-entry-point">the entry point<a class="anchor" href="#the-entry-point" aria-label="link to this section">#</a></h2>
<p>Compromised publishing credentials. The specifics of how they were obtained matter less than the pattern: a maintainer's token, or a CI workflow with publishing rights, was reachable by an attacker.</p>
<p>This is the dominant supply chain attack shape. Not typosquatting, not dependency confusion, not a malicious contribution. <strong>Take over the account of a legitimate maintainer of a package people already depend on.</strong></p>
<h2 id="the-controls-that-would-have-stopped-it">the controls that would have stopped it<a class="anchor" href="#the-controls-that-would-have-stopped-it" aria-label="link to this section">#</a></h2>
<p>In order of effectiveness:</p>
<p><strong>1. <code>--ignore-scripts</code>.</strong> The payload was in <code>postinstall</code>. An install that does not run lifecycle scripts does not run the payload.</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">npm ci --ignore-scripts</code></pre></div>
<p>Set it in <code>.npmrc</code> for your CI and see what breaks. For most projects the answer is: nothing, or one package that needs a native build, which you allowlist.</p>
<p><strong>2. Trusted publishing.</strong> OIDC-based publishing from a verified CI workflow rather than a long-lived token. A token that does not exist cannot be stolen. npm supports this now and adoption is the bottleneck.</p>
<p><strong>3. Install cooldown.</strong> The malicious versions were live for hours before removal. A policy of not adopting a version until it is 24 to 72 hours old would have avoided this entirely, and avoids most incidents of this shape, because the window between publish and detection is short.</p>
<div class="code"><span class="code-lang">json</span><pre><code class="lang-json">{ "minimumReleaseAge": 4320 }</code></pre></div>
<p><strong>4. Short-lived credentials.</strong> The payload harvested long-lived tokens. If your CI uses OIDC federation to get a fifteen-minute cloud credential, there is nothing durable to steal.</p>
<h2 id="if-you-were-affected">if you were affected<a class="anchor" href="#if-you-were-affected" aria-label="link to this section">#</a></h2>
<p>The order matters:</p>
<ol><li><strong>Rotate everything</strong> the machine had access to. npm tokens, GitHub PATs, SSH keys, cloud credentials. Assume everything readable was read.</li><li><strong>Check for the public repository</strong> in your GitHub account and delete it — but capture its contents first for your incident record.</li><li><strong>Audit for persistence.</strong> Check <code>~/.bashrc</code>, <code>~/.zshrc</code>, <code>~/.profile</code>, shell history, cron, launch agents, and git hooks in your repositories.</li><li><strong>Review recent activity</strong> on every account whose credentials were on that machine.</li></ol>
<p>Rotation before cleanup. Cleanup on a compromised machine while live credentials are still valid is how you spend an afternoon on remediation and the attacker spends it on access.</p>
<h2 id="the-trend">the trend<a class="anchor" href="#the-trend" aria-label="link to this section">#</a></h2>
<p>This is one of several npm incidents this year and they are getting more sophisticated: better targeting, cleaner payloads, and now the use of local AI tooling as an attack primitive.</p>
<p>The ecosystem's structural weakness has not changed. A single maintainer's credentials protect a package that millions of machines install and execute automatically. Every control above mitigates; none of them fix that.</p>
<p>The fix requires the registry to make unattended publishing hard by default, and that is a change to how a lot of people work. It is going to happen anyway, because the alternative is this every few weeks.</p>]]></content:encoded></item>
</channel>
</rss>
