<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — engineering-management</title>
<link>https://readme.news/tags/engineering-management/</link>
<atom:link href="https://readme.news/tags/engineering-management/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged engineering-management.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>How to run a spike</title><link>https://readme.news/how-to-run-a-spike/</link><guid isPermaLink="true">https://readme.news/how-to-run-a-spike/</guid><pubDate>Tue, 18 Aug 2026 09:00:00 +0000</pubDate><description>Timeboxed investigation is the cheapest way to buy certainty, and most teams do it in a way that produces neither certainty nor code.</description><content:encoded><![CDATA[<p>A spike is a timeboxed investigation whose output is a decision, not a feature. Done well it is the highest-leverage day in a project. Done the way most teams do it, it produces a half-finished prototype and an opinion nobody trusts.</p>
<p>The difference is entirely in the setup.</p>
<h2 id="write-the-question-down-first">write the question down first<a class="anchor" href="#write-the-question-down-first" aria-label="link to this section">#</a></h2>
<p>A spike without a written question becomes exploration, and exploration has no end condition.</p>
<p>The question has to be answerable and specific:</p>
<ul><li>Bad: "look into whether we can use X"</li><li>Good: "can X ingest 50k events/second on a single node with our event shape, and what does its failure behaviour look like when the disk fills?"</li></ul>
<p>The second one tells you when you are finished, what to measure, and what would count as a no. The first one runs until someone gets bored.</p>
<h2 id="state-the-decision-it-feeds">state the decision it feeds<a class="anchor" href="#state-the-decision-it-feeds" aria-label="link to this section">#</a></h2>
<p>A spike exists to unblock a decision. Name the decision explicitly: <em>we will use X for the ingest path, or we will stay with Y and shard it.</em></p>
<p>If you cannot name the decision, you are not running a spike, you are learning something — which is fine and should be called that, because it has a different budget.</p>
<h2 id="set-the-box-and-mean-it">set the box, and mean it<a class="anchor" href="#set-the-box-and-mean-it" aria-label="link to this section">#</a></h2>
<p>Two days is the usual right answer. Long enough to get past setup, short enough that being wrong is cheap.</p>
<p>The rule that makes it work: <strong>when the box ends, you stop and report, even if you are nearly there.</strong> "Nearly there" is where spikes go to become three-week projects. If the answer genuinely needs more time, that is itself a finding — report it and ask for a second box with a narrower question.</p>
<h2 id="the-output-is-a-document-not-a-branch">the output is a document, not a branch<a class="anchor" href="#the-output-is-a-document-not-a-branch" aria-label="link to this section">#</a></h2>
<p>This is the part teams skip. The deliverable is one page:</p>
<div class="code"><span class="code-lang">markdown</span><pre><code class="lang-markdown">## Question
Can X ingest 50k events/s on one node with our event shape?

## Answer
Yes for the steady state, no for our burst profile.
32k/s sustained, 61k/s for ~40s before the write buffer backs up.

## Evidence
Harness: spike/x-ingest/ (throwaway). Ran on c7g.4xlarge, 20 min,
production event sample from 2026-08-10. Numbers: spike/x-ingest/results.md

## What surprised me
Backpressure is silent — it drops rather than erroring, and the drop
counter is only in a debug endpoint. That is a production risk with X
regardless of throughput.

## Recommendation
Stay with Y and shard. Revisit X if their 3.x backpressure work lands.

## What I did not test
Multi-node, recovery after disk-full, or the managed offering.</code></pre></div>
<p>Six sections, twenty minutes to write, and it is still useful in a year when somebody asks why the decision went the way it did. The "what I did not test" section is the one that keeps the document honest.</p>
<h2 id="throw-the-code-away">throw the code away<a class="anchor" href="#throw-the-code-away" aria-label="link to this section">#</a></h2>
<p>Spike code is written to answer a question, not to be maintained. It has no tests, no error handling, hard-coded credentials and a <code>main</code> function that does everything.</p>
<p>That is correct — the speed comes from those omissions. The failure is letting it become the foundation of the real implementation, at which point you have shipped a prototype and will spend a year discovering what it does not handle.</p>
<p>Put it in a directory named <code>spike/</code>, or a branch you never merge. Say in the document where it is. Then write the real thing properly, informed by what you learned.</p>
<p>The strongest signal that a team's spikes are working: the spike code never ships.</p>
<h2 id="when-a-spike-is-the-wrong-tool">when a spike is the wrong tool<a class="anchor" href="#when-a-spike-is-the-wrong-tool" aria-label="link to this section">#</a></h2>
<ul><li><strong>When the answer is already written down.</strong> Read the documentation, the <a class="xref" href="/the-unreasonable-effectiveness-of-a-changelog/" title="The unreasonable effectiveness of a changelog">changelog</a>, the issue tracker. A surprising amount of "we need to try it" is answerable in an hour of reading.</li><li><strong>When the question is about production behaviour under real load.</strong> A spike cannot tell you that. A canary can.</li><li><strong>When the decision is reversible and cheap.</strong> Then just pick one. A spike to choose between two libraries you could swap in a day costs more than being wrong.</li></ul>
<p>Spikes are for expensive, hard-to-reverse decisions where the deciding information does not exist yet. That is a narrower set than it feels like, and it is exactly where two days is a bargain.</p>]]></content:encoded></item><item><title>What "done" actually means</title><link>https://readme.news/what-done-actually-means/</link><guid isPermaLink="true">https://readme.news/what-done-actually-means/</guid><pubDate>Sat, 08 Aug 2026 09:00:00 +0000</pubDate><description>Every team has a definition of done and most of them are a checklist nobody reads. Here is the version that changes behaviour.</description><content:encoded><![CDATA[<p>Ask five engineers on the same team whether a ticket is done and you will get three answers: the code is written, the code is merged, the code is in production. All three are defensible and they differ by days.</p>
<p>That gap is where most delivery confusion lives, and it is cheap to close.</p>
<h2 id="the-honest-definition">the honest definition<a class="anchor" href="#the-honest-definition" aria-label="link to this section">#</a></h2>
<p>A change is done when <strong>the outcome it was meant to produce is observable, and the team would find out if it stopped.</strong></p>
<p>Not "the code is merged." Merged code that is not deployed is inventory. Deployed code that nobody looked at is a hypothesis.</p>
<p>That framing sounds demanding and mostly is not, because for a large share of work the observation is trivial — the feature is behind a flag, the flag is on for the team, someone clicked it. The point is that <em>someone did</em>.</p>
<h2 id="the-checklist-that-actually-helps">the checklist that actually helps<a class="anchor" href="#the-checklist-that-actually-helps" aria-label="link to this section">#</a></h2>
<p>Most definition-of-done checklists fail because they list activities rather than properties. "Tests written" is an activity. "The change is covered by a test that would fail if it regressed" is a property, and it is checkable.</p>
<p>The version I have seen work is short:</p>
<ul><li><strong>The change does what the ticket asked</strong>, verified by someone other than the author — a test, a reviewer, or the person who asked for it.</li><li><strong>It is in production</strong>, or in a deliberate queue with a named release date.</li><li><strong>A regression would be caught</strong> by a test, a metric, or an alert. If none of those apply, that is a decision, not an oversight.</li><li><strong>Its failure mode is known.</strong> Someone can say what happens when it breaks.</li><li><strong>It is documented where the next person will look</strong> — which is usually the code, sometimes the runbook, rarely the wiki.</li><li><strong>The flag, the branch, and the dead code are cleaned up</strong>, or ticketed with a date.</li></ul>
<p>Six lines. If a change satisfies them, it is done in a way that survives the author going on holiday.</p>
<h2 id="the-two-failure-modes">the two failure modes<a class="anchor" href="#the-two-failure-modes" aria-label="link to this section">#</a></h2>
<p><strong>Done-done-done.</strong> Teams that discover the ambiguity often respond by inventing stages: "dev done", "QA done", "really done". This makes the confusion explicit without removing it, and adds ceremony. The fix is one definition, not three adjectives.</p>
<p><strong>Definition as theatre.</strong> A twenty-item checklist copied from a blog post, pasted into the wiki, referenced never. If your definition of done is not enforced by something — a pull request template, a merge check, a question someone actually asks in standup — it is decoration.</p>
<p>The test: pick a ticket closed last week and walk the list. If it fails two items and nobody noticed, the definition is not operating.</p>
<h2 id="the-part-about-flags-and-branches">the part about flags and branches<a class="anchor" href="#the-part-about-flags-and-branches" aria-label="link to this section">#</a></h2>
<p>The most common way "done" quietly is not: the feature ships behind a flag at 10%, works, and then nobody finishes the rollout. Three months later the flag is still at 10%, the old code path is still there, and the ticket has been closed since March.</p>
<p>That is not done. It is half-deployed with the cleanup unfunded, and it is how codebases accumulate the dual implementations that make everything else harder.</p>
<p><strong>Make the rollout part of the ticket, not a follow-up.</strong> A change is not finished when it is enabled for some users; it is finished when the decision has been made either way and the losing path is deleted.</p>
<h2 id="why-it-is-worth-the-argument">why it is worth the argument<a class="anchor" href="#why-it-is-worth-the-argument" aria-label="link to this section">#</a></h2>
<p>A shared definition of done is what makes "how much is left" answerable. Without it, the burn-down is measuring something nobody agrees on, estimates are uncomparable between people, and "almost done" means whatever the speaker wants.</p>
<p>It costs one conversation and a paragraph in the repository. It is the cheapest process improvement available to most teams and it is skipped because it sounds like process rather than engineering.</p>]]></content:encoded></item><item><title>The onboarding document that actually works</title><link>https://readme.news/the-onboarding-document-that-actually-works/</link><guid isPermaLink="true">https://readme.news/the-onboarding-document-that-actually-works/</guid><pubDate>Mon, 13 Jul 2026 09:00:00 +0000</pubDate><description>Most onboarding docs are written by people who already know. Here is what a new engineer needs on day one, in order.</description><content:encoded><![CDATA[<p>Onboarding documentation is written by someone who already knows the system, which means it is written from the wrong side of the knowledge gap.</p>
<p>The result is consistently: an architecture overview that is meaningless without context, a list of tools with no explanation of why, and no answer to any question a new person actually has.</p>
<h2 id="what-a-new-engineer-actually-needs-in-order">what a new engineer actually needs, in order<a class="anchor" href="#what-a-new-engineer-actually-needs-in-order" aria-label="link to this section">#</a></h2>
<p><strong>Day one: get something running.</strong></p>
<p>Not understanding. Running. A new engineer who has the application running locally on day one is in a completely different position from one who is still fighting a dependency on day three.</p>
<p>This section should be a numbered list of commands that work. Tested. On a clean machine. Recently.</p>
<div class="code"><span class="code-lang">markdown</span><pre><code class="lang-markdown">## get it running

1. Install prerequisites:  `brew bundle`   (or the equivalent — see below)
2. Copy the environment:   `cp .env.example .env`
3. Get secrets:            `./scripts/fetch-dev-secrets` (needs VPN)
4. Start dependencies:     `docker compose up -d`
5. Migrate:                `./scripts/db-reset`
6. Run:                    `pnpm dev`

You should see the app at http://localhost:5173 with seeded data.
Log in with dev@example.com / password.

If step 3 fails with "unauthorized", you need to be added to the
dev-secrets group — ask in #eng-onboarding.</code></pre></div>
<p>That last paragraph — the anticipated failure with its resolution — is the part that separates a document that works from one that does not. Every step that has ever failed for anyone should have a note.</p>
<p><strong>Day two: make a change and see it.</strong></p>
<p>A guided first change. Something trivial and real: change a label, add a field, fix a typo in a template. All the way through: edit, test, review, merge, deploy.</p>
<p>The point is not the change. It is that they have now executed the entire delivery pipeline once, so every subsequent change is a variation on something they have done.</p>
<p><strong>Day three to five: the map.</strong></p>
<p><em>Now</em> the architecture overview, and now it means something, because they have seen the system run.</p>
<p>Keep it to a page. What are the major pieces, what does each do, how do they talk. A diagram. Where the code for each piece lives.</p>
<p>Not a complete description. A map, at the resolution of "which building do I go to."</p>
<p><strong>Week two: the why.</strong></p>
<p>The decisions that are not obvious from the code. Why the database is structured that way. Why there is a queue between those two services. Why that module is frozen. Why the obvious approach to X does not work.</p>
<p>This is the highest-value and least-written documentation in any organization, because it exists only in the heads of people who were there. When they leave, it is gone, and the next person spends a year rediscovering it — usually by proposing the obvious approach and being told no.</p>
<h2 id="the-sections-everyone-forgets">the sections everyone forgets<a class="anchor" href="#the-sections-everyone-forgets" aria-label="link to this section">#</a></h2>
<p><strong>The glossary.</strong> Every organization has jargon: internal product names, acronyms, words used with a specific local meaning. A new person hears twenty of these in their first week and cannot ask about all of them without feeling stupid.</p>
<p>Write them down. This is a thirty-minute task with an outsized payoff.</p>
<p><strong>Who to ask about what.</strong> Not the org chart. "Payments: ask Priya. Deploy pipeline: #platform-help. Anything about the legacy importer: Marcus, and be warned it is complicated."</p>
<p><strong>The things that will surprise you.</strong> The test suite that fails on the first run until you seed a fixture. The service that takes four minutes to start. The one flaky test everyone knows about. The staging environment that resets on Sundays.</p>
<p>Every codebase has these. Writing them down converts "this is broken and I do not want to admit I cannot fix it" into "oh, that is expected."</p>
<p><strong>What not to touch.</strong> Frozen modules, generated files, anything requiring a specific review.</p>
<h2 id="the-maintenance-mechanism">the maintenance mechanism<a class="anchor" href="#the-maintenance-mechanism" aria-label="link to this section">#</a></h2>
<p>Onboarding docs rot faster than any other documentation, because the people who would notice the errors are the ones who no longer read it.</p>
<p><strong>The fix: the last person onboarded owns it.</strong></p>
<p>Every new engineer's first task is to follow the document, fix everything that was wrong, and add the things they had to ask about. Then they own it until the next person arrives.</p>
<p>This works because their frustration is fresh and their perspective is exactly the target reader's. It is the only mechanism I have seen that keeps these documents accurate.</p>
<h2 id="the-test">the test<a class="anchor" href="#the-test" aria-label="link to this section">#</a></h2>
<p>Hand it to a new engineer and do not help them.</p>
<p>Time how long until they have the application running. Note every question they had to ask. Each question is a gap, with a measured cost.</p>
<p>Then fix them, and repeat with the next person.</p>
<p>Most teams have never done this and would be surprised by the result.</p>]]></content:encoded></item><item><title>What a staff engineer does all day</title><link>https://readme.news/what-a-staff-engineer-does-all-day/</link><guid isPermaLink="true">https://readme.news/what-a-staff-engineer-does-all-day/</guid><pubDate>Fri, 19 Jun 2026 09:00:00 +0000</pubDate><description>The role is genuinely ambiguous and that ambiguity is load-bearing. An attempt at a concrete description.</description><content:encoded><![CDATA[<p>"Staff engineer" is the most poorly-defined common title in the industry. It means different things at different companies, and even within one company two staff engineers may do almost nothing in common.</p>
<p>Here is an attempt at what the role actually is, based on watching people do it well.</p>
<h2 id="what-it-is-not">what it is not<a class="anchor" href="#what-it-is-not" aria-label="link to this section">#</a></h2>
<p><strong>Not "a senior engineer who has been there longer."</strong> Time in seat produces a senior engineer with more context, which is valuable and is not this.</p>
<p><strong>Not "the best coder."</strong> The highest-output individual contributor is a valuable role and it is a different one. Staff engineers frequently write less code than seniors.</p>
<p><strong>Not "a manager who did not want to manage."</strong> The scope is comparable to a manager's; the mechanism is entirely different.</p>
<h2 id="what-it-is">what it is<a class="anchor" href="#what-it-is" aria-label="link to this section">#</a></h2>
<p>The concise version: <strong>a staff engineer is responsible for the technical success of work that spans more than one team, without having authority over those teams.</strong></p>
<p>Everything distinctive about the role follows from "without authority." You cannot assign work. You cannot approve headcount. You cannot make a decision stick by deciding it. Every outcome has to be achieved through information, credibility, and persuasion.</p>
<h2 id="the-actual-activities">the actual activities<a class="anchor" href="#the-actual-activities" aria-label="link to this section">#</a></h2>
<p><strong>Finding the problem nobody owns.</strong> The most common form of staff-level impact. Every organization has problems that fall between teams: the shared library nobody maintains, the integration that fails and each side thinks is the other's, the performance issue that is caused by the interaction of three services.</p>
<p>Nobody owns these, so nobody fixes them, so they persist for years. Identifying one, proving it matters, and getting it fixed is a substantial contribution and it usually requires a person who is not on any of the teams involved.</p>
<p><strong>Making a decision that spans teams.</strong> Which of three approaches, when each team has a preference and none of them can see the whole picture. This is where the "no authority" constraint bites hardest: the decision has to be made in a way that the people affected accept, which means the reasoning has to be visible and the objections have to be genuinely addressed.</p>
<p><strong>Writing the document that ends the argument.</strong> A recurring disagreement that resurfaces every quarter because nobody wrote down the resolution. One good document — with the options, the trade-offs, the decision, and the reasoning — can end a multi-year debate.</p>
<p><strong>Being the person who read the whole system.</strong> Most engineers know their team's code. Somebody needs to know how it all fits together, including the parts nobody has touched in three years. That knowledge is what makes the cross-cutting problems visible.</p>
<p><strong>Raising the floor.</strong> Not by writing better code. By making the good pattern easy: a library that removes a class of bug, a template that encodes the right defaults, a lint rule that prevents the mistake, documentation that means nobody has to ask. Impact through leverage rather than output.</p>
<p><strong>Mentoring, specifically on judgment.</strong> Not "how do I use this API." "Should we build this at all," "how do I disagree with my manager about a technical decision," "how do I tell if this design will be a problem in a year."</p>
<p><strong>Saying no with a reason, at a level where it lands.</strong> Frequently the most valuable thing a staff engineer does, and the reason it requires seniority is that the no has to come with an alternative and with credibility behind it.</p>
<h2 id="the-day-concretely">the day, concretely<a class="anchor" href="#the-day-concretely" aria-label="link to this section">#</a></h2>
<p>A real week looks roughly like:</p>
<ul><li>20% writing code, usually the hard or risky part of something, or a prototype that settles an argument.</li><li>25% writing documents — designs, decisions, analyses, postmortems.</li><li>25% in conversations — reviews, one-on-ones, arguing about designs, being asked "does this seem right to you."</li><li>15% reading — code, incident reports, other people's designs, the thing everyone is complaining about.</li><li>15% on whatever is currently on fire.</li></ul>
<p>The proportion of coding is the part that surprises people moving into the role, and the discomfort of not shipping visible code is the most common reason people bounce out of it.</p>
<h2 id="how-to-tell-if-someone-is-good-at-it">how to tell if someone is good at it<a class="anchor" href="#how-to-tell-if-someone-is-good-at-it" aria-label="link to this section">#</a></h2>
<p><strong>Do things get decided?</strong> Not "do they have opinions." Do arguments they are involved in reach a resolution that holds.</p>
<p><strong>Do other engineers get better?</strong> Look at the people around them over a year.</p>
<p><strong>Do they work on things nobody asked them to?</strong> The highest-value staff work is usually self-directed, because if it were obvious and assigned it would already have an owner.</p>
<p><strong>Are they trusted by people who disagree with them?</strong> This is the real test. Someone who is only trusted by people who already agree has influence, not credibility.</p>
<h2 id="the-failure-modes">the failure modes<a class="anchor" href="#the-failure-modes" aria-label="link to this section">#</a></h2>
<p><strong>Becoming an architecture astronaut.</strong> Designs, diagrams, opinions, no contact with running code. The credibility runs out within about a year, and it is very hard to get back.</p>
<p><strong>Becoming a very expensive senior engineer.</strong> Doing excellent work with single-team scope. Comfortable, valuable, and not the job.</p>
<p><strong>Spreading too thin.</strong> Involved in twelve things, effective in none. The scope is tempting and the constraint is real: two or three significant efforts at a time is the realistic maximum.</p>
<p><strong>Losing the ability to build.</strong> The role requires enough hands-on work to stay credible and calibrated. An engineer who has not shipped in a year is guessing.</p>]]></content:encoded></item><item><title>Platform teams that don't get resented</title><link>https://readme.news/platform-teams-that-dont-get-resented/</link><guid isPermaLink="true">https://readme.news/platform-teams-that-dont-get-resented/</guid><pubDate>Fri, 22 May 2026 09:00:00 +0000</pubDate><description>Internal platforms fail for predictable reasons. The successful ones share four properties.</description><content:encoded><![CDATA[<p>Most internal platform teams end up resented by the engineers they serve. The pattern is consistent enough that the causes are identifiable.</p>
<h2 id="the-failure-pattern">the failure pattern<a class="anchor" href="#the-failure-pattern" aria-label="link to this section">#</a></h2>
<ol><li>Platform team forms to reduce duplicated infrastructure work.</li><li>They build an abstraction over the cloud provider.</li><li>The abstraction covers 80% of cases well.</li><li>The remaining 20% is impossible, and the escape hatch is either absent or punished.</li><li>Product teams work around the platform.</li><li>Platform team responds by mandating the platform.</li><li>Everyone is unhappy and the platform is now a tax.</li></ol>
<p>Every step follows from the previous one. The root is step four.</p>
<h2 id="the-four-properties-of-platforms-that-work">the four properties of platforms that work<a class="anchor" href="#the-four-properties-of-platforms-that-work" aria-label="link to this section">#</a></h2>
<p><strong>1. An escape hatch that is not punished.</strong></p>
<p>The platform covers the common case. It cannot cover every case, and pretending otherwise is what breaks trust.</p>
<p>There must be a supported path for "I need something the platform does not do," and taking that path must not require an exception process, a meeting, or an apologetic Slack message.</p>
<p>The best platforms make the escape hatch cheap and then compete on being better than it. The worst make it forbidden, which does not eliminate the need — it drives it underground.</p>
<p><strong>2. Adoption is voluntary, at least at first.</strong></p>
<p>A platform that teams choose is a platform that is good. A platform teams are required to use never gets the feedback that would make it good, because the feedback mechanism — people leaving — has been disabled.</p>
<p>If you cannot get voluntary adoption, that is information. Mandating it does not fix the underlying problem; it hides it and converts a product problem into a political one.</p>
<p>Mandate later, when it is genuinely better, and the mandate will be uncontroversial because everyone already uses it.</p>
<p><strong>3. The abstraction leaks deliberately, not accidentally.</strong></p>
<p>Every abstraction leaks. The question is whether you planned for it.</p>
<p>A good platform lets you drop a level when you need to: use the paved path for the deployment, and reach the underlying resource directly when you need something specific. A bad one hides the underlying system entirely, so that when it fails you cannot debug it and neither can the platform team, because now there are two systems to understand.</p>
<p><strong>Concretely:</strong> if your platform generates infrastructure configuration, let people see it. If it wraps a cloud API, let people access the underlying resource. If it runs their container, give them the logs from the actual runtime, not a filtered view.</p>
<p><strong>4. The platform team is measured on adoption and satisfaction, not on compliance.</strong></p>
<p>If the platform team's metric is "percentage of services on the platform," they will optimize for mandating it.</p>
<p>If the metric is "would you use this if you had a choice," they will optimize for making it good.</p>
<p>Ask that question quarterly, anonymously, and publish the answer.</p>
<h2 id="the-specific-things-that-generate-resentment">the specific things that generate resentment<a class="anchor" href="#the-specific-things-that-generate-resentment" aria-label="link to this section">#</a></h2>
<p><strong>Slow escape.</strong> A team needs something the platform does not support. The answer is "file a request, we will look at it next quarter." Their deadline is Friday.</p>
<p><strong>Breaking changes without migration paths.</strong> The platform is infrastructure. Break it and every team stops. Platform teams frequently hold themselves to a lower compatibility standard than they would accept from a vendor.</p>
<p><strong>Opaque failures.</strong> The deploy failed. The error is a platform-internal message. The product engineer cannot debug it and must escalate, which means waiting.</p>
<p><strong>Being a gate rather than a service.</strong> A platform that must approve things is a bureaucracy. A platform that makes the right thing easy is infrastructure.</p>
<p><strong>Solving the platform team's problems.</strong> Standardization is valuable to the platform team and is not automatically valuable to product teams. If the pitch for a migration is "this makes our lives easier," expect a cool reception.</p>
<h2 id="the-framing-that-works">the framing that works<a class="anchor" href="#the-framing-that-works" aria-label="link to this section">#</a></h2>
<p><strong>You are building a product. Your users are engineers. They have alternatives.</strong></p>
<p>That framing produces the right behaviors automatically: user research before building, documentation that assumes nothing, onboarding that works, support that responds, and a roadmap driven by what users need rather than by architectural preference.</p>
<p>The platform teams I have seen work best behave exactly like a startup selling to a skeptical market, and they say so out loud.</p>
<p>The ones that fail behave like an internal standards body, and they are usually correct about the standards and wrong about how to get them adopted.</p>
<h2 id="the-measurement-that-matters">the measurement that matters<a class="anchor" href="#the-measurement-that-matters" aria-label="link to this section">#</a></h2>
<p>Time from "a new engineer joins" to "their code is running in production."</p>
<p>That single number captures most of what a platform is for, it is measurable, and it is the thing product teams actually care about. If it is going down, the platform is working, regardless of what the adoption <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> says.</p>]]></content:encoded></item><item><title>Hiring juniors in 2026</title><link>https://readme.news/hiring-juniors-in-2026/</link><guid isPermaLink="true">https://readme.news/hiring-juniors-in-2026/</guid><pubDate>Thu, 30 Apr 2026 09:00:00 +0000</pubDate><description>The entry-level pipeline is breaking in a way that will be expensive in five years. Some of the fixes are cheap.</description><content:encoded><![CDATA[<p>Entry-level software hiring has contracted sharply. The reasons are partly cyclical and partly a genuine belief among hiring managers that AI tooling has reduced the need for junior engineers.</p>
<p>The first part will recover. The second part is a mistake, and it is the kind of mistake that is invisible for four years and then extremely expensive.</p>
<h2 id="the-argument-for-not-hiring-juniors">the argument for not hiring juniors<a class="anchor" href="#the-argument-for-not-hiring-juniors" aria-label="link to this section">#</a></h2>
<p>Stated honestly, because it is not stupid:</p>
<p>A junior engineer's first-year output is mostly small, well-specified tasks — the CRUD endpoint, the test coverage, the bug in a file someone pointed them at. That work is now substantially automatable. Meanwhile the junior requires mentoring time from senior engineers, which is the scarcest resource.</p>
<p>So the ROI on a junior looks worse than it did. That reasoning is coherent.</p>
<h2 id="why-it-is-wrong">why it is wrong<a class="anchor" href="#why-it-is-wrong" aria-label="link to this section">#</a></h2>
<p><strong>Seniors come from juniors.</strong> There is no other supply. An organization that hires only seniors is free-riding on other organizations' training, and if everyone does it, the pipeline empties. This is a classic collective action failure and the industry is walking into it with open eyes.</p>
<p><strong>The judgment that makes seniors valuable comes from doing the work.</strong> The ability to look at plausible code and know it is wrong comes from having written the wrong version and debugged it at 3 a.m. You cannot read your way to it and you cannot prompt your way to it.</p>
<p>If the apprenticeship stops, the next generation of senior engineers does not exist, and the current one retires.</p>
<p><strong>Juniors are better at the new tools.</strong> Consistently, in my experience. They have no prior workflow to defend and they explore. A team of only senior engineers adopts new tooling slowly and grudgingly.</p>
<p><strong>Mentoring makes seniors better.</strong> The engineer who has to explain why a design is wrong understands it better afterward. Teams with no juniors lose that forcing function and get sloppier about articulating their own reasoning.</p>
<h2 id="what-actually-has-to-change">what actually has to change<a class="anchor" href="#what-actually-has-to-change" aria-label="link to this section">#</a></h2>
<p>The old model — hire a junior, give them small tickets for a year, gradually increase scope — does not work as well when the small tickets are automated. The model has to change, not the hiring.</p>
<p><strong>Start them on reading, not writing.</strong> Give a new engineer a real system and a week to understand and explain it. Have them write the architecture document that does not exist. This builds the skill that actually matters now and it produces something useful.</p>
<p><strong>Give them debugging, not features.</strong> Debugging is the skill that generalizes, that AI is least reliable at, and that cannot be learned from a course. Pair them on incidents. Give them the flaky test nobody wants.</p>
<p><strong>Make them review agent output.</strong> Reviewing machine-generated code with a senior engineer walking through what is wrong with it is an extraordinarily efficient teaching mechanism. You get a stream of plausible-but-flawed code, which is exactly the training material you want and which used to be expensive to produce.</p>
<p><strong>Require them to write the tests first.</strong> Specifying behavior before implementing teaches design, and it is the part of the workflow that has become more important rather than less.</p>
<p><strong>Do not let them delegate the hard part.</strong> For the first year, some things get done by hand, deliberately, because the point is the learning rather than the output. Say this out loud so it does not feel like an arbitrary restriction.</p>
<h2 id="the-hiring-signal-that-works-now">the hiring signal that works now<a class="anchor" href="#the-hiring-signal-that-works-now" aria-label="link to this section">#</a></h2>
<p>Traditional junior screens — implement this algorithm, complete this take-home — are substantially defeated and were never good predictors anyway.</p>
<p>What works better:</p>
<p><strong>A code review exercise.</strong> Give them a pull request with three problems. Watch what they find and how they talk about it. This is the job.</p>
<p><strong>A debugging exercise on a real repository.</strong> Failing test, thirty minutes, any tools they want including AI. Watch the process, not the outcome. Do they read the error? Form a hypothesis? Check it? Notice when the model's suggestion is wrong?</p>
<p><strong>A conversation about something they built.</strong> Follow-up questions until you hit the edge of their understanding. Where that edge sits, and how they handle reaching it, tells you almost everything.</p>
<h2 id="the-case-to-make-internally">the case to make internally<a class="anchor" href="#the-case-to-make-internally" aria-label="link to this section">#</a></h2>
<p>If you are arguing for junior headcount:</p>
<p>The cost of a junior is roughly a senior's partial attention for a year plus a below-market salary. The cost of a senior hire in three years, in a market where nobody trained anyone, is going to be considerably higher than it is now.</p>
<p>Every organization that stopped training in 2009 spent 2013 through 2016 paying enormous premiums for the engineers who had been trained elsewhere. It is the same trade and it is being made again.</p>
<p>The organizations that keep training through this period will have a meaningful advantage in five years, and it will be very hard to catch up to them quickly.</p>]]></content:encoded></item><item><title>The cost of a meeting, in engineering terms</title><link>https://readme.news/the-cost-of-a-meeting-in-engineering-terms/</link><guid isPermaLink="true">https://readme.news/the-cost-of-a-meeting-in-engineering-terms/</guid><pubDate>Wed, 11 Mar 2026 09:00:00 +0000</pubDate><description>Not the hourly rate. The fragmentation. Here&#x27;s why a 30-minute meeting costs four hours.</description><content:encoded><![CDATA[<p>The standard argument against meetings is arithmetic: eight people, thirty minutes, four person-hours, multiply by salary.</p>
<p>That understates it by a large factor, and the reason is the same reason context switching is expensive for a CPU.</p>
<h2 id="the-fragmentation-cost">the fragmentation cost<a class="anchor" href="#the-fragmentation-cost" aria-label="link to this section">#</a></h2>
<p>Deep engineering work requires holding a system in your head: the call graph, the invariants, the specific thing you were about to check, the four hypotheses you had narrowed to two.</p>
<p>Building that state takes time — call it twenty to forty minutes for nontrivial work. Losing it takes one interruption.</p>
<p>So a meeting at 2 p.m. does not cost thirty minutes. It costs:</p>
<ul><li>The twenty minutes before it, where you cannot start anything substantial because you will be interrupted.</li><li>The thirty minutes of the meeting.</li><li>The twenty to forty minutes after it, rebuilding the state you lost.</li></ul>
<p>That is roughly ninety minutes for a thirty-minute meeting, per person, and it is the optimistic case.</p>
<p>Worse: a meeting placed in the middle of a morning does not remove ninety minutes from a four-hour block. It removes the <em>block</em>. Two ninety-minute fragments are not equivalent to one three-hour stretch, because the hardest work requires depth that ninety minutes cannot reach.</p>
<h2 id="the-schedule-shapes-that-work">the schedule shapes that work<a class="anchor" href="#the-schedule-shapes-that-work" aria-label="link to this section">#</a></h2>
<p><strong>Meeting-free days.</strong> Two full days a week with no recurring meetings. Not "try to keep them clear" — a calendar policy. This is the single highest-impact scheduling change available and it is nearly free.</p>
<p><strong>Meetings at the edges.</strong> Cluster them at the start or end of the day. A day with three meetings from 9 to 11 and nothing after is dramatically more productive than one with three meetings at 10, 1, and 3.</p>
<p><strong>Default to 25 and 50 minutes.</strong> Not 30 and 60. The buffer prevents the cascade where every meeting starts late and the person coming from the previous one misses the first five minutes.</p>
<p><strong>Async by default for status.</strong> Standup is a meeting to exchange information that could be a written update. The argument for synchronous standup is that it surfaces blockers, which it does, and a written update with a "blocked on" field surfaces them too, in a searchable form, without costing everyone their morning.</p>
<h2 id="when-meetings-are-actually-right">when meetings are actually right<a class="anchor" href="#when-meetings-are-actually-right" aria-label="link to this section">#</a></h2>
<p>They are, frequently, and the anti-meeting position overcorrects.</p>
<p><strong>High-bandwidth disagreement.</strong> Two people who disagree about a design will resolve it in twenty minutes of conversation and in nine days of comment threads. Real-time is enormously better for anything with back-and-forth.</p>
<p><strong>Ambiguity resolution.</strong> When nobody is sure what the problem even is, a conversation converges much faster than writing, because you can ask the clarifying question immediately.</p>
<p><strong>Anything with emotional content.</strong> Feedback, conflict, bad news. Writing is the wrong medium and using it is usually avoidance.</p>
<p><strong>Building relationships.</strong> Real, unmeasurable, and the reason fully-async organizations struggle in ways they cannot diagnose.</p>
<p>The rule: <strong>meet for things that require interaction. Write for things that require information.</strong></p>
<h2 id="the-tests">the tests<a class="anchor" href="#the-tests" aria-label="link to this section">#</a></h2>
<p>Before scheduling, three questions:</p>
<p><strong>What decision will be made?</strong> If none, this is a status update. Write it.</p>
<p><strong>Who must be there for that decision?</strong> The answer is usually two to four people. Everyone else can read the outcome.</p>
<p><strong>What would happen if we did not have it?</strong> If the answer is "nothing," you have found a recurring meeting that outlived its purpose. Every organization has several.</p>
<h2 id="the-audit-worth-doing">the audit worth doing<a class="anchor" href="#the-audit-worth-doing" aria-label="link to this section">#</a></h2>
<p>Take a recurring meeting and cancel it for a month. Not "make it optional" — cancel it.</p>
<p>If someone notices and asks for it back, with a reason, reinstate it. Most of the time nobody notices, which tells you what you needed to know.</p>
<p>This works because recurring meetings are created for a reason that expires and never re-evaluated, and the social cost of proposing cancellation is high enough that nobody does it. A trial cancellation removes the social cost.</p>
<h2 id="the-part-managers-should-hear">the part managers should hear<a class="anchor" href="#the-part-managers-should-hear" aria-label="link to this section">#</a></h2>
<p>Your calendar is not your team's calendar. A manager's day is legitimately made of meetings — that is the job, and thirty-minute chunks are the natural unit.</p>
<p>An engineer's day is not, and scheduling as if it were is the single most common way well-meaning managers destroy their team's output while looking at metrics that do not show it.</p>
<p>Protect the blocks. It is most of what schedule-level management can do.</p>]]></content:encoded></item><item><title>The engineer's guide to saying no</title><link>https://readme.news/the-engineers-guide-to-saying-no/</link><guid isPermaLink="true">https://readme.news/the-engineers-guide-to-saying-no/</guid><pubDate>Sat, 28 Feb 2026 09:00:00 +0000</pubDate><description>Refusing work badly is a career problem. Refusing it well is one of the most valuable things a senior engineer does.</description><content:encoded><![CDATA[<p>Most engineers are bad at saying no. They either cannot do it — and end up with a commitment they cannot meet — or they do it in a way that reads as obstruction, and get routed around.</p>
<p>Both failures come from the same mistake: treating "no" as a verdict rather than as the opening of a conversation about trade-offs.</p>
<h2 id="what-the-request-actually-is">what the request actually is<a class="anchor" href="#what-the-request-actually-is" aria-label="link to this section">#</a></h2>
<p>When someone asks you to build something, they are not asking for the thing. They are asking for an outcome, and the thing is their guess at how to get it.</p>
<p>That means the highest-value response is frequently not yes or no. It is a better guess.</p>
<blockquote><p>"You are asking for a real-time <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>. What decision are you going to make from it? If it is 'should we page someone,' an alert is better than a dashboard and it is a day of work instead of a month."</p></blockquote>
<p>That is a no to the request and a yes to the outcome, and nobody experiences it as obstruction.</p>
<p>Ask what the outcome is before you evaluate the request. Half the time the request evaporates.</p>
<h2 id="the-four-honest-noes">the four honest noes<a class="anchor" href="#the-four-honest-noes" aria-label="link to this section">#</a></h2>
<p><strong>"Not this, that."</strong> The alternative approach that achieves the same outcome for less. This is the best one and it requires actually understanding the problem.</p>
<p><strong>"Yes, and here is what it displaces."</strong> Not a no. A statement about capacity, which is factual and is the other person's decision to make.</p>
<blockquote><p>"I can do that in this cycle. It means the API migration slips to next quarter. Which do you want?"</p></blockquote>
<p>This is enormously more effective than "we do not have time," because it hands the prioritization decision to the person whose job it is, with the information they need.</p>
<p><strong>"Yes, after X."</strong> A sequencing objection. "We can build this on top of the new data model. Building it on the old one means we build it twice."</p>
<p><strong>"No, and here is the risk I am flagging."</strong> The real no, reserved for correctness, security, legal, or ethical problems. Use it rarely so that it lands when you do.</p>
<blockquote><p>"I am not going to implement this the way it is specified because it stores plaintext credentials. I will implement it with a token exchange, which takes three extra days. If that is unacceptable, I want the decision documented and made by someone who can accept the risk."</p></blockquote>
<p>That is a hard sentence to say and it is the sentence the job sometimes requires.</p>
<h2 id="the-ones-that-do-not-work">the ones that do not work<a class="anchor" href="#the-ones-that-do-not-work" aria-label="link to this section">#</a></h2>
<p><strong>"That is not possible."</strong> Almost always false, and the person will find someone who says it is. Say "that would take six months" instead, which is the real constraint and is checkable.</p>
<p><strong>"That is a bad idea."</strong> Without an alternative, this is just friction.</p>
<p><strong>Silence.</strong> Not responding is a no that damages trust, because it looks like you did not care rather than that you disagreed.</p>
<p><strong>"Sure"</strong> followed by not doing it. The worst one. It destroys your reliability, which is the only currency you actually have.</p>
<h2 id="the-timing">the timing<a class="anchor" href="#the-timing" aria-label="link to this section">#</a></h2>
<p>Say no early. The cost of a no rises with every day of planning that assumed a yes.</p>
<p>An objection raised in the design review is a discussion. The same objection raised two weeks before launch is a crisis, and people will remember that you could have said it earlier — correctly.</p>
<p>If you have doubts, voice them while they are cheap.</p>
<h2 id="the-part-about-capital">the part about capital<a class="anchor" href="#the-part-about-capital" aria-label="link to this section">#</a></h2>
<p>Every no spends something. Every yes earns something. If you never say yes, your noes stop landing, because you have become the person who says no.</p>
<p>Deliver reliably on what you agree to. That is what makes the refusal credible when it matters. Engineers who are trusted to ship get enormous latitude to push back, and engineers who are not, do not — regardless of whether they are right.</p>
<h2 id="the-hardest-case">the hardest case<a class="anchor" href="#the-hardest-case" aria-label="link to this section">#</a></h2>
<p>Sometimes you are overruled on something you believe is wrong, and it is not a correctness or ethics issue — it is a judgment call and someone with the authority made a different one.</p>
<p>Disagree and commit is the right practice here and it is genuinely hard. Say your piece once, clearly, in writing. Then build the thing well.</p>
<p>Do not build it badly to prove a point. Do not relitigate it in every standup. Do not say "I told you so" if it goes wrong — the written record already said it, and gloating costs you the ability to be listened to next time.</p>
<p>You will be wrong about some of these. That is the actual reason to commit gracefully: you are not always right, and a culture where disagreement is followed by good-faith execution is one where being wrong is survivable for everyone, including you.</p>]]></content:encoded></item><item><title>On-call is a design problem</title><link>https://readme.news/on-call-is-a-design-problem/</link><guid isPermaLink="true">https://readme.news/on-call-is-a-design-problem/</guid><pubDate>Wed, 25 Feb 2026 09:00:00 +0000</pubDate><description>If your rotation is painful, that&#x27;s information about your architecture, not about your people&#x27;s resilience.</description><content:encoded><![CDATA[<p>Bad on-call is treated as a fact of life, a personal endurance test, or a staffing problem. It is none of those. It is a measurement of your system's design, delivered directly to your team's sleep.</p>
<h2 id="what-the-pages-are-telling-you">what the pages are telling you<a class="anchor" href="#what-the-pages-are-telling-you" aria-label="link to this section">#</a></h2>
<p>Every page is a statement about the system.</p>
<p><strong>A page that requires a human to restart something</strong> says the system cannot recover from a condition it will encounter again. That is automatable and the automation is usually a supervisor and a <a class="xref" href="/your-monitoring-is-measuring-the-wrong-nines/" title="Your monitoring is measuring the wrong nines">health check</a>.</p>
<p><strong>A page for a transient issue that resolved itself</strong> says your alert threshold is wrong, or your alert lacks a duration condition, or the underlying flakiness needs a retry with backoff.</p>
<p><strong>A page where the runbook is "look at the <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> and see if it is bad"</strong> is not a page. That is a dashboard, and it should be checked during business hours.</p>
<p><strong>A page that only one person can resolve</strong> is a knowledge distribution failure and a single point of human failure. The fix is documentation and rotation of the work, not heroism.</p>
<p><strong>A page at 3 a.m. for something that could have waited until 9 a.m.</strong> says nobody has classified alerts by urgency. Not everything that is wrong is urgent.</p>
<h2 id="the-audit">the audit<a class="anchor" href="#the-audit" aria-label="link to this section">#</a></h2>
<p>Pull the last ninety days of pages. For each one:</p>
<ul><li>What was the user impact? (Often: none.)</li><li>Could it have been automated?</li><li>Could it have waited?</li><li>Did the runbook exist and was it correct?</li><li>What was the actual fix?</li></ul>
<p>Then categorize:</p>
<p><strong>Should not have paged</strong> — no user impact or no urgency. Delete the alert or downgrade it. This is usually a large fraction and deleting alerts is the highest value hour available to most teams.</p>
<p><strong>Should have been automatic</strong> — the fix was mechanical. Automate it. If the runbook says "run this command," a computer can run that command.</p>
<p><strong>Genuine incidents</strong> — real user impact requiring judgment. These are the ones on-call exists for, and there should not be many.</p>
<p>A healthy rotation is a small number of genuine incidents. If you are getting paged more than a couple of times per week, the problem is not that your systems are complex.</p>
<h2 id="the-design-changes-that-reduce-pages">the design changes that reduce pages<a class="anchor" href="#the-design-changes-that-reduce-pages" aria-label="link to this section">#</a></h2>
<p><strong>Make things self-healing.</strong> Restart on crash. Retry with exponential backoff and jitter. Circuit-break to a degraded mode. Shed load rather than falling over. Each of these converts a page into a metric.</p>
<p><strong>Make degradation graceful and explicit.</strong> Decide in advance which features can be disabled. <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">Feature flags</a> that turn off the expensive path let a page become "turn off recommendations, investigate Monday."</p>
<p><strong>Make everything reversible fast.</strong> A large fraction of incidents are caused by a deploy. If rollback is one command and takes ninety seconds, the incident is ninety seconds long. If it requires reversing a migration, it is four hours.</p>
<p><strong>Add capacity headroom.</strong> Running at 85% utilization to save money means every spike is a page. Headroom is cheaper than the human cost, and much cheaper than the turnover.</p>
<h2 id="the-human-side">the human side<a class="anchor" href="#the-human-side" aria-label="link to this section">#</a></h2>
<p><strong>Pay for it.</strong> On-call is work performed outside working hours with real cost to the person's life. Compensate it explicitly — money or time off, and enough that the cost is visible to whoever decides whether to fix the alert noise.</p>
<p>Unpaid on-call means the cost is borne entirely by the person and is invisible to the organization, which guarantees it never improves.</p>
<p><strong>Time off after a bad night.</strong> Non-negotiable, automatic, not something the person has to ask for.</p>
<p><strong>Never one person.</strong> A primary and a secondary, always. The primary needs to be able to escalate without feeling like a failure.</p>
<p><strong>Rotate wide.</strong> If only three people can be on-call, they will burn out and leave, and the knowledge leaves with them.</p>
<p><strong>The person who was paged decides what gets fixed.</strong> Give the on-call engineer authority to prioritize the follow-up work. They have the information and the motivation, and nobody else has both.</p>
<h2 id="the-metric-to-track">the metric to track<a class="anchor" href="#the-metric-to-track" aria-label="link to this section">#</a></h2>
<p>Pages per on-call shift, trended over quarters.</p>
<p>If it is flat or rising, nobody is fixing the causes. That is a management failure, not a resilience failure, and the outcome of ignoring it is attrition — which is more expensive than every fix combined.</p>]]></content:encoded></item><item><title>Documentation is a product and you should staff it like one</title><link>https://readme.news/documentation-is-a-product-and-you-should-staff-it-like-one/</link><guid isPermaLink="true">https://readme.news/documentation-is-a-product-and-you-should-staff-it-like-one/</guid><pubDate>Fri, 20 Feb 2026 09:00:00 +0000</pubDate><description>Every team says docs matter. Almost none of them assign an owner, a budget, or a metric.</description><content:encoded><![CDATA[<p>Ask any engineering team whether documentation matters and they will say yes. Ask who owns it and you will get a pause.</p>
<p>That pause is the entire problem.</p>
<h2 id="the-four-kinds-and-why-mixing-them-fails">the four kinds, and why mixing them fails<a class="anchor" href="#the-four-kinds-and-why-mixing-them-fails" aria-label="link to this section">#</a></h2>
<p>The taxonomy that fixed documentation for me — and it is not mine, it is the Diátaxis framework — is that there are four distinct kinds and they have incompatible goals.</p>
<p><strong>Tutorials</strong> teach a beginner by having them do something that works. The goal is a successful experience, not completeness. A tutorial that mentions every option has failed. It should be prescriptive, opinionated, and it should work exactly as written, every time.</p>
<p><strong>How-to guides</strong> help someone accomplish a specific task they already understand. "How to configure TLS." Goal-oriented, assumes competence, skips explanation.</p>
<p><strong>Reference</strong> describes the machinery precisely and completely. Every parameter, every return value, every error. Boring by design. Generated where possible.</p>
<p><strong>Explanation</strong> provides understanding. Why is it designed this way? What are the trade-offs? What is the mental model? This is the kind that is almost always missing and the kind that most reduces support burden.</p>
<p>Most documentation fails because it tries to be all four at once. A tutorial that stops to explain architecture loses the beginner. A reference page with a narrative is hard to scan. Separate them, label them, and each one gets better.</p>
<h2 id="what-staff-it-like-a-product-means">what "staff it like a product" means<a class="anchor" href="#what-staff-it-like-a-product-means" aria-label="link to this section">#</a></h2>
<p><strong>One named owner.</strong> Not "the team." A person whose review includes it.</p>
<p><strong>A budget in the sprint.</strong> Documentation work sized and scheduled alongside features, not appended to the end of a ticket where it gets cut.</p>
<p><strong>Metrics.</strong> Support tickets that a doc would have prevented. Search queries with no results. Time-to-first-successful-request for a new user. Page-level feedback. Every one of these is measurable and almost nobody measures them.</p>
<p><strong>A definition of done that includes it.</strong> A feature is not shipped until it is documented. This is either enforced or it is a slogan; there is no middle.</p>
<h2 id="the-practices-that-actually-move-the-needle">the practices that actually move the needle<a class="anchor" href="#the-practices-that-actually-move-the-needle" aria-label="link to this section">#</a></h2>
<p><strong>Docs live with the code.</strong> Same repository, same pull request, same review. Documentation in a separate wiki drifts within one quarter, guaranteed, without exception.</p>
<p><strong>Test the examples.</strong> Every code sample in your docs should be extracted and run in CI. Broken examples are worse than no examples — they destroy trust in the whole document, and every set of docs has them, because they were correct when written.</p>
<p><strong>Write the failure cases.</strong> The single highest-value section in any documentation is "common errors and what they mean." This is what <a class="xref" href="/stack-overflows-traffic-fell-off-a-cliff-and-it-is-not-coming-back/" title="Stack Overflow&#x27;s traffic fell off a cliff and it is not coming back">Stack Overflow</a> existed to provide and it is the thing your docs almost certainly lack.</p>
<p>Go read your support queue. Every recurring question is a documentation gap with a measured frequency attached.</p>
<p><strong>Date and version everything.</strong> "This page describes v4.2, last updated 2026-01-15." Undated documentation is untrustworthy documentation, because the reader cannot tell whether it is current.</p>
<p><strong>Make the first example work.</strong> The single most common documentation failure: the quickstart does not run. Someone changed a default, renamed a parameter, required a new config field. Test the quickstart in CI, on a clean environment, on every release.</p>
<h2 id="the-argument-that-gets-budget">the argument that gets budget<a class="anchor" href="#the-argument-that-gets-budget" aria-label="link to this section">#</a></h2>
<p>Documentation is deflection. Every question answered by a doc is a question not asked of an engineer.</p>
<p>Count your support load. Estimate the fraction that is documentation-shaped — "how do I," "what does this error mean," "does it support." In most organizations it is more than half.</p>
<p>Now price that in engineer-hours. That is your documentation ROI, and it is usually large enough to fund a technical writer, which is the actual right answer and which almost nobody does.</p>
<h2 id="the-new-reason-it-matters">the new reason it matters<a class="anchor" href="#the-new-reason-it-matters" aria-label="link to this section">#</a></h2>
<p>Your documentation is now also a model's training data and a model's retrieval corpus.</p>
<p>When a developer asks an assistant about your library, the answer is synthesized from your docs. If your docs are wrong, incomplete, or ambiguous, the assistant confidently produces wrong code, and the user blames your library.</p>
<p>You have less control over how your project is explained than you did three years ago, and the only lever you have is the quality of the source material.</p>
<p>That is a strange new incentive and it is the strongest argument for good documentation that has ever existed.</p>]]></content:encoded></item><item><title>The postmortem that changes something</title><link>https://readme.news/the-postmortem-that-changes-something/</link><guid isPermaLink="true">https://readme.news/the-postmortem-that-changes-something/</guid><pubDate>Tue, 16 Dec 2025 09:00:00 +0000</pubDate><description>Most incident reviews produce a document and a ticket that never gets done. Here&#x27;s the difference.</description><content:encoded><![CDATA[<p>Most organizations run blameless postmortems. Most organizations also have the same class of incident repeatedly. Both of those things are true simultaneously and nobody finds it strange.</p>
<h2 id="the-failure-pattern">the failure pattern<a class="anchor" href="#the-failure-pattern" aria-label="link to this section">#</a></h2>
<p>The typical postmortem:</p>
<ol><li>Timeline of what happened. Accurate, detailed, useful.</li><li>Root cause. Usually a single technical fact.</li><li>Action items. Five to twelve of them.</li><li>Filed. Two action items get done. The rest age out.</li></ol>
<p>Six months later, a similar incident, with a similar document.</p>
<p>Three things are wrong here.</p>
<h2 id="problem-one-root-cause-is-singular">problem one: "root cause" is singular<a class="anchor" href="#problem-one-root-cause-is-singular" aria-label="link to this section">#</a></h2>
<p>Complex systems do not fail because of one thing. They fail because several conditions aligned, each of which was individually survivable.</p>
<p>The Cloudflare outage in November: a permissions change, a query that returned duplicates, a fixed buffer size, an error path that panicked instead of degrading, and a config pipeline that propagated globally in minutes. Remove any one and it does not happen, or it is much smaller.</p>
<p>Picking one and calling it "the root cause" means you fix one and leave four.</p>
<p>Better framing: <strong>contributing factors</strong>, plural, each with its own assessment of whether it is worth addressing. Some will not be — that is a legitimate decision if it is made explicitly.</p>
<h2 id="problem-two-action-items-without-owners-and-dates-are-wishes">problem two: action items without owners and dates are wishes<a class="anchor" href="#problem-two-action-items-without-owners-and-dates-are-wishes" aria-label="link to this section">#</a></h2>
<p>An action item that says "improve monitoring for the config pipeline" with no owner, no date, and no definition of done is a sentence, not a plan.</p>
<p>The fix is unglamorous:</p>
<ul><li><strong>Every action item has one named person.</strong> Not a team. A person.</li><li><strong>Every action item has a date.</strong> If nobody will commit to a date, it is not going to happen and you should delete it and say so.</li><li><strong>Every action item has a definition of done</strong> that someone else could verify.</li><li><strong>They go in the same backlog as feature work</strong>, prioritized against it. An action item in a separate "incident follow-up" list that is never sprint planned is a list of things that will not be done.</li></ul>
<p>And the one that actually forces it: <strong>review the open action items at the start of the next postmortem.</strong> Nothing motivates completion like a room full of people looking at your undone item from last quarter's incident while discussing this quarter's similar one.</p>
<h2 id="problem-three-nobody-asks-about-the-near-misses">problem three: nobody asks about the near misses<a class="anchor" href="#problem-three-nobody-asks-about-the-near-misses" aria-label="link to this section">#</a></h2>
<p>The incidents you review are the ones that broke through. For every one, there were several that did not — a bad deploy caught by a canary, a config error someone noticed in review, a query that would have taken down the database if it had run on Monday instead of Sunday.</p>
<p>Those contain the same information at a fraction of the cost, and almost nobody collects them.</p>
<p>Add a lightweight channel for it. "I nearly broke prod today, here is how." No document, no meeting, no blame. Just a note. The pattern that emerges over a quarter is more valuable than any individual postmortem.</p>
<h2 id="the-questions-that-produce-useful-findings">the questions that produce useful findings<a class="anchor" href="#the-questions-that-produce-useful-findings" aria-label="link to this section">#</a></h2>
<p>Replace "what was the root cause" with:</p>
<ul><li><strong>"What made this hard to detect?"</strong> Detection time is usually the largest component of impact and is the most improvable.</li><li><strong>"What made this hard to diagnose?"</strong> Usually missing observability. This produces the highest-value action items.</li><li><strong>"What made recovery slow?"</strong> Often a missing runbook, a missing <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">kill switch</a>, or a rollback that was not actually tested.</li><li><strong>"Who knew something that would have helped, and why did that not reach the responders?"</strong> This is an organizational question and it is frequently the real finding.</li><li><strong>"What did we do that helped?"</strong> Genuinely important. Practices that worked should be named so they get kept.</li></ul>
<h2 id="the-blameless-part-done-correctly">the blameless part, done correctly<a class="anchor" href="#the-blameless-part-done-correctly" aria-label="link to this section">#</a></h2>
<p>Blameless does not mean nobody made a mistake. It means the analysis focuses on why the mistake was possible and easy, rather than on the person.</p>
<p>"Alice deployed without running the tests" is blame and it is also useless. "The deploy path does not require tests to pass, and the shortcut that skips them is the fastest way to deploy" is the same fact stated in a way you can act on.</p>
<p>If the answer to "why did they do that" is "because the system made it easy and the correct path was hard," you have found something.</p>
<p>If the answer is genuinely "they were careless," you still fix the system, because the next person will also be careless eventually. Humans are the constant; the system is the variable.</p>]]></content:encoded></item><item><title>The technical interview is measuring the wrong thing, again</title><link>https://readme.news/the-technical-interview-is-measuring-the-wrong-thing-again/</link><guid isPermaLink="true">https://readme.news/the-technical-interview-is-measuring-the-wrong-thing-again/</guid><pubDate>Fri, 28 Nov 2025 09:00:00 +0000</pubDate><description>Every hiring process eventually optimizes for the wrong signal. The current one has a new failure mode.</description><content:encoded><![CDATA[<p>The technical interview has been broken in a rotating set of ways for twenty years. Each fix creates the next problem.</p>
<ul><li><strong>Brainteasers</strong> measured whether you had heard the brainteaser. Replaced by algorithms.</li><li><strong>Algorithm puzzles</strong> measured whether you had ground a puzzle site. Replaced, partially, by take-homes.</li><li><strong>Take-homes</strong> measured how much unpaid time you had, which selects against people with children and second jobs. Replaced, partially, by pairing.</li><li><strong>Pairing</strong> is the best of them and it measures your comfort being watched, which is correlated with experience and confidence, which is correlated with demographics.</li></ul>
<p>Now there is a new one, and it is genuinely novel.</p>
<h2 id="the-current-problem">the current problem<a class="anchor" href="#the-current-problem" aria-label="link to this section">#</a></h2>
<p>A candidate with a model available can pass most remote technical screens. Not "can cheat" — <em>can pass</em>, because the tasks we set are exactly the tasks these tools are good at.</p>
<p>The industry response has been mostly bad. Return-to-office for interviews. Proctoring software. Increasingly hostile monitoring. All of it makes the experience worse for honest candidates and is defeated by anyone determined.</p>
<p>The response has been bad because the question was framed wrong. The question is not "how do we stop candidates using AI." It is "what are we actually trying to measure, given that the job now includes using these tools."</p>
<h2 id="what-the-job-actually-is">what the job actually is<a class="anchor" href="#what-the-job-actually-is" aria-label="link to this section">#</a></h2>
<p>If you hired someone today, their work would involve:</p>
<ul><li>Understanding an existing system well enough to change it safely.</li><li>Deciding what to build, which is mostly deciding what not to build.</li><li>Using AI tools effectively, including knowing when the output is wrong.</li><li>Communicating a technical decision to people who will be affected by it.</li><li>Debugging something under time pressure with incomplete information.</li><li>Reviewing someone else's code — increasingly, a machine's — and catching the problem.</li></ul>
<p>Not one of those is measured by "implement an LRU cache in forty-five minutes."</p>
<h2 id="interviews-that-measure-the-real-thing">interviews that measure the real thing<a class="anchor" href="#interviews-that-measure-the-real-thing" aria-label="link to this section">#</a></h2>
<p><strong>Code review.</strong> Give them a 300-line pull request with three deliberate problems: one obvious bug, one subtle design issue, one thing that is fine but looks wrong. Ask them to review it.</p>
<p>This is excellent. It is exactly the job, AI does not obviously help, and the conversation about the third item — where they explain why the suspicious thing is actually correct — tells you more about their judgment than any implementation task.</p>
<p><strong>Debugging a real system.</strong> Give them a repository, a failing test, and thirty minutes. Let them use whatever tools they want, including models. Watch how they narrow it down. Do they read the error? Do they form a hypothesis? Do they check it? Do they notice when the model's suggestion is wrong?</p>
<p>Watching someone debug with AI assistance is a much better signal than watching them code without it, because it is the actual work.</p>
<p><strong>Design discussion on their own past work.</strong> "Tell me about a system you built. What would you change?" The follow-up questions are where the signal is. People who genuinely understood their system can answer six levels deep. People who did not, cannot, and it becomes clear quickly.</p>
<p><strong>A short, paid, scoped project.</strong> Four hours, paid at a real rate, on something close to the actual work. This is the highest-signal option and the least scalable, and it is worth it for senior roles.</p>
<h2 id="what-to-stop-doing">what to stop doing<a class="anchor" href="#what-to-stop-doing" aria-label="link to this section">#</a></h2>
<p><strong>Stop asking people to implement data structures from memory.</strong> They will not do this in the job, and if they need one they will look it up, correctly.</p>
<p><strong>Stop the six-round loop.</strong> Every round is a coin flip with a false-negative rate. Six rounds does not make the signal six times better; it makes the process long enough that good candidates take another offer.</p>
<p><strong>Stop pretending the whiteboard measures anything but whiteboard performance.</strong></p>
<h2 id="the-thing-nobody-wants-to-hear">the thing nobody wants to hear<a class="anchor" href="#the-thing-nobody-wants-to-hear" aria-label="link to this section">#</a></h2>
<p>Interviews have a low ceiling on signal. The correlation between interview performance and job performance is weak in every study anyone has run.</p>
<p>The highest-signal thing is working with someone. Everything else is a proxy. Which argues for: shorter processes, more willingness to take a chance, and robust ways to correct the mistake — a real probation practice, honest early feedback, and the organizational nerve to act on it.</p>
<p>That is a harder cultural change than redesigning the interview loop, which is why everyone redesigns the interview loop instead.</p>]]></content:encoded></item><item><title>Estimation is a communication problem</title><link>https://readme.news/estimation-is-a-communication-problem/</link><guid isPermaLink="true">https://readme.news/estimation-is-a-communication-problem/</guid><pubDate>Tue, 04 Nov 2025 09:00:00 +0000</pubDate><description>You are not bad at predicting the future. You are bad at explaining uncertainty to people who want a number.</description><content:encoded><![CDATA[<p>Every few years the industry rediscovers that software estimates are bad and proposes a new methodology. Story points. T-shirt sizes. Ideal days. No estimates. Reference class forecasting. Each one works for a while, mostly through the Hawthorne effect, and then stops.</p>
<p>The methodologies keep failing because they are solving the wrong problem.</p>
<h2 id="what-is-actually-happening">what is actually happening<a class="anchor" href="#what-is-actually-happening" aria-label="link to this section">#</a></h2>
<p>When a manager asks "how long will this take," they are usually not asking for a prediction. They are asking one of these:</p>
<ul><li><em>Can I promise this to a customer for Q1?</em></li><li><em>Should we do this or the other thing?</em></li><li><em>Do I need to hire?</em></li><li><em>Is this a two-week thing or a two-quarter thing, because those go in different plans?</em></li></ul>
<p>Every one of those is answerable. None of them requires a precise duration. And a precise duration, given confidently, actively harms all four — because it gets recorded, propagated, and treated as a commitment by people three steps removed who never saw the assumptions.</p>
<p>The engineer knows the estimate is uncertain. The number arrives downstream stripped of that uncertainty. That is the failure, and it is a communication failure, not a forecasting one.</p>
<h2 id="what-to-say-instead">what to say instead<a class="anchor" href="#what-to-say-instead" aria-label="link to this section">#</a></h2>
<p><strong>Give a range with an explicit confidence.</strong></p>
<blockquote><p>"Two to six weeks. I'd take even money on three. The spread is because I don't know yet whether the legacy import path can be reused — if it can, it's two weeks; if I have to rewrite it, it's five or six."</p></blockquote>
<p>That sentence contains everything the asker needs: a planning number, a worst case, and the specific unknown that determines which. It is also honest, which the single number was not.</p>
<p><strong>Name the decisive uncertainty.</strong> Almost every estimate has one thing that dominates the variance. Say what it is. That converts "the engineer is being vague" into "there is a specific question we could go answer," which is actionable.</p>
<p><strong>Offer to reduce the uncertainty.</strong></p>
<blockquote><p>"Give me two days to spike the import path and I'll come back with a much tighter number."</p></blockquote>
<p>Two days to turn a 3× range into a 1.3× range is almost always worth it, and proposing it changes the conversation from negotiation to problem-solving.</p>
<p><strong>Separate the estimate from the commitment.</strong> "It will probably take three weeks" and "I commit to delivering it in three weeks" are different statements, and conflating them is where trust gets destroyed. If someone wants a commitment, they are asking for buffer, and you should say what buffer you need.</p>
<h2 id="why-estimates-are-actually-wrong">why estimates are actually wrong<a class="anchor" href="#why-estimates-are-actually-wrong" aria-label="link to this section">#</a></h2>
<p>Not because engineers are optimistic — though they are. The dominant causes, in my experience, in order:</p>
<ol><li><strong>Unknown unknowns in existing code.</strong> The thing you have to modify does something you did not expect. This is the biggest one by a wide margin and it scales with the age and size of the codebase.</li><li><strong>Work that is not the work.</strong> Code review latency, CI flakes, environment problems, a dependency upgrade you did not plan, the meeting.</li><li><strong>Scope discovered during implementation.</strong> "Also it has to work for the enterprise tier" arrives in week two.</li><li><strong>Interruption.</strong> An estimate of three focused days is a lie if the person has two days of meetings and an <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> rotation.</li></ol>
<p>Note that only the first is about the difficulty of the task. Most estimation error is organizational, which is why individual estimation skill improves results much less than people expect.</p>
<h2 id="the-thing-that-actually-works">the thing that actually works<a class="anchor" href="#the-thing-that-actually-works" aria-label="link to this section">#</a></h2>
<p>Track your own history. Not for a methodology — for calibration.</p>
<p>Write down the estimate. Write down the actual. Do it for three months. You will discover you have a personal multiplier, it is remarkably stable, and it is probably between 1.5 and 3.</p>
<p>Then apply it. Silently. That single practice does more than any team-level process, and it requires no one else's participation.</p>
<h2 id="the-part-managers-should-read">the part managers should read<a class="anchor" href="#the-part-managers-should-read" aria-label="link to this section">#</a></h2>
<p>If you punish an engineer for a missed estimate, you will get padded estimates forever, and you will have destroyed the information content of the number.</p>
<p>If you ask "what would make you more confident," you will get real answers and a team that tells you when something is going wrong early, which is the only thing that actually helps.</p>
<p>You cannot have both. Pick.</p>]]></content:encoded></item><item><title>Shopify's memo and the new hiring question</title><link>https://readme.news/shopifys-memo-and-the-new-hiring-question/</link><guid isPermaLink="true">https://readme.news/shopifys-memo-and-the-new-hiring-question/</guid><pubDate>Fri, 25 Apr 2025 09:00:00 +0000</pubDate><description>&quot;Prove AI can&#x27;t do this job&quot; as a requirement before headcount. A reasonable policy with an unreasonable failure mode.</description><content:encoded><![CDATA[<p>Tobi Lütke published an internal memo this month stating that Shopify teams must demonstrate why AI cannot do a job before requesting headcount, and that "reflexive AI usage" is now a baseline expectation in performance reviews.</p>
<p>The memo is better written and more reasonable than the headlines about it. It is also going to be misapplied everywhere it gets copied, and it is going to get copied everywhere.</p>
<h2 id="the-defensible-core">the defensible core<a class="anchor" href="#the-defensible-core" aria-label="link to this section">#</a></h2>
<p>Two claims in the memo are correct.</p>
<p><strong>"Try the AI first" is good engineering hygiene.</strong> Before you build the internal tool, before you write the script, before you file the ticket — spend ten minutes seeing whether an existing model handles it. Often it does. The cost of checking is near zero and the expected value is high.</p>
<p><strong>Stagnation is a choice.</strong> A team that has not changed how it works in two years during a period of rapid tooling change is not being careful, it is being incurious. That is a fair thing to name.</p>
<h2 id="the-failure-mode">the failure mode<a class="anchor" href="#the-failure-mode" aria-label="link to this section">#</a></h2>
<p>"Prove AI cannot do it" is an unfalsifiable standard, and unfalsifiable standards in an organization become political tools.</p>
<p>You cannot prove a negative about a capability that changes monthly. Any manager who wants to deny headcount now has an infinitely flexible reason, and any manager who wants to grant it will produce a document explaining why AI cannot do it. The document is not evidence. It is theater. Every organization that has ever required a justification memo for headcount has produced a genre of justification-memo prose, and this is just the newest style.</p>
<p>The real question — "what is the highest-leverage use of an additional engineer" — was always the question. Adding an AI framing does not make it easier to answer, it makes it easier to obscure.</p>
<h2 id="the-second-order-effect-nobody-plans-for">the second-order effect nobody plans for<a class="anchor" href="#the-second-order-effect-nobody-plans-for" aria-label="link to this section">#</a></h2>
<p>If teams are evaluated on AI usage, teams will use AI, including in places where it is worse. This is Goodhart's law with a new coat of paint. You will get:</p>
<ul><li><a class="xref" href="/code-review-comments-that-change-things/" title="Code review comments that change things">Code review comments</a> generated by a model that read fine and check nothing.</li><li>Documentation nobody reads, generated because generating it is cheap.</li><li>Test suites with impressive coverage numbers and no assertions that would fail.</li><li>Postmortems written by a model that faithfully summarize the incident and identify no real cause.</li></ul>
<p>All of that is measurable AI adoption. None of it is value.</p>
<h2 id="what-a-better-version-looks-like">what a better version looks like<a class="anchor" href="#what-a-better-version-looks-like" aria-label="link to this section">#</a></h2>
<p>If you want the outcome the memo is aiming at, ask for the outcome directly:</p>
<ul><li><strong>Cycle time</strong>, not tool adoption. If the team ships faster with the same quality, they are using their tools well, and it does not matter which ones.</li><li><strong>Toil reduction as a named goal</strong>, with a quarterly review of what got automated. That surfaces the same opportunities without the unfalsifiable test.</li><li><strong>A budget and permission, not a mandate.</strong> The blocker in most organizations is not enthusiasm, it is that the good tools are not approved and the data policy is unclear. Fix that and adoption happens without a memo.</li></ul>
<h2 id="the-part-i-actually-agree-with">the part I actually agree with<a class="anchor" href="#the-part-i-actually-agree-with" aria-label="link to this section">#</a></h2>
<p>The memo says learning to use these tools well is now part of the job. That is true and it is not controversial in the way people are treating it.</p>
<p>Getting good output from a model is a skill — knowing what context to provide, how to decompose a task, when the answer is wrong in a way that looks right. It is closer to being a good technical lead than to typing a search query. People who have developed it are meaningfully more effective, and people who dismissed the whole category in 2023 and never revisited it are falling behind in a way that is going to be uncomfortable to talk about at review time.</p>
<p>That is worth saying out loud. It just does not require a headcount policy.</p>]]></content:encoded></item>
</channel>
</rss>
