<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>README — testing</title>
<link>https://readme.news/tags/testing/</link>
<atom:link href="https://readme.news/tags/testing/feed.xml" rel="self" type="application/rss+xml"/>
<description>README pieces tagged testing.</description>
<language>en-us</language>
<lastBuildDate>Thu, 01 Oct 2026 13:20:31 +0000</lastBuildDate>
<item><title>Why your tests are slow</title><link>https://readme.news/why-your-tests-are-slow/</link><guid isPermaLink="true">https://readme.news/why-your-tests-are-slow/</guid><pubDate>Wed, 16 Sep 2026 09:00:00 +0000</pubDate><description>Six causes, ranked by how often they are the real problem, and the fix for each.</description><content:encoded><![CDATA[<p>A test suite that takes twenty minutes is not run locally. A suite that is not run locally is a suite that fails in CI, which means the feedback loop is now measured in pipeline runs.</p>
<p>Here is where the time actually goes, roughly in order of how often each is the dominant cause.</p>
<h2 id="1-everything-talks-to-a-real-database">1. everything talks to a real database<a class="anchor" href="#1-everything-talks-to-a-real-database" aria-label="link to this section">#</a></h2>
<p>The most common cause by a wide margin. Each test sets up a database, inserts fixtures, runs, and tears down.</p>
<p>Fixes, in order of return:</p>
<p><strong>Roll back instead of truncating.</strong> Wrap each test in a transaction and roll it back. Orders of magnitude faster than deleting rows, and it needs no cleanup code.</p>
<p><strong>Share the schema, not the data.</strong> Create the schema once per run, not per test.</p>
<p><strong>Use a tmpfs for the test database.</strong> The data does not need to survive; durable writes are pure cost. On Postgres, a data directory in memory plus <code>fsync=off</code> and <code>synchronous_commit=off</code> is dramatically faster and completely inappropriate for anything but tests.</p>
<p><strong>Move the logic out of the database's reach.</strong> The deepest fix: if business logic is a pure function of its inputs, its tests do not need a database at all. Tests that are hard to write without infrastructure are telling you about your design.</p>
<h2 id="2-the-suite-is-serial">2. the suite is serial<a class="anchor" href="#2-the-suite-is-serial" aria-label="link to this section">#</a></h2>
<p>Most runners parallelise and most projects have not turned it on, usually because it exposed shared state once and someone reverted it.</p>
<p>That shared state is a bug. Tests that pass only in a particular order have an ordering dependency, and ordering dependencies in tests usually mirror a real one in the code.</p>
<p>Turn on parallelism, fix what breaks, and you will find at least one genuine issue.</p>
<h2 id="3-sleeps">3. sleeps<a class="anchor" href="#3-sleeps" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">time.sleep(2)   # wait for the worker to pick it up</code></pre></div>
<p>Every one of these is pure latency, and they are always tuned to the slowest machine anyone has run on. Fifty of them is a hundred seconds of doing nothing.</p>
<p>Replace with polling on the actual condition:</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">wait_until(lambda: job.reload().status == "done", timeout=5)</code></pre></div>
<p>Same worst case, typically a hundredth of the elapsed time, and it fails with a useful message instead of a mysterious assertion.</p>
<h2 id="4-the-fixtures-are-enormous">4. the fixtures are enormous<a class="anchor" href="#4-the-fixtures-are-enormous" aria-label="link to this section">#</a></h2>
<p>A test that needs one user builds an organisation, three teams, forty users and a year of history, because it reuses the "standard" fixture.</p>
<p>Build the minimum. Factories with sensible defaults, overridden per test, beat shared fixture files that grow to serve every case.</p>
<h2 id="5-it-is-doing-real-network-io">5. it is doing real network I/O<a class="anchor" href="#5-it-is-doing-real-network-io" aria-label="link to this section">#</a></h2>
<p>Tests hitting real HTTP endpoints — even internal ones, even mock servers over a socket — pay connection setup and scheduling per call.</p>
<p>Intercept at the client layer rather than the network layer. Most languages have a way to stub the HTTP client in-process, and it is both faster and more deterministic than a local server.</p>
<p>The exception is contract tests, which exist to catch exactly what stubs hide. Keep those, run them separately, and do not let them into the fast suite.</p>
<h2 id="6-everything-is-an-end-to-end-test">6. everything is an end-to-end test<a class="anchor" href="#6-everything-is-an-end-to-end-test" aria-label="link to this section">#</a></h2>
<p>Browser-driving tests are two to three orders of magnitude slower than unit tests. A suite made mostly of them is slow no matter what you do.</p>
<p>The ratio that works: a large number of fast unit tests, a moderate number of integration tests around real boundaries, and a small number of end-to-end tests covering the handful of journeys that must never break.</p>
<p>Inverting that ratio is the most expensive testing mistake a team can make, and it is usually the result of not trusting the lower layers rather than a deliberate choice.</p>
<h2 id="the-measurement-first">the measurement first<a class="anchor" href="#the-measurement-first" aria-label="link to this section">#</a></h2>
<p>Before fixing anything, get the distribution:</p>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">pytest --durations=25
go test ./... -json | jq -r 'select(.Action=="pass") | "\(.Elapsed) \(.Test)"' | sort -rn | head -25</code></pre></div>
<p>Almost always, a small number of tests dominate. Fixing the slowest twenty is usually the whole job, and it is an afternoon rather than a project.</p>
<h2 id="the-target">the target<a class="anchor" href="#the-target" aria-label="link to this section">#</a></h2>
<p><strong>Under 10 seconds for the unit suite</strong>, run on every save. Under two minutes for everything that gates a merge.</p>
<p>The ten-second number is not arbitrary — it is roughly the threshold past which people stop running tests as they work and start batching. Once that happens you have lost the feedback loop, and the suite's speed stops being a convenience question and starts being a quality one.</p>]]></content:encoded></item><item><title>Property-based testing deserves its moment</title><link>https://readme.news/property-based-testing-deserves-its-moment/</link><guid isPermaLink="true">https://readme.news/property-based-testing-deserves-its-moment/</guid><pubDate>Mon, 20 Jul 2026 09:00:00 +0000</pubDate><description>State an invariant, let the machine find the counterexample. A twenty-year-old technique that fits this moment exactly.</description><content:encoded><![CDATA[<p>Example-based tests check the cases you thought of. Property-based tests check the cases you did not.</p>
<p>The technique has been around since QuickCheck in 1999 and has stayed niche. It fits the current moment unusually well, and it is worth another look.</p>
<h2 id="the-idea">the idea<a class="anchor" href="#the-idea" aria-label="link to this section">#</a></h2>
<p>Instead of "for this input, expect this output," you state a property that should hold for <em>all</em> inputs, and the framework generates hundreds of inputs trying to break it.</p>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">from hypothesis import given, strategies as st

@given(st.lists(st.integers()))
def test_sort_is_idempotent(xs):
    assert sorted(sorted(xs)) == sorted(xs)

@given(st.lists(st.integers()))
def test_sort_preserves_elements(xs):
    assert sorted(xs).count == xs.count   # same multiset

@given(st.text())
def test_roundtrip(s):
    assert decode(encode(s)) == s</code></pre></div>
<p>The framework generates empty lists, single elements, lists with duplicates, huge values, negative numbers, and — critically — when it finds a failure, it <em>shrinks</em> the input to the minimal case that still fails.</p>
<p>That shrinking is what makes the technique practical. A failure on a 400-element list is not actionable. The same failure shrunk to <code>[0, 0]</code> tells you exactly what is wrong.</p>
<h2 id="the-properties-that-are-actually-useful">the properties that are actually useful<a class="anchor" href="#the-properties-that-are-actually-useful" aria-label="link to this section">#</a></h2>
<p>The hard part is not the tooling. It is identifying properties, and there is a small catalogue that covers most cases.</p>
<p><strong>Round trip.</strong> <code>decode(encode(x)) == x</code>. Serialization, <a class="xref" href="/compression-is-underrated/" title="Compression is underrated">compression</a>, parsing, encryption. This one property catches an enormous number of real bugs and applies almost everywhere.</p>
<p><strong>Invariants.</strong> Something that is always true regardless of operations. A balanced tree stays balanced. A sorted list stays sorted. Account balances sum to the same total across a transfer. A cache never returns stale data past its TTL.</p>
<p><strong>Oracle.</strong> Compare against a simpler, slower, obviously-correct implementation. You have an optimized version; write the naive one, and assert they agree. This is extremely powerful and underused — it is how you test the fast path against the version you can reason about.</p>
<p><strong>Idempotence.</strong> <code>f(f(x)) == f(x)</code>. Normalization, deduplication, and — importantly — any API operation that claims to be idempotent. If you have <a class="xref" href="/idempotency-is-the-only-distributed-systems-concept-you-need/" title="Idempotency is the only distributed systems concept you need">idempotency</a> keys, this is how you actually verify them.</p>
<p><strong>Commutativity and associativity.</strong> Order should not matter. Merge operations, set operations, CRDT merges.</p>
<p><strong>Metamorphic relations.</strong> When you cannot state the correct output, state how the output should <em>change</em>. Adding an item to a cart should increase the total by that item's price. Searching for a more specific query should return a subset. Sorting descending should be the reverse of sorting ascending.</p>
<p>This last category is the one that unlocks property testing for business logic, where there is no obvious oracle.</p>
<h2 id="where-it-fits-best">where it fits best<a class="anchor" href="#where-it-fits-best" aria-label="link to this section">#</a></h2>
<p><strong>Parsers and serializers.</strong> Round trip, always.</p>
<p><strong>Data structures.</strong> Invariants after every operation sequence.</p>
<p><strong>State machines.</strong> Generate random valid operation sequences, assert the invariants hold throughout. This finds ordering bugs that no hand-written test would.</p>
<p><strong>Anything with an obvious naive implementation.</strong> Oracle testing.</p>
<p><strong>Financial and unit arithmetic.</strong> Rounding, currency, conversions. The edge cases are numerous and boring, which is exactly what a generator is for.</p>
<p><strong>Concurrent code</strong>, with a framework that generates interleavings. This is the hardest category to test any other way.</p>
<h2 id="where-it-does-not-fit">where it does not fit<a class="anchor" href="#where-it-does-not-fit" aria-label="link to this section">#</a></h2>
<p><strong>UI.</strong> The properties are aesthetic.</p>
<p><strong>Anything where the correct output requires human judgment.</strong></p>
<p><strong>Integration tests against real systems.</strong> Generation implies many runs, and many runs against a real database is slow.</p>
<p><strong>Code where you cannot state a property.</strong> If you genuinely cannot, that may itself be information about the design — code with no statable invariants is code with no contract.</p>
<h2 id="the-practical-advice">the practical advice<a class="anchor" href="#the-practical-advice" aria-label="link to this section">#</a></h2>
<p><strong>Do not replace example tests.</strong> Keep them. They document intent and they are the fastest way to communicate what a function is for. Add properties alongside.</p>
<p><strong>Start with round-trip properties.</strong> Easiest to state, highest hit rate for real bugs.</p>
<p><strong>Save the failing seed.</strong> When a property fails, the framework gives you the counterexample. Add it as a regression example test so it is checked deterministically forever.</p>
<p><strong>Bound the input space.</strong> Unbounded generation produces absurd inputs and slow tests. Constrain to realistic ranges: strings up to a reasonable length, numbers in a plausible domain.</p>
<p><strong>Run more cases in CI than locally.</strong> A hundred examples locally for fast feedback, a thousand in CI, ten thousand in a nightly run.</p>
<h2 id="why-now">why now<a class="anchor" href="#why-now" aria-label="link to this section">#</a></h2>
<p>Two reasons this fits the current moment.</p>
<p><strong>Verification is the bottleneck.</strong> Property testing is verification you write once and that checks a large input space forever. That is exactly the leverage the moment calls for.</p>
<p><strong>Stating a property is a task worth doing carefully by hand</strong>, and generating the implementation is not. Writing "the total must always equal the sum of line items" requires understanding the domain; nothing else in the test does.</p>
<p>That is a good division of labor, and it points at where human effort should concentrate: on specifying what must be true, and letting everything else be derived.</p>]]></content:encoded></item><item><title>Static analysis is finally worth the false positives</title><link>https://readme.news/static-analysis-is-finally-worth-the-false-positives/</link><guid isPermaLink="true">https://readme.news/static-analysis-is-finally-worth-the-false-positives/</guid><pubDate>Fri, 05 Jun 2026 09:00:00 +0000</pubDate><description>The tools got dramatically better while everyone was ignoring them because of a bad experience in 2015.</description><content:encoded><![CDATA[<p>A lot of engineers formed their opinion of static analysis from a tool that produced four thousand warnings on first run, 95% of which were noise, and got disabled within a month.</p>
<p>That was an accurate assessment of the tools at the time. The tools are substantially different now and the assessment has not updated.</p>
<h2 id="what-changed">what changed<a class="anchor" href="#what-changed" aria-label="link to this section">#</a></h2>
<p><strong>Flow-sensitive analysis became standard.</strong> Older linters matched patterns in the syntax tree. Modern analyzers track values through the control flow graph, which means they can tell that a variable was checked for null on line 12 and therefore is not null on line 40. That single capability eliminates the largest source of false positives.</p>
<p><strong>Language servers made it interactive.</strong> A warning in your editor while you type is a different product from a report generated in CI. You fix it in context, in seconds, instead of triaging a list a week later.</p>
<p><strong><a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">Type systems</a> absorbed much of it.</strong> A lot of what static analysis used to catch is now caught by the compiler in a typed language, for free, with no false positives at all.</p>
<p><strong>The defaults got sane.</strong> Modern tools ship with a curated recommended set rather than everything enabled. <code>clippy</code>, <code>ruff</code>, <code>biome</code>, <code>staticcheck</code> and their peers are opinionated about what is worth reporting.</p>
<p><strong>They got fast.</strong> Analyzers written in compiled languages run over a large codebase in seconds. Speed matters more than people credit — a check that takes two minutes gets run in CI, and a check that takes two seconds gets run on every save, which is where it actually changes behavior.</p>
<h2 id="what-to-actually-run">what to actually run<a class="anchor" href="#what-to-actually-run" aria-label="link to this section">#</a></h2>
<p><strong>A fast linter with a good default set</strong>, on save, in the editor. <code>ruff</code> for Python, <code>clippy</code> for Rust, <code>biome</code> or <code>eslint</code> for JavaScript, <code>staticcheck</code> for Go.</p>
<p><strong>A typechecker in strict mode</strong>, in CI. This is the highest-value item on the list for a gradually-typed language, and the non-strict modes permit exactly the holes that make the guarantees unreliable.</p>
<p><strong>A security-focused analyzer</strong> if you handle untrusted input. Taint tracking — does data from a request reach a SQL query, a shell command, or a template without sanitization — is the specific capability worth having, and it is a genuinely different analysis from ordinary linting.</p>
<p><strong>A dependency scanner</strong>, ranked by reachability if your tooling supports it. A critical CVE in code you never call is lower priority than a medium in your request path, and a scanner that cannot tell you which is which produces a queue nobody reads.</p>
<h2 id="the-adoption-sequence">the adoption sequence<a class="anchor" href="#the-adoption-sequence" aria-label="link to this section">#</a></h2>
<p>Turning on a full rule set against an existing codebase produces thousands of warnings and gets the tool disabled. The sequence that works:</p>
<p><strong>1. Run it in report-only mode.</strong> Get the number. Do not fix anything yet.</p>
<p><strong>2. Enable a small subset that has near-zero false positives</strong> and fix those. Usually: unused variables, unreachable code, obviously wrong comparisons, missing awaits. Twenty rules, not four hundred.</p>
<p><strong>3. Make it blocking for new and changed code only.</strong> Most tools support this, either natively or through a diff-aware wrapper. This is the key move — the existing violations do not block anyone, and the codebase stops getting worse immediately.</p>
<p><strong>4. Burn down the backlog opportunistically.</strong> When you touch a file, fix its warnings. No cleanup sprint, no dedicated project.</p>
<p><strong>5. Add rules gradually</strong>, one at a time, each with a burn-down.</p>
<p>Steps three and four are where most adoptions succeed or fail. A tool that blocks the whole team on a pre-existing backlog gets turned off; one that only blocks new violations is uncontroversial.</p>
<h2 id="the-rules-worth-arguing-about">the rules worth arguing about<a class="anchor" href="#the-rules-worth-arguing-about" aria-label="link to this section">#</a></h2>
<p>Some checks are genuinely contested and you should decide deliberately rather than accepting the default:</p>
<p><strong>Cyclomatic complexity limits.</strong> Sometimes a function is legitimately complex because the domain is. A hard limit produces artificially split functions that are harder to read, not easier.</p>
<p><strong>Line length.</strong> Real disagreement, formatter should handle it, not worth a rule.</p>
<p><strong>Naming conventions.</strong> Worth enforcing, and pick your convention rather than the tool's default if they differ.</p>
<p><strong>Anything with more than a few percent false positives.</strong> A rule that is wrong one time in ten trains people to ignore it, and that habit generalizes to the rules that are right.</p>
<h2 id="the-honest-limits">the honest limits<a class="anchor" href="#the-honest-limits" aria-label="link to this section">#</a></h2>
<p>Static analysis finds a specific class of bug: local, syntactic, pattern-matchable. It does not find logic errors, wrong business rules, race conditions in most cases, or performance problems.</p>
<p>It is not a substitute for tests, review, or thought. It is a way to spend zero human attention on the errors that do not require human attention, which frees attention for the ones that do.</p>
<p>That framing — attention allocation rather than bug finding — is the one that makes it worth the setup.</p>
<h2 id="the-new-reason-it-matters">the new reason it matters<a class="anchor" href="#the-new-reason-it-matters" aria-label="link to this section">#</a></h2>
<p>Machine-generated code has a characteristic error profile: plausible, syntactically valid, and wrong in specific recurring ways. Unchecked errors, missing awaits, resource leaks, off-by-one in boundary conditions.</p>
<p>Those are exactly the errors static analysis is good at. Running a strict analyzer over generated code is verification you get for free, and it is one of the few places where the verification bottleneck has an automated answer.</p>
<p>Turn it on.</p>]]></content:encoded></item><item><title>Model evaluation for people who ship</title><link>https://readme.news/model-evaluation-for-people-who-ship/</link><guid isPermaLink="true">https://readme.news/model-evaluation-for-people-who-ship/</guid><pubDate>Wed, 27 May 2026 09:00:00 +0000</pubDate><description>Not research benchmarks. A practical harness you can build in a day that makes every future model decision an hour instead of a week.</description><content:encoded><![CDATA[<p>Every few weeks a new model ships and someone asks whether you should switch. Without an evaluation harness, answering that takes a week of impressions and you will get it wrong. With one, it takes an hour.</p>
<p>This is the highest-return day of engineering available to anyone building on models, and a surprising number of teams have not done it.</p>
<h2 id="what-it-is-not">what it is not<a class="anchor" href="#what-it-is-not" aria-label="link to this section">#</a></h2>
<p>Not MMLU. Not a leaderboard. Not "vibes after twenty prompts."</p>
<p>Public benchmarks tell you which models are worth testing. They do not predict performance on your task, because your task is not in them and because contamination is universal.</p>
<h2 id="what-it-is">what it is<a class="anchor" href="#what-it-is" aria-label="link to this section">#</a></h2>
<p>Fifty to two hundred examples from your actual production traffic, with expected outputs or a grading rubric, run automatically, producing a number.</p>
<div class="code"><pre><code>evals/
  cases/
    001-refund-request.json
    002-ambiguous-address.json
    ...
  run.py
  results/
    2026-05-27-model-a.json</code></pre></div>
<p>Each case:</p>
<div class="code"><span class="code-lang">json</span><pre><code class="lang-json">{
  "id": "042",
  "input": { "ticket": "my order never arrived and I want my money back" },
  "expect": { "category": "refund", "urgency": "high", "needs_human": false },
  "notes": "the word 'never' should not trigger the fraud path"
}</code></pre></div>
<p>That is it. The whole thing is a test suite where the assertions are fuzzier.</p>
<h2 id="where-the-cases-come-from">where the cases come from<a class="anchor" href="#where-the-cases-come-from" aria-label="link to this section">#</a></h2>
<p><strong>Your production failures.</strong> This is the best source by a wide margin. Every time the system gets something wrong, that becomes a case. Your eval set grows into a precise map of your problem's difficulty.</p>
<p>Set this up as a workflow: a thumbs-down in the product, or a support escalation, creates a candidate case that someone reviews and adds.</p>
<p><strong>Your edge cases.</strong> The weird inputs. The empty ones. The ones in another language. The adversarial ones. The ones with an injection attempt.</p>
<p><strong>A stratified sample of normal traffic.</strong> So you notice when a change breaks the common case while fixing an edge case.</p>
<p><strong>Cases with no correct answer.</strong> Where the right behavior is to refuse, escalate, or ask a clarifying question. Models are frequently bad at this and it is rarely tested.</p>
<h2 id="grading">grading<a class="anchor" href="#grading" aria-label="link to this section">#</a></h2>
<p>Three approaches, and you will use all three.</p>
<p><strong>Exact or structural match.</strong> For classification, extraction, and structured output. Cheap, deterministic, unambiguous. Use it wherever you can.</p>
<p><strong>Programmatic checks.</strong> For generated code: does it compile, do the tests pass. For SQL: does it run, does it return the right shape. This is the strongest form of grading and it is available more often than people realize — if you can verify mechanically, do.</p>
<p><strong>Model-as-judge.</strong> For open-ended output. A second model grades against a rubric.</p>
<p>Use it carefully:</p>
<ul><li><strong>Write a specific rubric</strong>, not "is this good." Score each dimension separately.</li><li><strong>Validate the judge against human ratings</strong> on a sample. If the judge disagrees with you, the judge is wrong and the rubric needs work.</li><li><strong>Use a different model than the one being evaluated</strong>, or at minimum be aware of self-preference bias, which is well documented and large.</li></ul>
<h2 id="the-metrics">the metrics<a class="anchor" href="#the-metrics" aria-label="link to this section">#</a></h2>
<p><strong>Accuracy on your set</strong>, obviously.</p>
<p><strong>Cost per case.</strong> Tokens in and out, at current prices. Track this — a model that is 2% better and 4× the cost is usually the wrong choice.</p>
<p><strong>Latency, at percentiles.</strong> p50 and p95. Reasoning models have high variance and the average hides it.</p>
<p><strong>Failure mode distribution.</strong> Not just how many wrong — <em>how</em> wrong. A model that fails by refusing is very different from one that fails by confidently fabricating, and the aggregate score treats them identically.</p>
<h2 id="the-workflow">the workflow<a class="anchor" href="#the-workflow" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">bash</span><pre><code class="lang-bash">python evals/run.py --model model-a --model model-b --parallel 8</code></pre></div>
<p>Output a table. Commit the results. Diff them across runs.</p>
<p>Run it:</p>
<ul><li><strong>When any model ships.</strong> Within an hour of the announcement, you know.</li><li><strong>When you change a prompt.</strong> Prompt changes are code changes with no <a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">type system</a> and no compiler; the eval is your only regression check.</li><li><strong>On a schedule.</strong> Providers update models behind stable names. Behavior drifts. You want to know from your <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a>, not from your support queue.</li></ul>
<p>That last one catches something most teams never notice: the model you deployed against is not the model serving your traffic today.</p>
<h2 id="the-one-that-matters-most">the one that matters most<a class="anchor" href="#the-one-that-matters-most" aria-label="link to this section">#</a></h2>
<p><strong>Version your prompts and store the eval result with them.</strong></p>
<p>A prompt is production configuration. It should be in version control, reviewed, and associated with a measured quality number. "Someone edited the prompt and something got worse three weeks ago" is a debugging session that should not be possible.</p>
<h2 id="the-payoff">the payoff<a class="anchor" href="#the-payoff" aria-label="link to this section">#</a></h2>
<p>Once this exists:</p>
<ul><li>Model migrations are an afternoon.</li><li>Prompt changes are safe to make.</li><li>You can argue about model choice with data instead of anecdotes.</li><li>You detect provider-side drift.</li><li>Onboarding a new engineer to the AI parts of your system means handing them the eval set, which is the best available documentation of what the system is supposed to do.</li></ul>
<p>It is a day of work. It is the single highest-leverage day available in this space and it has been for three years.</p>]]></content:encoded></item><item><title>Run the incident before the incident</title><link>https://readme.news/run-the-incident-before-the-incident/</link><guid isPermaLink="true">https://readme.news/run-the-incident-before-the-incident/</guid><pubDate>Wed, 13 May 2026 09:00:00 +0000</pubDate><description>Game days, failure injection, and the specific reason your untested runbook is wrong.</description><content:encoded><![CDATA[<p>Every runbook that has never been executed is wrong. Not might be — is. The commands have changed, the <a class="xref" href="/the-dashboard-nobody-looks-at/" title="The dashboard nobody looks at">dashboard</a> moved, the person who wrote it left, and the system it describes has been modified fourteen times.</p>
<p>The only way to find out is to run it, and the only good time to run it is when nothing is actually broken.</p>
<h2 id="what-a-game-day-is">what a game day is<a class="anchor" href="#what-a-game-day-is" aria-label="link to this section">#</a></h2>
<p>A scheduled exercise where you deliberately break something in a controlled way and have the <a class="xref" href="/on-call-is-a-design-problem/" title="On-call is a design problem">on-call</a> rotation respond as if it were real.</p>
<p>Two hours. A stated scenario. The people who would actually respond. Real tools, real runbooks, real dashboards. Someone taking notes on everything that did not work.</p>
<p>That is the whole practice. It is not chaos engineering in the automated-random sense, though that is a good adjacent practice. It is a rehearsal.</p>
<h2 id="what-you-find-reliably">what you find, reliably<a class="anchor" href="#what-you-find-reliably" aria-label="link to this section">#</a></h2>
<p>I have never run one of these that did not find at least four of the following:</p>
<p><strong>The runbook references something that does not exist.</strong> A dashboard that was renamed, a script in a repository that was archived, an alias nobody has.</p>
<p><strong>Nobody has the access.</strong> The runbook says to restart the service. The on-call engineer does not have permission and does not know who does. This is the single most common finding.</p>
<p><strong>The dashboard does not show the thing.</strong> You built the alert. You never built the view that tells you what to do about it.</p>
<p><strong>The escalation path is a person, not a rotation.</strong> "Ask Sarah." Sarah is on vacation. Sarah left last year.</p>
<p><strong>Nobody knows the customer impact.</strong> The service is degraded. Which customers? Which features? Nobody can answer, so nobody can decide how urgent it is.</p>
<p><strong>Recovery has an undocumented step.</strong> The service restarts and does not work because a cache must be cleared first, which is knowledge that lives in one person's head.</p>
<p><strong>The communication path is unclear.</strong> Who tells the customers? Who updates the status page? Who has the credentials for the status page?</p>
<p>Every one of those is cheap to fix on a Tuesday afternoon and extremely expensive to discover at 3 a.m.</p>
<h2 id="the-scenarios-worth-running">the scenarios worth running<a class="anchor" href="#the-scenarios-worth-running" aria-label="link to this section">#</a></h2>
<p>Start with the boring ones. Exotic <a class="xref" href="/boring-technology-revisited/" title="Boring technology, revisited">failure modes</a> are fun and the common ones are what actually happens.</p>
<p><strong>The database primary fails over.</strong> Does the application reconnect? How long? Does anything need a manual restart?</p>
<p><strong>A dependency returns 500s.</strong> Not down — erroring. Does your circuit breaker work? Does your retry policy make it worse?</p>
<p><strong>A dependency gets slow.</strong> Harder than down and much more common. Does your timeout fire? Do connections exhaust? Does the slowness propagate to callers?</p>
<p><strong>Disk fills up.</strong> On the database host, on the log host, on the application host. This is the single most common self-inflicted outage.</p>
<p><strong>A bad deploy.</strong> Deploy something broken to staging and time the rollback. Not the theoretical rollback — the actual one, executed by the actual on-call person.</p>
<p><strong>Certificate expiration.</strong> Set one to expire in staging. Watch what happens. Most teams discover their monitoring does not cover this.</p>
<p><strong>The person who knows is unavailable.</strong> Run a game day where the subject matter expert is explicitly not allowed to help. This finds the knowledge concentration problems that nothing else does.</p>
<h2 id="how-to-run-one-that-works">how to run one that works<a class="anchor" href="#how-to-run-one-that-works" aria-label="link to this section">#</a></h2>
<p><strong>Announce it.</strong> Do not surprise people. Surprise exercises generate resentment and teach people to distrust the process. Everyone should know it is a drill.</p>
<p><strong>Staging first, production eventually.</strong> Staging finds most of the runbook problems. Production finds the ones that only exist because staging is not production, which are real and are the ones that matter most.</p>
<p><strong>Have a stop condition.</strong> A named person who can call it off, and a defined way to revert whatever you broke.</p>
<p><strong>Write down every friction point</strong>, including small ones. "It took four minutes to find the right dashboard" is a real finding.</p>
<p><strong>Fix things within a week.</strong> A game day that produces a list nobody acts on is theater, and the second one will have lower attendance.</p>
<p><strong>Do it quarterly.</strong> Systems change. A runbook validated a year ago is a runbook that has not been validated.</p>
<h2 id="the-cultural-part">the cultural part<a class="anchor" href="#the-cultural-part" aria-label="link to this section">#</a></h2>
<p>The purpose is to find gaps in the system, not gaps in people.</p>
<p>If someone cannot resolve the scenario, that is a finding about documentation, tooling, or access — not about them. Say this explicitly before you start, and mean it, or people will optimize for looking competent rather than for surfacing problems.</p>
<p>The best outcome of a game day is a long list of things that went wrong, discovered by people who were not stressed, on a schedule, with time to fix them.</p>]]></content:encoded></item><item><title>Accessibility is a testing problem</title><link>https://readme.news/accessibility-is-a-testing-problem/</link><guid isPermaLink="true">https://readme.news/accessibility-is-a-testing-problem/</guid><pubDate>Wed, 22 Apr 2026 09:00:00 +0000</pubDate><description>Most accessibility failures are mechanical and catchable. Treating it as a specialist concern is why it doesn&#x27;t get fixed.</description><content:encoded><![CDATA[<p>Accessibility gets treated as a specialist discipline requiring expert audits. Parts of it are. Most of what actually breaks is mechanical, catchable automatically, and would be fixed if it failed the build.</p>
<h2 id="what-automated-tooling-catches">what automated tooling catches<a class="anchor" href="#what-automated-tooling-catches" aria-label="link to this section">#</a></h2>
<p>Roughly a third to a half of real accessibility issues are detectable by a linter or a test:</p>
<ul><li>Images without alternative text.</li><li>Form inputs without associated labels.</li><li>Insufficient color contrast.</li><li>Missing document language.</li><li>Heading levels that skip.</li><li>Buttons and links with no accessible name.</li><li>ARIA attributes that are invalid or applied to the wrong role.</li><li>Interactive elements not reachable by keyboard.</li><li>Duplicate or missing landmark regions.</li></ul>
<p>Every one of those is a fixed rule with a fixed check, and every one of them is a real barrier for someone.</p>
<p>Getting the automatable third fixed is not the whole job and it is a much better place than most applications are.</p>
<h2 id="the-setup">the setup<a class="anchor" href="#the-setup" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">js</span><pre><code class="lang-js">// component test
import { axe } from 'vitest-axe';

test('checkout form has no automatically detectable a11y violations', async () =&gt; {
  const { container } = render(&lt;CheckoutForm /&gt;);
  expect(await axe(container)).toHaveNoViolations();
});</code></pre></div>
<p>Plus a linter for the static cases:</p>
<div class="code"><span class="code-lang">json</span><pre><code class="lang-json">{ "extends": ["plugin:jsx-a11y/recommended"] }</code></pre></div>
<p>Plus a crawl of key pages in CI against the built site.</p>
<p>That is an afternoon of setup. It will find violations on day one — every codebase has them — so run it in report-only mode first, fix the backlog, then make it blocking. Making it blocking with a hundred existing violations gets it disabled within a week.</p>
<h2 id="what-tooling-cannot-catch">what tooling cannot catch<a class="anchor" href="#what-tooling-cannot-catch" aria-label="link to this section">#</a></h2>
<p>Being honest about the limits matters, because "we run axe" becomes a claim of compliance that it does not support.</p>
<p>**Whether the alt text is <em>good</em>.** <code>alt="image"</code> passes. It is useless. A screenshot of a chart needs a description of what the chart shows, not the word "chart."</p>
<p><strong>Whether the reading order makes sense.</strong> CSS can visually reorder content while the DOM order stays wrong, and a screen reader follows the DOM.</p>
<p><strong>Whether focus management works.</strong> Open a modal — where does focus go? Close it — does it return? Navigate to a new route — is focus reset and announced? These are the most common real-world failures in single-page applications and no automated check catches them.</p>
<p><strong>Whether dynamic content is announced.</strong> A live region that updates too often is worse than one that does not update at all.</p>
<p><strong>Whether it is actually usable.</strong> A page can pass every automated check and be completely impractical to navigate.</p>
<h2 id="the-manual-checks-worth-doing-routinely">the manual checks worth doing routinely<a class="anchor" href="#the-manual-checks-worth-doing-routinely" aria-label="link to this section">#</a></h2>
<p>Three, each a few minutes, on every significant feature:</p>
<p><strong>1. Unplug the mouse.</strong> Navigate the entire flow with <code>Tab</code>, <code>Shift+Tab</code>, <code>Enter</code>, <code>Space</code>, arrows, and <code>Escape</code>. Can you complete the task? Can you always see where focus is? Does focus ever get trapped, or jump somewhere unexpected?</p>
<p>This single test finds more real problems than any automated tool.</p>
<p><strong>2. Zoom to 400%.</strong> Browser zoom, at a 1280px viewport. Does content reflow, or does it require horizontal scrolling? Is anything cut off? This is a WCAG requirement and it is failed constantly.</p>
<p><strong>3. Turn on a screen reader for five minutes.</strong> VoiceOver on macOS (<code>Cmd+F5</code>), NVDA on Windows. You will be bad at it and that is fine — you are not evaluating your skill, you are listening to whether your page announces anything coherent.</p>
<p>Most developers who do this once are permanently changed by it, because the experience of hearing your own <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a> read aloud as "button, button, button, link, clickable" is clarifying.</p>
<h2 id="the-framing-that-gets-it-prioritized">the framing that gets it prioritized<a class="anchor" href="#the-framing-that-gets-it-prioritized" aria-label="link to this section">#</a></h2>
<p>Not "it is the right thing to do." That is true and it does not survive a prioritization meeting.</p>
<p><strong>It is a legal requirement</strong> in many jurisdictions, with an enforcement trend that is going up rather than down, and in the EU the European Accessibility Act now applies to a broad category of consumer-facing digital services.</p>
<p><strong>It is a larger market than most product managers assume.</strong> Roughly one in six people has a disability. Not all of those affect software use; enough do.</p>
<p><strong>It overlaps with quality generally.</strong> Semantic HTML, keyboard support, and clear focus states make interfaces better for everyone. Keyboard shortcuts are an accessibility feature that power users love.</p>
<p><strong>It is much cheaper to build in than to retrofit.</strong> An audit that finds two hundred issues in a shipped product is a quarter of remediation work. Catching them at the component level costs minutes.</p>
<h2 id="the-single-highest-value-practice">the single highest-value practice<a class="anchor" href="#the-single-highest-value-practice" aria-label="link to this section">#</a></h2>
<p>Use semantic HTML.</p>
<p>A <code>&lt;button&gt;</code> is focusable, keyboard-activatable, announced as a button, and works with every assistive technology, for free. A <code>&lt;div onclick&gt;</code> requires you to add a role, a tabindex, key handlers for Enter and Space, and focus styles — and you will get one of them wrong.</p>
<p>The overwhelming majority of accessibility problems in modern web applications come from reimplementing native elements badly. Use the element. It already works.</p>]]></content:encoded></item><item><title>Feature flags and the state space nobody tests</title><link>https://readme.news/feature-flags-and-the-state-space-nobody-tests/</link><guid isPermaLink="true">https://readme.news/feature-flags-and-the-state-space-nobody-tests/</guid><pubDate>Mon, 02 Mar 2026 09:00:00 +0000</pubDate><description>Twenty flags is a million configurations. You are testing one of them. Here&#x27;s how to keep that from being a problem.</description><content:encoded><![CDATA[<p>Feature flags decouple deploy from release, enable gradual rollout, and give you a kill switch. All of that is genuinely valuable.</p>
<p>They also multiply your state space by two for every flag you add, and almost nobody accounts for that.</p>
<h2 id="the-arithmetic">the arithmetic<a class="anchor" href="#the-arithmetic" aria-label="link to this section">#</a></h2>
<p>Twenty boolean flags is 2^20 possible configurations — over a million. Your test suite exercises the default configuration. Your staging environment exercises one other. Production is running dozens simultaneously, because different flags are on for different user segments.</p>
<p>Which means: <strong>most of the configurations your users are running have never been executed anywhere before.</strong></p>
<p>Most of the time this is fine, because most flags are independent. The failures come from the ones that are not, and you find out about those from a support ticket.</p>
<h2 id="the-failure-modes">the failure modes<a class="anchor" href="#the-failure-modes" aria-label="link to this section">#</a></h2>
<p><strong>Interaction bugs.</strong> Flag A changes the data format. Flag B reads that data. Both work alone. Together, one of them is reading a shape it does not expect. This is the classic and it is very hard to catch because neither change is wrong.</p>
<p><strong>Flag debt.</strong> A flag that has been at 100% for a year is still in the code, with its dead branch, and nobody remembers whether it can be removed. The dead branch does not compile-error and does not test-fail; it just sits there being wrong.</p>
<p><strong>The stale branch.</strong> Once a flag is at 100%, the disabled path stops being exercised. Six months later someone flips it as a rollback and discovers the old path no longer works with the current schema. Your kill switch is broken and you find out during an incident.</p>
<p><strong>Configuration drift.</strong> Flags set differently in staging and production means staging tests a configuration nobody runs.</p>
<h2 id="the-rules-that-keep-this-manageable">the rules that keep this manageable<a class="anchor" href="#the-rules-that-keep-this-manageable" aria-label="link to this section">#</a></h2>
<p><strong>Every flag has an owner and an expiry date at creation.</strong> Not a suggestion — a required field. A flag that has passed its expiry shows up in a report and somebody has to either extend it with a reason or delete it.</p>
<p><strong>Flags are deleted, not left at 100%.</strong> The cleanup is part of the work, not a follow-up ticket. A flag rollout is not done when it reaches 100%; it is done when the flag and the dead branch are gone.</p>
<p><strong>Flags never nest.</strong> If flag A only means something when flag B is on, you have created a configuration nobody can reason about. Combine them into one flag with three states, or restructure.</p>
<p><strong>Kill switches are a separate category</strong> with separate rules. They live forever, they are documented, and they are <em>tested on a schedule</em> — a quarterly exercise where you flip each one in staging and verify it works. An untested kill switch is not a kill switch.</p>
<p><strong>Test both branches.</strong> Your test suite should run the critical path with each significant flag both on and off. Not the full combinatorial space — that is impossible — but each flag independently against the default configuration. That catches the "old path rotted" failure, which is the expensive one.</p>
<h2 id="the-taxonomy-that-helps">the taxonomy that helps<a class="anchor" href="#the-taxonomy-that-helps" aria-label="link to this section">#</a></h2>
<p>Four kinds of flag, with different lifecycles:</p>
<div class="table-wrap"><table><thead><tr><th style="text-align:left">kind</th><th style="text-align:left">lifetime</th><th style="text-align:left">example</th></tr></thead><tbody><tr><td style="text-align:left">release</td><td style="text-align:left">days to weeks</td><td style="text-align:left">shipping a new checkout</td></tr><tr><td style="text-align:left">experiment</td><td style="text-align:left">weeks</td><td style="text-align:left">A/B test</td></tr><tr><td style="text-align:left">ops / kill switch</td><td style="text-align:left">permanent</td><td style="text-align:left">disable recommendations</td></tr><tr><td style="text-align:left">permission</td><td style="text-align:left">permanent</td><td style="text-align:left">enterprise-tier feature</td></tr></tbody></table></div>
<p>The first two must expire. The last two must be documented and tested. Conflating them is how you get a thousand flags and no idea which matter.</p>
<h2 id="the-observability-part">the observability part<a class="anchor" href="#the-observability-part" aria-label="link to this section">#</a></h2>
<p>Every event you log should carry the flag configuration that produced it.</p>
<div class="code"><span class="code-lang">json</span><pre><code class="lang-json">{ "event": "checkout", "status": 500, "flags": ["new_pricing", "fast_path"] }</code></pre></div>
<p>Without this, a bug that only affects one flag combination is undiagnosable — you see errors, you cannot correlate them to anything, and you spend a day guessing. With it, the correlation is a single query.</p>
<p>This is a fifteen-minute change and it is the single highest-value thing you can do if you use flags at all.</p>
<h2 id="the-counterargument">the counterargument<a class="anchor" href="#the-counterargument" aria-label="link to this section">#</a></h2>
<p>Some teams respond to all of this by using fewer flags and doing more trunk-based deployment with fast rollback.</p>
<p>That is a legitimate position and it works when your rollback is genuinely fast and your changes are genuinely reversible. Flags are a tool for when they are not — when a change is expensive to revert, when you need per-segment control, or when release timing must be decoupled from deployment for business reasons.</p>
<p>Use them for those. Do not use them because they feel safer, because a flag you do not test is not safety, it is the appearance of it.</p>]]></content:encoded></item><item><title>Verification is the whole job now</title><link>https://readme.news/verification-is-the-whole-job-now/</link><guid isPermaLink="true">https://readme.news/verification-is-the-whole-job-now/</guid><pubDate>Thu, 08 Jan 2026 09:00:00 +0000</pubDate><description>Producing code got cheap. Knowing whether it&#x27;s right did not. Everything downstream follows from that asymmetry.</description><content:encoded><![CDATA[<p>Here is the single fact that explains most of what has happened to software engineering in the last two years:</p>
<p><strong>Generation got roughly two orders of magnitude cheaper. Verification did not get cheaper at all.</strong></p>
<p>Everything else — the review bottleneck, the arguments about junior hiring, the sudden importance of tests, the unease people cannot articulate about agent-written code — is downstream of that asymmetry.</p>
<h2 id="why-verification-did-not-get-cheaper">why verification did not get cheaper<a class="anchor" href="#why-verification-did-not-get-cheaper" aria-label="link to this section">#</a></h2>
<p>You might expect it to. Surely a model can check code as well as it can write it?</p>
<p>It can, sort of, and not in the way that matters. Three problems:</p>
<p><strong>The checker has the same blind spots as the writer.</strong> A model that does not know your system produced code with an assumption baked in. The same model reviewing it holds the same assumption and does not see it. Independent errors can be filtered by voting; correlated errors cannot.</p>
<p><strong>Verification requires knowing what correct means.</strong> That is a property of your domain, your users, your obligations — not of the code. A model can verify that code does what it says. It cannot verify that what it says is what you needed, because that information was never written down anywhere.</p>
<p><strong>Confidence is decoupled from correctness.</strong> Generated code is fluent. Fluency is the signal humans use to assess competence, and it has been severed from correctness. Reviewing text that reads well but is subtly wrong is much harder than reviewing text that reads badly, and everyone underestimates this.</p>
<h2 id="what-this-changes">what this changes<a class="anchor" href="#what-this-changes" aria-label="link to this section">#</a></h2>
<p><strong>Tests stop being a chore and become the primary artifact.</strong> If verification is the constraint, then anything that makes verification automatic is worth enormous investment. A test suite is executable verification. It was always valuable; it is now the thing that determines your throughput.</p>
<p>The practical implication: write tests first, review those carefully, and let the implementation be cheap. Invert the effort. The specification is the part you have to get right; the code is increasingly a derived artifact.</p>
<p><strong><a class="xref" href="/type-systems-and-the-cost-of-being-right/" title="Type systems and the cost of being right">Type systems</a> get a promotion.</strong> Every constraint the compiler can check is a constraint you do not have to verify by reading. Languages with expressive type systems and strong static guarantees are worth more than they were, because they convert human verification into machine verification.</p>
<p>I have watched teams that were ambivalent about strict TypeScript become evangelists in the space of a year, and the reason is always the same: it catches the class of error that agent-written code produces most often.</p>
<p><strong>Small diffs become non-negotiable.</strong> Verification cost scales superlinearly with diff size, because the number of interactions you have to reason about grows faster than the lines. A machine can produce a 2,000-line change effortlessly. Accepting it is not a favor to anyone.</p>
<p><strong>Property-based testing gets its moment.</strong> If you can state an invariant, you can verify a very large input space cheaply. Invariants are exactly the kind of thing that a human should specify and a machine should check. This technique has been niche for twenty years and it is the right shape for this moment.</p>
<p><strong>Observability becomes a verification tool.</strong> If you cannot verify everything before deploy, you verify in production: <a class="xref" href="/feature-flags-and-the-state-space-nobody-tests/" title="Feature flags and the state space nobody tests">feature flags</a>, canaries, and metrics that detect wrongness quickly. Fast detection is a substitute for perfect pre-verification, and it is often a better investment.</p>
<h2 id="the-skill-that-matters">the skill that matters<a class="anchor" href="#the-skill-that-matters" aria-label="link to this section">#</a></h2>
<p>The engineers getting genuine leverage out of these tools all share one trait, and it is not prompting.</p>
<p>It is that they can look at a plausible-looking piece of code and say "that retry loop will hammer the upstream on a 429" or "that will deadlock under concurrent writes" or "that assumes the list is sorted and nothing sorts it."</p>
<p>That skill comes from having debugged those exact failures. It does not come from reading about them.</p>
<p>Which produces the uncomfortable question everybody is circling: the traditional way to acquire it was to write a lot of code badly and then fix it. If that apprenticeship is being automated away, what replaces it?</p>
<h2 id="the-answer-i-have-which-is-partial">the answer I have, which is partial<a class="anchor" href="#the-answer-i-have-which-is-partial" aria-label="link to this section">#</a></h2>
<p>Deliberately do the hard part yourself sometimes.</p>
<p>Not out of nostalgia. Because the judgment is the product, and the judgment is built by doing. If you delegate every debugging session, you will be worse at debugging in two years, and debugging is the thing you are being paid for now.</p>
<p>Pick the gnarliest bug of the week and do it by hand. Read the code the agent wrote in the module you own, all of it, once a month. Write the tricky concurrency code yourself and let the machine do the CRUD.</p>
<p>That is not a workflow recommendation for efficiency. It is a training regimen, and treating it as one is the honest framing.</p>]]></content:encoded></item><item><title>The test suite is a design document</title><link>https://readme.news/the-test-suite-is-a-design-document/</link><guid isPermaLink="true">https://readme.news/the-test-suite-is-a-design-document/</guid><pubDate>Fri, 09 May 2025 09:00:00 +0000</pubDate><description>If your tests are hard to write, your design is wrong. That&#x27;s the signal, and most teams ignore it.</description><content:encoded><![CDATA[<p>The most useful thing tests do is not catch regressions. It is tell you, before you have committed to anything, that your design is bad.</p>
<p>That signal is available for free, it is extremely reliable, and most teams respond to it by making the tests worse instead of the design better.</p>
<h2 id="the-tells">the tells<a class="anchor" href="#the-tells" aria-label="link to this section">#</a></h2>
<p><strong>You need six mocks to test one function.</strong> The function has six dependencies. That is the actual problem. The mocks are the messenger.</p>
<p><strong>You are testing private methods.</strong> The class is doing more than one thing and you know it, which is why the interesting behavior is not reachable from the public <a class="xref" href="/the-interface-is-the-product/" title="The interface is the product">interface</a>. Extract the thing you want to test into its own unit with its own public interface.</p>
<p><strong>Setup is longer than the assertion.</strong> The unit under test requires a large amount of world to exist before it does anything. That is coupling, expressed in lines of setup code.</p>
<p><strong>You cannot test it without a database.</strong> Sometimes true and unavoidable. Usually it means business logic and persistence are interleaved, and separating them would make both testable and both clearer.</p>
<p><strong>The test breaks whenever you refactor.</strong> You are asserting on implementation rather than behavior. Either the test is bad, or the unit has no meaningful behavioral contract, which is worse.</p>
<p><strong>You cannot describe what the test verifies in one sentence.</strong> Then neither can the code.</p>
<h2 id="the-mechanism">the mechanism<a class="anchor" href="#the-mechanism" aria-label="link to this section">#</a></h2>
<p>Tests are the first <em>consumer</em> of your interface. They are the only place, before production, where you have to actually use the thing you built rather than look at it.</p>
<p>Everything that is awkward about consuming your interface shows up in the test as friction: too many parameters, unclear ordering, hidden global state, temporal coupling between calls, error conditions that cannot be triggered deliberately.</p>
<p>This is why "test-first" works for people it works for, and it is not because of any ritual. It is because writing the usage before the implementation forces you to design the interface from the caller's perspective, which is the only perspective that matters.</p>
<p>You do not have to write tests first to get this. You have to <em>notice</em> when the test is hard and treat that as data.</p>
<h2 id="the-common-wrong-response">the common wrong response<a class="anchor" href="#the-common-wrong-response" aria-label="link to this section">#</a></h2>
<p>Team hits testing friction. Team adds a mocking framework with more power. Now you can mock static methods, patch constructors, intercept module loading, and rewrite the world.</p>
<p>The friction goes away. The design problem does not. It is now permanently invisible, encased in tooling, and the test suite has become a second implementation of the system that must be maintained in parallel with the first.</p>
<p>Powerful mocking frameworks are a tool for testing legacy code you cannot change. Reaching for one in new code is choosing not to hear the signal.</p>
<h2 id="what-good-looks-like">what good looks like<a class="anchor" href="#what-good-looks-like" aria-label="link to this section">#</a></h2>
<div class="code"><span class="code-lang">python</span><pre><code class="lang-python">def test_order_total_applies_bulk_discount():
    order = Order(items=[Item(price=10_00, qty=100)])
    assert order.total() == 900_00</code></pre></div>
<p>No mocks. No setup. No database. One sentence describes it. It breaks if and only if the discount rule changes.</p>
<p>That test is easy to write because <code>Order.total</code> is a pure function of its data. Making it a pure function of its data was the design decision. The test just confirmed it was a good one.</p>
<h2 id="the-one-exception-worth-stating">the one exception worth stating<a class="anchor" href="#the-one-exception-worth-stating" aria-label="link to this section">#</a></h2>
<p>Sometimes the system is genuinely complex and the tests are genuinely complicated, because the domain is. Distributed consensus, a compiler backend, a scheduler. Difficulty is not always a design smell.</p>
<p>The distinguishing question: is the test complicated because the <em>problem</em> is hard, or because <em>getting to the problem</em> is hard? Complexity in the assertions is fine. Complexity in the setup is the smell.</p>]]></content:encoded></item>
</channel>
</rss>
