tech, developers, and the code underneath

issue 221· essay·

Why your tests are slow

Six causes, ranked by how often they are the real problem, and the fix for each.

A test suite that takes twenty minutes is not run locally. A suite that is not run locally is a suite that fails in CI, which means the feedback loop is now measured in pipeline runs.

Here is where the time actually goes, roughly in order of how often each is the dominant cause.

1. everything talks to a real database#

The most common cause by a wide margin. Each test sets up a database, inserts fixtures, runs, and tears down.

Fixes, in order of return:

Roll back instead of truncating. Wrap each test in a transaction and roll it back. Orders of magnitude faster than deleting rows, and it needs no cleanup code.

Share the schema, not the data. Create the schema once per run, not per test.

Use a tmpfs for the test database. The data does not need to survive; durable writes are pure cost. On Postgres, a data directory in memory plus fsync=off and synchronous_commit=off is dramatically faster and completely inappropriate for anything but tests.

Move the logic out of the database's reach. The deepest fix: if business logic is a pure function of its inputs, its tests do not need a database at all. Tests that are hard to write without infrastructure are telling you about your design.

2. the suite is serial#

Most runners parallelise and most projects have not turned it on, usually because it exposed shared state once and someone reverted it.

That shared state is a bug. Tests that pass only in a particular order have an ordering dependency, and ordering dependencies in tests usually mirror a real one in the code.

Turn on parallelism, fix what breaks, and you will find at least one genuine issue.

3. sleeps#

python
time.sleep(2)   # wait for the worker to pick it up

Every one of these is pure latency, and they are always tuned to the slowest machine anyone has run on. Fifty of them is a hundred seconds of doing nothing.

Replace with polling on the actual condition:

python
wait_until(lambda: job.reload().status == "done", timeout=5)

Same worst case, typically a hundredth of the elapsed time, and it fails with a useful message instead of a mysterious assertion.

4. the fixtures are enormous#

A test that needs one user builds an organisation, three teams, forty users and a year of history, because it reuses the "standard" fixture.

Build the minimum. Factories with sensible defaults, overridden per test, beat shared fixture files that grow to serve every case.

5. it is doing real network I/O#

Tests hitting real HTTP endpoints — even internal ones, even mock servers over a socket — pay connection setup and scheduling per call.

Intercept at the client layer rather than the network layer. Most languages have a way to stub the HTTP client in-process, and it is both faster and more deterministic than a local server.

The exception is contract tests, which exist to catch exactly what stubs hide. Keep those, run them separately, and do not let them into the fast suite.

6. everything is an end-to-end test#

Browser-driving tests are two to three orders of magnitude slower than unit tests. A suite made mostly of them is slow no matter what you do.

The ratio that works: a large number of fast unit tests, a moderate number of integration tests around real boundaries, and a small number of end-to-end tests covering the handful of journeys that must never break.

Inverting that ratio is the most expensive testing mistake a team can make, and it is usually the result of not trusting the lower layers rather than a deliberate choice.

the measurement first#

Before fixing anything, get the distribution:

bash
pytest --durations=25
go test ./... -json | jq -r 'select(.Action=="pass") | "\(.Elapsed) \(.Test)"' | sort -rn | head -25

Almost always, a small number of tests dominate. Fixing the slowest twenty is usually the whole job, and it is an afternoon rather than a project.

the target#

Under 10 seconds for the unit suite, run on every save. Under two minutes for everything that gates a merge.

The ten-second number is not arbitrary — it is roughly the threshold past which people stop running tests as they work and start batching. Once that happens you have lost the feedback loop, and the suite's speed stops being a convenience question and starts being a quality one.

— Dom, September 16, 2026

get README in your inbox

One dispatch, no noise. Tech and developer news, plus the occasional long piece on the craft.

subscribe →