tech, developers, and the code underneath

issue 216· essay·

The debugging notebook

Writing down what you tried is the cheapest debugging technique there is, and almost nobody does it.

Two hours into a hard bug, most people cannot remember which of the six things they tried actually changed the behaviour, or whether they have already checked the thing they are about to check again.

The fix is a text file. It costs nothing and it is the difference between an investigation and a random walk.

what goes in it#

14:02  Symptom: 502s on /api/orders, ~4% of requests, started ~13:40.
       Not correlated with deploy (last one 11:20).

14:08  FACT: only POST. GET is clean. (checked: 30min of access logs)
14:15  FACT: all failures have Content-Length > 64KB.
       → hypothesis: body size limit somewhere
14:22  Checked nginx client_max_body_size = 10m. Not it.
14:31  Checked the app's own limit — 100MB. Not it.
14:40  FACT: failures all hit pod-7 and pod-11 (of 12). Others clean.
       → hypothesis abandoned: not size, size just correlates with
         which client sends big bodies, and that client is sticky
14:52  pod-7 and pod-11 started 13:38. Two minutes before symptoms.
       → new hypothesis: bad config in the new pods
15:01  CONFIRMED: those pods have UPSTREAM_TIMEOUT unset → defaulting to 1s

An hour of work, ten lines. Note what it captures: facts with how they were checked, hypotheses explicitly raised and explicitly abandoned, and the moment a correlation turned out to be a red herring.

why it works#

It stops you re-checking things. The single biggest time sink in a long debugging session is doing the same check twice because you cannot remember the result.

It forces hypotheses to be explicit. Written down, "it's a size limit" is a claim you can test and discard. Held in your head, it quietly steers every subsequent action.

It separates fact from guess. The most common way a debugging session goes wrong is a guess getting promoted to a fact through repetition. Labelling them differently prevents it.

It makes handoff possible. When you hit the end of your day, or the end of your knowledge, the notebook is the handoff. Without it the next person starts from zero.

It becomes the postmortem. The timeline is already written, with the dead ends included — which are the most instructive part and the first thing lost to memory.

the three prompts#

When stuck, the notebook gives you somewhere to answer these:

"What changed?" Deploys, config, data volume, dependency versions, traffic shape, time of day. Most bugs in a previously-working system are caused by a change, and enumerating them beats staring at code.

"What do I actually know?" Read back your own FACT lines. Frequently the answer is that you know much less than you thought, and one of your assumptions was never verified.

"What would prove me wrong?" Not what would confirm the theory — what would kill it. If you cannot name that, the theory is not yet a hypothesis.

the format does not matter#

A scratch file, a comment thread on the ticket, a channel where you narrate to yourself. What matters is that it is written, timestamped, and distinguishes what you checked from what you suspect.

The one thing that does not work is keeping it in your head, because the whole point is that your head is where the confusion is.

the compounding version#

Keep the notebooks. A directory of them, named by date and symptom.

Six months in, you have a searchable record of how this system actually fails, which is knowledge that otherwise exists only as a vague feeling in whoever was on call. When the same class of bug returns — and it does — the search takes a minute and the fix takes five.

That is the cheapest institutional memory available to an engineering team, and it is built one text file at a time.

— Dom, September 5, 2026

get README in your inbox

One dispatch, no noise. Tech and developer news, plus the occasional long piece on the craft.

subscribe →