Timeouts: every one of them
A request crosses a dozen components with a timeout each, and almost nobody has ever added them up.
Every layer of your stack has a timeout. Most of them are defaults. Almost nobody has written them down in one place and checked that they make sense together.
The result is systems where the client gives up at 10 seconds, the server keeps working for 60, and the database holds a lock for 300.
the inventory#
For a single HTTP request, in rough order:
| layer | typical default |
|---|---|
| browser / client library | 30s or none |
| DNS resolution | 5s per attempt |
| TCP connect | 20–75s (OS) |
| TLS handshake | inherits connect |
| load balancer idle | 60s |
| reverse proxy read | 60s |
| application server request | often none |
| HTTP client to downstream | often none |
| connection pool acquire | 30s |
| database statement | often none |
| database lock wait | often none |
Two things stand out. "Often none" appears five times — most application-level timeouts are unset by default. And the values are not coordinated with each other in any way.
the rule that fixes most of it#
Timeouts must decrease as you go deeper.
client 10s
└─ load balancer 9s
└─ application 8s
└─ downstream call 3s (with 1 retry → 6s worst case)
└─ connection acquire 1s
└─ database statement 2sEach layer must be shorter than its caller, with room for retries. If an inner layer can outlast its caller, the caller gives up while the work continues — which means you are burning capacity on results nobody will receive. Under load, that is most of your capacity.
deadline propagation#
The better version of the rule: do not configure each timeout independently. Pass the deadline down.
// caller has 10s; every downstream inherits what remains
ctx, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()
// 3s into the request, this call gets at most 7s, automatically
resp, err := downstream.Fetch(ctx, id)Go's context, gRPC's deadlines and equivalents elsewhere exist for this. Once you propagate deadlines, a slow first step automatically shortens the budget for later ones, and work never continues past the point where anyone is waiting.
This is the single highest-value change available in most service codebases, and it is usually a few days of threading a parameter through.
the ones people forget#
Lock wait timeouts. A transaction waiting on a row lock with no timeout waits forever. Set lock_timeout in Postgres, innodb_lock_wait_timeout in MySQL.
DDL statements. A migration that cannot get its lock will queue behind a long transaction — and everything else queues behind it. SET lock_timeout = '3s' before DDL is the line that prevents a large share of migration incidents.
Idle-in-transaction. A connection that opened a transaction and went away holds locks and blocks vacuum indefinitely. idle_in_transaction_session_timeout is the safety net.
Client-side connect vs. read. These are different, and most libraries let you set them separately. A short connect timeout with a long read timeout is usually what you want.
DNS. Rarely configured, occasionally the whole problem.
write them down#
One table, in the repository, listing every timeout in the request path with its current value and where it is configured.
The exercise takes an afternoon and reliably finds at least one inversion — an inner layer waiting longer than the outer one — and at least two places where the value is a framework default nobody chose.
That document then becomes something you can review when latency changes, rather than a set of numbers scattered across six config files and three languages.
the failure mode this prevents#
Without coordinated timeouts, a slow dependency does not degrade your service — it exhausts it. Requests pile up waiting on something their callers abandoned long ago, connections stay held, the pool empties, and healthy requests start failing for want of a connection.
That is a full outage caused by one slow downstream, and the difference between it and a brief latency blip is entirely whether the numbers were coordinated.
— Dom, September 10, 2026