The graph nobody drew
Your service dependency graph exists whether or not anyone has written it down. Writing it down takes an afternoon.
Ask an engineering team to list what their service depends on to serve a request. You will get the database and two or three obvious APIs.
The real list is usually fifteen things, and the difference between the two lists is where the surprising outages come from.
the full inventory#
For one service, one request path:
- Datastores. Primary, replicas, cache, search index, object storage.
- Internal services it calls synchronously.
- Third-party APIs — payments, email, geocoding, feature flags, auth.
- Infrastructure control planes. Service discovery, DNS, the secret manager, the config service. These are the ones nobody lists and they are the ones that take everything down.
- Identity. The token issuer, the JWKS endpoint it fetches keys from.
- Observability. If your telemetry agent blocks when the collector is down — and some do — it is in your request path.
- The container registry, at start-up. Fine until you need to scale during an incident and cannot pull an image.
- Certificate infrastructure, including OCSP and the ACME endpoint.
- NTP. Rarely, spectacularly.
The pattern: the dependencies people list are the ones they call. The ones that hurt are the ones they need in order to start, authenticate, or scale.
the two questions per dependency#
For each item, two answers, written down:
What happens when it is slow? Slow is worse than down and much more common. Down fails fast; slow holds your connections, exhausts your pool, and turns one dependency's bad day into your outage.
What happens when it is unavailable? Three possible answers, and only one is a decision:
- We fail. Fine, if the dependency is genuinely essential — a database for a write path.
- We degrade. Serve stale cache, hide the feature, use a default. This is the answer for most non-essential dependencies and it requires code that usually does not exist yet.
- We do not know. This is the real answer for most dependencies, and it is the finding.
the classification that matters#
Sort every dependency into two buckets:
Hard — the request genuinely cannot be served without it. Should be a very short list. Each one caps your maximum availability at its own.
Soft — the request can be served in degraded form. Every soft dependency needs a timeout, a fallback, and a circuit breaker, or it is a hard dependency that nobody has admitted to.
The arithmetic is the reason this matters: five hard dependencies at 99.9% each gives you at best 99.5%, before any failure of your own. If your availability target is higher than that, some of those dependencies have to become soft, and that is an engineering project rather than a target you can declare.
the feature flag trap#
Worth naming specifically because it catches good teams. Feature flag services are adopted as a safety mechanism — a way to turn things off when they go wrong.
Then the flag SDK becomes a hard dependency on the request path, and when the flag service has an outage your service does too. The tool you adopted to make incidents smaller has made one bigger.
The fix is standard for this whole category and rarely applied: cache flag values locally, evaluate from the cache, refresh in the background, and ship a default in the binary. Then the flag service being down means flags are stale, which is survivable.
how to draw it#
Do not buy anything. Open a file:
orders-api
HARD postgres-primary write path; no fallback
HARD auth-jwks cached 1h, so 1h of grace
SOFT redis cache; miss → primary, 3x latency
SOFT inventory-svc timeout 300ms → show "check availability"
SOFT pricing-svc timeout 200ms → list price, no promo
SOFT flags-svc local cache + baked defaults
BOOT vault secrets at start; running pods unaffected
BOOT registry image pull; blocks scale-up onlyTwenty minutes per service. The BOOT category — needed to start but not to serve — is the one that catches people out during incidents, because those dependencies are invisible right up until you need to replace a pod.
Then, for every SOFT line, check that the fallback described actually exists in code. That check is where the real findings are, and it is usually not the timeout that is missing but the fallback behind it.
— Dom, September 14, 2026