tech, developers, and the code underneath

issue 220· essay·

The graph nobody drew

Your service dependency graph exists whether or not anyone has written it down. Writing it down takes an afternoon.

Ask an engineering team to list what their service depends on to serve a request. You will get the database and two or three obvious APIs.

The real list is usually fifteen things, and the difference between the two lists is where the surprising outages come from.

the full inventory#

For one service, one request path:

  • Datastores. Primary, replicas, cache, search index, object storage.
  • Internal services it calls synchronously.
  • Third-party APIs — payments, email, geocoding, feature flags, auth.
  • Infrastructure control planes. Service discovery, DNS, the secret manager, the config service. These are the ones nobody lists and they are the ones that take everything down.
  • Identity. The token issuer, the JWKS endpoint it fetches keys from.
  • Observability. If your telemetry agent blocks when the collector is down — and some do — it is in your request path.
  • The container registry, at start-up. Fine until you need to scale during an incident and cannot pull an image.
  • Certificate infrastructure, including OCSP and the ACME endpoint.
  • NTP. Rarely, spectacularly.

The pattern: the dependencies people list are the ones they call. The ones that hurt are the ones they need in order to start, authenticate, or scale.

the two questions per dependency#

For each item, two answers, written down:

What happens when it is slow? Slow is worse than down and much more common. Down fails fast; slow holds your connections, exhausts your pool, and turns one dependency's bad day into your outage.

What happens when it is unavailable? Three possible answers, and only one is a decision:

  • We fail. Fine, if the dependency is genuinely essential — a database for a write path.
  • We degrade. Serve stale cache, hide the feature, use a default. This is the answer for most non-essential dependencies and it requires code that usually does not exist yet.
  • We do not know. This is the real answer for most dependencies, and it is the finding.

the classification that matters#

Sort every dependency into two buckets:

Hard — the request genuinely cannot be served without it. Should be a very short list. Each one caps your maximum availability at its own.

Soft — the request can be served in degraded form. Every soft dependency needs a timeout, a fallback, and a circuit breaker, or it is a hard dependency that nobody has admitted to.

The arithmetic is the reason this matters: five hard dependencies at 99.9% each gives you at best 99.5%, before any failure of your own. If your availability target is higher than that, some of those dependencies have to become soft, and that is an engineering project rather than a target you can declare.

the feature flag trap#

Worth naming specifically because it catches good teams. Feature flag services are adopted as a safety mechanism — a way to turn things off when they go wrong.

Then the flag SDK becomes a hard dependency on the request path, and when the flag service has an outage your service does too. The tool you adopted to make incidents smaller has made one bigger.

The fix is standard for this whole category and rarely applied: cache flag values locally, evaluate from the cache, refresh in the background, and ship a default in the binary. Then the flag service being down means flags are stale, which is survivable.

how to draw it#

Do not buy anything. Open a file:

orders-api
  HARD  postgres-primary        write path; no fallback
  HARD  auth-jwks               cached 1h, so 1h of grace
  SOFT  redis                   cache; miss → primary, 3x latency
  SOFT  inventory-svc           timeout 300ms → show "check availability"
  SOFT  pricing-svc             timeout 200ms → list price, no promo
  SOFT  flags-svc               local cache + baked defaults
  BOOT  vault                   secrets at start; running pods unaffected
  BOOT  registry                image pull; blocks scale-up only

Twenty minutes per service. The BOOT category — needed to start but not to serve — is the one that catches people out during incidents, because those dependencies are invisible right up until you need to replace a pod.

Then, for every SOFT line, check that the fallback described actually exists in code. That check is where the real findings are, and it is usually not the timeout that is missing but the fallback behind it.

get README in your inbox

One dispatch, no noise. Tech and developer news, plus the occasional long piece on the craft.

subscribe →