Partial failure is not an edge case in distributed systems — it is the steady state you design around. The teams that recover fastest are not the ones with the most resilience patterns; they are the ones who decided in advance what breaks, what degrades, and what still answers.
Why failure is normal
Networks drop packets. Processes restart. Dependencies throttle. Clocks drift. Once a request crosses more than one machine, you no longer have a single failure domain — you have a graph of them. A design that assumes healthy dependencies is a design that fails closed when reality disagrees.
The goal is not “zero downtime.” The goal is bounded blast radius and predictable degradation: when something fails, users experience a limited, intentional product behavior — not an opaque 500 page or a hung request that exhausts your thread pool.
Timeouts are a design decision
Every remote call needs a timeout. Not eventually. At the call site, with a number chosen by a human. Default framework timeouts are often longer than your users will wait, which means a slow dependency becomes a slow product.
Pick three budgets
- Connect timeout — how long to establish the socket (often 100–500ms).
- Request timeout — total time for one attempt (often 1–3s for interactive APIs).
- Deadline propagation — the remaining budget from the caller, so the last service does not start work that will be abandoned.
If a user-facing request has a 2s budget and you call three services in series, each service must know it has less than the whole 2s. Pass deadlines downstream.
Retries without thundering herds
Retries help with transient faults and multiply load during partial outages. Treat retries as a budget, not a habit.
async function withRetry(fn, { retries = 2, baseMs = 80 } = {}) {
for (let attempt = 0; attempt <= retries; attempt++) {
try {
return await fn();
} catch (err) {
if (attempt === retries || !isRetryable(err)) throw err;
const jitter = 0.5 + Math.random();
await sleep(baseMs * 2 ** attempt * jitter);
}
}
}
Rules that hold up in production:
- Retry only idempotent operations by default.
- Cap retries at a small number (1–3) and always add jitter.
- Do not retry past the request deadline.
- Prefer a retry budget (e.g. 10% of traffic) over unbounded client retries.
Bulkheads and circuit breakers
Bulkheads isolate pools of concurrency so one slow dependency cannot consume every worker. Circuit breakers stop calling a dependency that is already failing, failing fast instead of piling latency on users.
A circuit breaker without good metrics is just a random outage. Measure success rates, latency percentiles, and open/close transitions or you will debug the breaker instead of the dependency.
Start with the simplest isolation that matches your failure mode: separate connection pools per dependency, or a dedicated worker pool for expensive calls. Full-blown adaptive concurrency control is rarely the first win.
Design observable degradation
For each dependency, write down the product behavior when it is slow or down. “Show cached data with a stale banner” is a design. “500 Internal Server Error” is an accident.
A small degradation menu
- Serve stale cache and mark freshness.
- Disable non-critical features; keep the core path.
- Queue writes and accept read-only mode.
- Return a partial response with explicit gaps.
Make the chosen mode visible in logs and metrics. If you cannot tell which degradation path ran, you cannot tune it.
Key takeaways
- Set explicit timeouts and propagate deadlines; never rely on defaults alone.
- Retry sparingly, idempotently, with jitter and a hard budget.
- Isolate dependencies so one failure does not become a platform failure.
- Define product-visible degradation paths before the incident, not during it.
- Instrument the failure paths; they are the ones you will be judged on.