Timeouts, retries and circuit breakers
Your dependencies will sometimes be slow or down. Resilience patterns decide whether that causes a small, contained error or takes your whole service with it.
Every network call needs a timeout
Without a timeout, a call to a hung dependency waits for a very long time, sometimes forever. Meanwhile it holds a thread, a connection and memory. As requests pile up, your service runs out of all three and stops answering, even for requests that never touch the broken dependency.
Many HTTP clients and database drivers have no timeout, or a very long one, by default. Set one explicitly on every call.
const res = await fetch(url, {
signal: AbortSignal.timeout(2000),
});
Base the value on the dependency’s normal latency, for example a little above its p99, not on a round number. Also set an overall deadline for the incoming request, and pass the remaining time down, so inner calls don’t keep working after the caller has given up.
Retry only what is safe to repeat
Many failures are brief: a dropped connection, a restarting instance. A retry often succeeds. But a retry sends the request again, so only retry operations that are safe to repeat: reads, idempotent writes, or POST requests protected by an idempotency key.
Retry errors that may be transient, such as timeouts, connection resets, 503 and 429. Don’t retry 400, 401, 403 or 404; the same request will fail the same way.
Exponential backoff with jitter
Retrying immediately hits a struggling service again at the worst moment. Use exponential backoff: double the wait after each attempt, up to a cap. Add jitter, a random component, so that many clients that failed together don’t retry together.
import random
def backoff(attempt, base=0.1, cap=5.0):
# "full jitter": random wait up to the exponential bound
return random.uniform(0, min(cap, base * 2 ** attempt))
Limit the number of attempts, and respect a Retry-After header if the server sends one.
Retry storms
Retries multiply. If each of three layers retries three times, one user request can become 27 calls to the bottom service, arriving exactly when it is overloaded. This is a retry storm, and it can keep a service down long after the original problem is gone.
Retry at one layer only, usually the one closest to the failing call. Some teams also use a retry budget: retries may add at most a small percentage on top of normal traffic, and beyond that they are dropped.
Circuit breakers
If a dependency is clearly down, waiting for every call to time out wastes time and resources. A circuit breaker tracks recent failures and, once they pass a threshold, opens: calls fail immediately without being sent. After a cool-down it lets a few trial calls through. If they succeed, it closes again; if not, it stays open.
| State | Behaviour |
|---|---|
| Closed | calls pass through; failures are counted |
| Open | calls fail fast without being sent |
| Half-open | a few trial calls decide whether to close or reopen |
Bulkheads
A bulkhead isolates resources so one dependency can’t consume all of them. Give each dependency its own connection pool or concurrency limit. If the recommendations service hangs, it fills its own small pool, and checkout keeps working with its separate one.
Graceful degradation and fallbacks
When a dependency fails, decide what the user should get instead of an error page. Options include a cached or default value, a reduced feature (the product page without recommendations), or queuing the work for later. Fallbacks must be simpler and more reliable than what they replace; a fallback that calls another fragile service just moves the problem. Make sure you can see in your metrics when you are running degraded.
Checklist for each dependency
- It has a timeout, and the request has an overall deadline.
- Retries are limited, use backoff with jitter, and apply only to safe operations.
- Only one layer retries.
- A circuit breaker and a bounded pool protect the caller.
- There is a defined fallback, and you have tested the failure path.