Resilience: Timeouts, Retries & Circuit Breakers Explained
How to keep a distributed system standing when one service fails. Timeouts, retries with backoff and jitter, the circuit breaker (closed/open/half-open), plus fallbacks and bulkheads — stopping cascading failures.
🟠Fundamentals · Senior → 🔴 super-senior deep dive
In a distributed system, things fail constantly — a service is slow, a network blips, a database is briefly overloaded. The senior question isn't "how do I prevent failures?" (you can't) but "how do I keep the whole system standing when one part wobbles?" That's resilience, and it comes down to a handful of patterns — timeouts, retries, and circuit breakers — that every serious backend uses. Get these right and one failing dependency doesn't take down everything. Let's build them up.
Why failure spreads (the cascade)
🟠Here's the nightmare a resilient design prevents. Service A calls Service B, which is overloaded and slow. A's threads all block waiting on B. Soon A runs out of threads and can't serve anyone — so A goes down too. Then whoever calls A blocks and fails. One slow service triggers a cascading failure that ripples across the system. The patterns below exist to stop that ripple at the first hop.
1. Timeouts — never wait forever
The most basic and most important rule: every network call must have a timeout. If Service B doesn't answer within, say, 2 seconds, give up and move on. Without a timeout, a hung dependency ties up your resources indefinitely — the exact starting point of a cascade. A timeout turns "wait forever" into "fail fast," freeing your threads to keep serving other work.
2. Retries — but carefully
🔴 Many failures are transient — a momentary blip that succeeds if you just try again. So on failure, retry. But naive retries are dangerous, and two refinements separate juniors from seniors:
- Exponential backoff: don't retry instantly and repeatedly — that adds load to an already-struggling service. Wait a bit more each time: 1s, then 2s, then 4s. Give it room to recover.
- Jitter (randomness): if 1,000 clients all fail at once and all back off by exactly 2s, they retry together — a synchronized stampede (the "thundering herd") that re-crushes the service. Add a small random delay so retries spread out.
- Only retry idempotent operations: retrying a "charge card" can double-charge. Retry safely only when repeating the operation is harmless (see idempotency), and cap the number of attempts.
Retries with backoff + jitter on idempotent calls: that's the phrase interviewers want to hear.
3. Circuit breaker — stop kicking a dead service
🔴 If a service is clearly down, retrying is pointless and harmful — you waste time and pile load on it. A circuit breaker (named after the electrical kind) watches the failure rate and, once it crosses a threshold, "trips": it stops calling the failing service entirely for a while and fails fast instead. It has three states:
- Closed: normal — requests pass through, failures are counted.
- Open: too many failures → trip. Reject calls immediately (fail fast) without touching the sick service, giving it time to recover.
- Half-open: after a cooldown, let one test request through. If it succeeds, close the circuit (recovered); if it fails, open again.
The circuit breaker is what stops the cascade: instead of thousands of threads piling onto a dead service, callers fail instantly and stay healthy.
Two more patterns worth naming
🔴 A couple of bonus patterns that round out a senior answer:
- Fallback / graceful degradation: when a call fails or the breaker is open, return something useful instead of an error — cached data, a default, or a reduced feature. "Recommendations service is down? Show generic popular items." The user barely notices.
- Bulkhead: like watertight compartments in a ship, isolate resources (e.g. separate thread pools) per dependency, so one flooded compartment can't sink the whole vessel. A failure in one integration can't consume all your threads.
Putting it together
A resilient call looks like this: set a timeout so you never hang; on transient failure, retry with exponential backoff + jitter (only if idempotent); wrap it in a circuit breaker so a truly-down service is skipped; and provide a fallback so users get a graceful experience instead of an error. Layer those and one failing dependency becomes a minor, contained blip rather than a site-wide outage.
The interview-ready summary
"Failures are constant in distributed systems, so I design to contain them. Every call gets a timeout to fail fast. Transient errors get retries with exponential backoff and jitter, but only on idempotent operations. A circuit breaker stops calling a service that's clearly down — closed, open, half-open — so I don't pile load on it and don't block my own threads. And I add a fallback for graceful degradation. Together these stop a single slow dependency from cascading into a full outage." That's a textbook senior answer."
What to read next
- Microservices Architecture — where these patterns become essential
- Distributed Transactions & Saga — idempotency, which safe retries depend on
- Rate Limiting — protecting a service from being overwhelmed in the first place
- ← The complete System Design guide (hub)
← Object storage & blobs · Latency numbers every engineer should know →