You Added Retries to Make Your API More Reliable… So Why Did the Outage Get Worse?
How immediate retries create retry storms and cascading failures, and how to use timeouts, exponential backoff, jitter, retry budgets, and circuit breakers safely.
7 min read
A Small Failure Became Three Times the Traffic
The inventory service starts timing out under load.
Your API is configured to retry every failed request twice:
Original traffic: 4,000 requests/sec
First retries: 4,000 requests/sec
Second retries: 4,000 requests/sec
Potential demand: 12,000 requests/secThe dependency was already struggling at 4,000 RPS. “Reliability” logic now asks it to handle up to 12,000.
Latency rises, more calls cross their timeout, and even successful slow operations get retried by impatient callers. The retry becomes part of the outage.
Why Immediate Retries Synchronize Failure
This implementation retries as fast as the process can loop:
async function fetchInventory(productId: string) {
for (let attempt = 0; attempt < 3; attempt++) {
try {
return await inventoryClient.get(productId);
} catch {
// Retry immediately.
}
}
throw new Error("Inventory unavailable");
}During an outage, thousands of callers fail at nearly the same time and retry at nearly the same time.
dependency slows
|
v
requests time out together
|
v
retries arrive together
|
v
dependency gets less recovery capacity
|
+---------- loop ----------+This synchronized wave is a retry storm.
First Decide Whether the Error Is Retryable
Usually retryable:
- Connection reset
- Temporary
503 Service Unavailable - Explicit
429 Too Many RequestswithRetry-After - A timeout where the operation is known to be idempotent
Usually not retryable without a change:
400validation failure401invalid credentials403authorization failure- Schema incompatibility
- A deterministic business rejection
- An unsafe write with unknown outcome
Classify the error. Retrying a permanent failure adds load without increasing the chance of success.
Exponential Backoff Creates Space
Backoff increases the delay after each failure:
attempt 1: wait about 100 ms
attempt 2: wait about 200 ms
attempt 3: wait about 400 ms
attempt 4: wait about 800 msBut identical backoff still synchronizes clients. Add jitter.
function fullJitterDelay(attempt: number, baseMs = 100, capMs = 5_000) {
const exponential = Math.min(capMs, baseMs * 2 ** attempt);
return Math.floor(Math.random() * exponential);
}
async function withRetry<T>(operation: () => Promise<T>): Promise<T> {
const maxAttempts = 3;
for (let attempt = 0; attempt < maxAttempts; attempt++) {
try {
return await operation();
} catch (error) {
if (!isTransient(error) || attempt === maxAttempts - 1) throw error;
await sleep(fullJitterDelay(attempt));
}
}
throw new Error("unreachable");
}Jitter spreads callers over time so the recovering service does not receive one coordinated wave.
Put Retries Inside a Total Time Budget
Suppose the incoming request has a 1-second service-level objective.
This cannot work:
attempt 1 timeout: 800 ms
attempt 2 timeout: 800 ms
attempt 3 timeout: 800 msThe retry policy requires 2.4 seconds before backoff and overhead.
Instead, propagate a deadline:
total budget: 1,000 ms
local processing: 100 ms
downstream budget: 700 ms
response reserve: 200 msEvery attempt consumes the same remaining budget. Do not start a retry that cannot finish before the caller’s deadline.
Retry at One Layer, Not Every Layer
Consider:
Browser retries 3 times
API gateway retries 3 times
Order service retries 3 times
Payment client retries 3 timesWorst-case attempts can multiply:
3 × 3 × 3 × 3 = 81 downstream attemptsChoose one layer that has enough context to retry safely. Disable overlapping library defaults unless they are intentionally part of the policy.
Writes Need Idempotency
The client times out while creating a payment. Did the server fail before charging, or did the response get lost after charging?
Retrying without an idempotency key can charge twice.
POST /payments
Idempotency-Key: checkout_81f2_paymentThe server stores the key and result:
CREATE TABLE idempotency_keys (
key text PRIMARY KEY,
request_hash text NOT NULL,
status text NOT NULL,
response jsonb,
expires_at timestamptz NOT NULL
);The same key and request return the original result. The same key with a different request must be rejected.
Circuit Breakers Stop Futile Calls
A circuit breaker watches recent outcomes.
CLOSED
calls flow normally
too many failures
|
v
OPEN
fail fast; no dependency call
after cooldown
|
v
HALF-OPEN
allow a few probes
success -> CLOSED
failure -> OPENFailing fast protects connection pools, threads, event-loop capacity, and the dependency itself.
A circuit breaker is not a substitute for timeouts. Without a timeout, calls may hang too long for the breaker to receive outcomes.
Use a Retry Budget
A retry budget limits retries relative to normal traffic.
For example, if the budget allows retries equal to at most 10% of original requests:
Original traffic: 5,000 RPS
Retry budget: 500 RPSOnce the budget is exhausted, fail fast or degrade. This prevents retries from becoming the dominant workload during a broad incident.
Respect Backpressure
When a server returns 429 or 503 with Retry-After, it is telling callers when capacity may be available.
Ignoring that signal and immediately retrying defeats load shedding.
For non-interactive work, a durable queue may be safer:
API accepts request
|
v
durable queue with delayed retry
|
v
worker processes within bounded concurrencyThat changes the contract from synchronous completion to asynchronous completion, so use it only when the product flow allows it.
Diagnose a Retry Storm
Correlate:
original_requests_total
retry_attempts_total by dependency and reason
timeout_total
circuit_state
dependency_latency
in_flight_requests
connection_pool_waitingA common timeline is:
10:00 dependency latency rises
10:01 timeouts increase
10:01 retry rate exceeds original request rate
10:02 connection pools saturate
10:03 unrelated endpoints slow downAdd an attempt field to traces and logs so retried requests are not mistaken for new user demand.
Trade-offs
| Control | Benefit | Cost |
|---|---|---|
| Backoff + jitter | Reduces synchronized load | Slower eventual success |
| Strict attempt limit | Bounds amplification | Some transient failures surface |
| Circuit breaker | Fast failure and recovery space | Requires careful thresholds |
| Retry budget | Protects global capacity | Rejects retries during incidents |
| Queue | Smooths asynchronous work | Adds delay and operational complexity |
| Idempotency key | Makes write retry safe | State, retention, and conflict handling |
Production Best Practices
- Retry only classified transient failures.
- Use exponential backoff with jitter.
- Keep attempt counts small.
- Enforce one end-to-end deadline.
- Retry at one deliberate layer.
- Require idempotency for retryable writes.
- Cap retries with a budget.
- Fail fast with a circuit breaker when recovery is unlikely.
- Respect
Retry-Afterand load-shedding signals. - Test dependency slowdown, not only total failure.
Conclusion
Retries improve reliability when failures are brief, operations are safe to repeat, and the system has spare recovery capacity.
They worsen outages when they are immediate, multiplied across layers, unbounded, or applied to permanent failures.
The safe model is:
classify
-> bound by deadline
-> back off with jitter
-> preserve idempotency
-> stop when capacity is goneThat same discipline is essential when applying backpressure across APIs, queues, and downstream dependencies, because retries are another source of producer traffic.
References
Related
Consistency Models in Distributed Systems: Strong, Eventual, and Read-Your-Writes
System Design27 min read
Caching Architecture: Cache-Aside, Write-Through, Write-Behind, and Redis Patterns
System Design34 min read
Your API Works Fine at 100 RPS… What Actually Breaks at 10,000 RPS?
System Design8 min read
Written by
Faisal
Software engineer writing about backend systems, Node.js, system design, scalable applications, and modern web and mobile development.