Retries & Backoff
Transient network and infrastructure failures often succeed on retry. Use bounded attempts, exponential backoff with jitter, and retry only idempotent operations - libraries like tenacity encode the pattern cleanly.
Search across all documentation pages
Transient network and infrastructure failures often succeed on retry. Use bounded attempts, exponential backoff with jitter, and retry only idempotent operations - libraries like tenacity encode the pattern cleanly.
import random
import time
def retry(callable_fn, attempts=3, base=0.5):
for attempt in range(1, attempts + 1):
try:
return callable_fn()
except TransientError:
if attempt == attempts:
raise
delay = base * (2 ** (attempt - 1)) + random.uniform(0, 0.1)
time.sleep(delay)When to reach for this:
Manual retry with jitter, then equivalent tenacity-style structure (stdlib-only demo).
import random
import time
from dataclasses import dataclass
class TransientError(Exception):
pass
@dataclass
class FakeClient:
calls: int = 0
fail_until: int = 2
def get(self) -> str:
self.calls += 1
if self.calls <= self.fail_until:
raise TransientError(f"attempt {self.calls}")
return "ok"
def with_retry(client: FakeClient, attempts: int = 5) -> str:
for attempt in range(1, attempts + 1):
try:
return client.get()
except TransientError as exc:
if attempt == attempts:
raise RuntimeError("exhausted retries") from exc
delay = min(2.0, 0.2 * (2 ** (attempt - 1))) + random.uniform(0, 0.05)
time.sleep(delay)
raise RuntimeError("unreachable")
client = FakeClient()
print(with_retry(client), "calls:", client.calls)What this demonstrates:
delay = min(cap, base * 2**attempt) + jitterstop_after_attempt, wait_exponential, retry_if_exception_type| Signal | Retry? |
|---|---|
| Timeout, 503 | Yes with backoff |
| 400, 422 validation | No |
| 409 conflict | Only with idempotent merge |
| Success | Stop |
stop condition burns CPU and masks outages. Fix: max attempts + circuit breaker.Retry-After header when present.| Alternative | Use When | Don't Use When |
|---|---|---|
| Circuit breaker | Sustained outage | Single blip |
| Queue + dead letter | Async workers | Sync request path |
| Client library built-in retry | boto/urllib3 configured | Need custom logging/metrics |
When decorators and composable stop/wait policies beat hand-rolled loops - most production services.
3-5 for HTTP with exponential backoff is common; tune with SLOs and p99 latency budgets.
Yes - include attempt, sleep, exception type, and correlation id for incident correlation.
Only for known transient errors (deadlock, connection lost) on idempotent statements.
Use asyncio.sleep instead of time.sleep; tenacity supports async callables.
Centralize policy for outbound HTTP clients; do not double-retry at multiple layers.
tenacity retry_if_result works when "soft failures" return error codes instead of raises.
Inject clock/sleep mock; assert attempt count without real delays (tenacity wait_fixed(0) in tests).
No - application or HTTP client (httpx transport) owns retry policy.
sleep = random.uniform(0, cap) - AWS recommended variant to spread load on recovery.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026