Rollback Strategies
Not every bad deploy can roll back. Delivery teams choose between reverting the previous image, forward-fixing with a patch, or killing a feature flag - based on migration compatibility and time to mitigate.
Search across all documentation pages
Not every bad deploy can roll back. Delivery teams choose between reverting the previous image, forward-fixing with a patch, or killing a feature flag - based on migration compatibility and time to mitigate.
Quick-reference recipe card - copy-paste ready.
## Rollback decision (2 minutes)
1. Customer impact active? → mitigate first
2. Deploy < 2h and image sha-prev exists? → candidate rollback
3. Migration destructive or contract? → forward-fix only
4. Bug isolated behind flag? → flag off, then plan fix
5. Record decision + commands in #incident threadWhen to reach for this:
"""rollback_decision.py - encode team policy as testable helpers."""
from __future__ import annotations
from dataclasses import dataclass
from enum import Enum
class Mitigation(str, Enum):
FLAG_OFF = "flag_off"
ROLLBACK_IMAGE = "rollback_image"
FORWARD_FIX = "forward_fix"
@dataclass(frozen=True)
class DeployState:
minutes_since_deploy: int
migration_contract: bool # removed column or destructive
migration_expand: bool
flag_isolates_bug: bool
previous_image_available: bool
def choose_mitigation(state: DeployState) -> Mitigation:
if state.flag_isolates_bug:
return Mitigation.FLAG_OFF
if state.migration_contract:
return Mitigation.FORWARD_FIX
if state.previous_image_available and state.minutes_since_deploy < 120:
return Mitigation.ROLLBACK_IMAGE
return Mitigation.FORWARD_FIX
if __name__ == "__main__":
s = DeployState(25, migration_contract=False, migration_expand=True, flag_isolates_bug=False, previous_image_available=True)
print(choose_mitigation(s))What this demonstrates:
rollout undo.| Situation | Preferred strategy |
|---|---|
| Bad logic, expand migration | Flag off or image rollback |
| Bad index lock | Forward-fix migration + maybe rollback |
| Dropped column in prod | Forward-fix only |
| Config typo | Revert config + restart |
| Worker incompatible | Pause workers, rollback API |
# Keep previous Gunicorn image handy
docker pull $REGISTRY/orders-api:sha-prev
kubectl set image deploy/orders-api api=$REGISTRY/orders-api:sha-prev| Alternative | Use When | Don't Use When |
|---|---|---|
| Auto rollback on canary fail | Mature metrics | Schema not compatible |
| Blue-green swap | Pre-warmed previous stack | Cost-sensitive tiny fleet |
| Traffic drain | Long-running requests | Need instant fix |
| Disable feature at LB | Path-based routing | Bug in shared middleware |
Target under 10 minutes from decision; practice in quarterly game day.
Rollback is promoted previous artifact; revert commit is next forward deploy - slower.
Destructive migrations, data migrations already applied, or rollback image unavailable.
Shift alias weight to previous version; same decision tree on schema.
Deploy to ephemeral env, run smoke, rollout undo, rerun smoke - nightly optional.
Application rollback does not undo DDL; use expand-contract and PITR for disasters.
IC during incident; release captain pre-prod. Document in thread.
Revert model artifact version in registry; may need shadow validation first.
Track change failure rate and MTTR; fast rollback improves both.
Image embeds lock hash; rollback restores previous dependencies automatically.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026