Incidents without guardrails repeat. Guardrails are alerts, tests, lint rules, and checklists that make the same failure harder or visible earlier. Python teams encode them in pytest, Ruff, Alembic review, and SLO burn alerts.
# pytest regression from post-mortem INC-8842def test_migration_uses_concurrently_for_large_tables(): sql = Path("alembic/versions/20260612_add_invoice_idx.py").read_text() assert "CONCURRENTLY" in sql or "op.create_index" not in sql
# tests/test_incident_9011_loop_block.pyimport astfrom pathlib import Pathdef test_no_time_sleep_in_async_functions(): src = Path("src/api/routes/checkout.py").read_text() tree = ast.parse(src) for node in ast.walk(tree): if isinstance(node, ast.AsyncFunctionDef): for child in ast.walk(node): if isinstance(child, ast.Call): func = child.func if isinstance(func, ast.Attribute) and func.attr == "sleep": raise AssertionError(f"time.sleep in async {node.name}")
What this demonstrates:
Guardrails link to incident IDs for traceability
Regression tests encode specific failure mechanics
CI scripts catch async blocking patterns
Ops review audits verification commands, not just ticket closure