Production bugs differ from local failures: traffic shape, data volume, and environment config hide issues your laptop never sees. Structured logs, distributed traces, and safe live profiling close that gap without risky restarts.
# Enable faulthandler on startup for native hangsimport faulthandlerfaulthandler.enable()# Temporary debug in one pod only (never commit)import logginglogging.getLogger("sqlalchemy.engine").setLevel(logging.INFO)
Restarting pods before capturing logs - You lose in-memory state and stack evidence. Fix: Snapshot logs, py-spy dump, and heap metrics first.
Missing correlation IDs - Thousands of identical errors are unsearchable. Fix: Require request_id/trace_id in middleware before handlers run.
Debugging with DEBUG level globally - Log volume explodes and PII leaks. Fix: Raise level on one pod or use dynamic sampling.
Assuming local == prod - Different PYTHONHASHSEED, locale, or dependency pins change behavior. Fix: Match image digest and env in a scratch environment.
Blocking the event loop in async apps - Sync ORM calls stall all requests; CPU looks low. Fix: Check span timelines for long synchronous segments; move I/O off the loop.
Use log platforms, APM, and kubectl exec/ecs exec only on a single canary pod. Prefer py-spy and env-print scripts checked into the repo over ad hoc edits.
Should I enable Python -X dev in production?
No. Dev mode adds expensive checks and warnings. Use staging with -X dev and keep prod on tuned logging and health probes.
What is the first query in an incident?
Filter errors by service and git_sha for the last two hours, then pivot on trace_id from the slowest successful request for contrast.
How do I trace a Celery task?
Pass trace_id in task kwargs and bind structlog contextvars in the task base class. Link worker logs to API logs via the same ID.
When is py-spy not enough?
When the process is blocked in native code without Python frames, or during startup before workers bind. Pair with eBPF or APM native profiling.
Can I use pprint in prod?
Only behind a feature flag and rate limit. Prefer structured fields so dashboards can aggregate.
How do I compare staging and prod configs?
Export redacted env keys (names only) and diff pool sizes, timeout values, and feature flags. Never paste secrets into tickets.
What about free-threaded Python 3.14?
GIL-related stalls diminish but I/O and lock contention remain. py-spy still helps; watch threading.Lock hot spots in metrics.
How long should I keep debug logging on?
Minutes to one pod. Revert immediately after capture; document findings in the incident channel.
Who should own prod debugging?
On-call engineer drives evidence; service owner interprets domain logic. Do not solo-debug SEV1 without an incident commander.