Observability Best Practices
Actionable signals, not noise, mean on-call wakes up for customer-impacting problems - not for log volume or unbounded metric cardinality.
Search across all documentation pages
Actionable signals, not noise, mean on-call wakes up for customer-impacting problems - not for log volume or unbounded metric cardinality.
X-Request-ID to clients.before_send./metrics endpoint. Internal network or auth - not public internet.country, plan).release to git SHA. Bisect regressions to deploy./health/live and /health/ready. Deps checked only on readiness.JSON logs, health endpoints, RED metrics, Sentry for exceptions, one dashboard, one on-call runbook.
Metrics for thresholds and SLOs; logs for investigation after alert fires - not vice versa.
Enough to debug weekly incidents - adjust sample rate vs observability bill.
Cron scripts: structured logs + exit code monitoring; metrics optional for duration histogram if critical.
Hash or omit; legal review for fields stored in Sentry and log retention period.
Service owners with platform template - consistency beats bespoke chart art.
One alert per symptom routed to one team; correlate Sentry + Prometheus with routing rules.
Pretty console logs locally; JSON in staging that mirrors prod pipeline.
De facto standard for new instrumentation - vendor agents secondary.
Logging everything at INFO with no structure - cannot query, cannot alert, high cost.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026