Performance Best Practices
Performance rules that prevent premature optimization and focus effort on measured bottlenecks.
Search across all documentation pages
Performance rules that prevent premature optimization and focus effort on measured bottlenecks.
asyncio.to_thread for legacy blocking code.Python is fast enough for most web and data workloads. Optimize hotspots, not the language.
Only if benchmarks on your specific workload show clear wins.
Readability first. Profile if hot. Differences are usually negligible.
When a profiled hotspot cannot be fixed with NumPy, better algorithms, or caching.
Benchmark critical paths in CI with pytest-benchmark. Alert on threshold breach.
No. Only I/O-bound services that benefit from concurrency. Adds complexity.
Indexes, query plans, and N+1 elimination often beat application-level tuning.
For millions of small objects with fixed attributes. Not by default.
Define p95 latency per endpoint. Block deploys that exceed budget in load tests.
Only in profiled hot loops. f-strings are fine for normal code.
Stack versions: This page was written for Python 3.14.0, FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026