Performance Basics
11 examples to get you started with Performance - 8 basic and 3 intermediate.
Search across all documentation pages
11 examples to get you started with Performance - 8 basic and 3 intermediate.
timeit, cProfile, tracemallocuv add numpy for vectorization examplesNever optimize without a baseline measurement.
import time
def slow_sum(n: int) -> int:
return sum(range(n))
start = time.perf_counter()
result = slow_sum(1_000_000)
elapsed = time.perf_counter() - start
print(f"Result: {result}, Time: {elapsed:.4f}s")time.perf_counter() is monotonic and high resolutionRelated: timeit & Microbenchmarks - precise timing
The timeit module runs code in tight loops for stable timing.
import timeit
elapsed = timeit.timeit("sum(range(1000))", number=10000)
print(f"Average: {elapsed / 10000 * 1e6:.2f} μs")number controls loop iterationsRelated: timeit & Microbenchmarks - methodology
Find which functions consume the most CPU time.
import cProfile
import pstats
def work():
sum(i * i for i in range(100_000))
cProfile.run("work()", "profile.out")
stats = pstats.Stats("profile.out")
stats.sort_stats("cumulative").print_stats(10)cumulative sorts by total time including subcallsRelated: cProfile & Profilers - flame graphs
Track memory allocations to find leaks and bloat.
import tracemalloc
tracemalloc.start()
data = [bytearray(1024) for _ in range(10_000)]
snapshot = tracemalloc.take_snapshot()
top = snapshot.statistics("lineno")[:5]
for stat in top:
print(stat)
tracemalloc.stop()tracemalloc is stdlib - no install neededRelated: Memory Profiling - memray deep dive
Comprehensions are faster for simple transforms.
# Faster
squares = [x * x for x in range(100_000)]
# Slower
squares = []
for x in range(100_000):
squares.append(x * x)list.append lookupsCache pure function results to avoid redundant computation.
from functools import lru_cache
@lru_cache(maxsize=256)
def fibonacci(n: int) -> int:
if n < 2:
return n
return fibonacci(n - 1) + fibonacci(n - 2)
print(fibonacci(100)) # instant with cachemaxsize bounds memory usage.cache_info() to inspect hit rateRelated: Caching & Memoization - cache strategies
Generators yield items lazily, saving memory on large datasets.
def read_lines(path: str):
with open(path) as f:
for line in f:
yield line.strip()
# Processes one line at a time - constant memory
for line in read_lines("huge.log"):
if "ERROR" in line:
print(line)sum(1 for _ in gen) still consumes the generator onceBind frequently accessed attributes to local variables in hot loops.
# Faster in tight loops
append = results.append
for item in data:
append(transform(item))
# Slower - repeated attribute lookup
for item in data:
results.append(transform(item))Replace Python loops with array operations for numeric work.
import numpy as np
data = np.arange(1_000_000, dtype=np.float64)
result = np.sqrt(data) * 2.0 + 1.0 # vectorized
# vs: [math.sqrt(x) * 2.0 + 1.0 for x in data] # 10-100x slowerRelated: Vectorization with NumPy - full guide
Parallelize waiting on network/disk without threads.
import asyncio
import httpx
async def fetch_all(urls: list[str]) -> list[int]:
async with httpx.AsyncClient() as client:
tasks = [client.get(url) for url in urls]
responses = await asyncio.gather(*tasks)
return [r.status_code for r in responses]
asyncio.run(fetch_all(["https://example.com"] * 10))asyncio.gather for concurrent coroutinesmultiprocessing or concurrent.futuresRelated: Async & Concurrency for Throughput - patterns
A systematic pass before claiming something is "optimized."
# 1. Measure baseline
# 2. Profile (cProfile or py-spy)
# 3. Fix the top hotspot only
# 4. Re-measure
# 5. Document the before/after numbers
baseline_ms = 450
optimized_ms = 120
improvement = (baseline_ms - optimized_ms) / baseline_ms * 100
print(f"Improvement: {improvement:.0f}%") # 73%Related: Performance Audit Checklist - full checklist
Stack versions: This page was written for Python 3.14.0, FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026