Async & Concurrency for Throughput
Python concurrency models target different bottlenecks: asyncio for I/O parallelism, multiprocessing for CPU parallelism, and threading for I/O with blocking libraries.
Search across all documentation pages
Python concurrency models target different bottlenecks: asyncio for I/O parallelism, multiprocessing for CPU parallelism, and threading for I/O with blocking libraries.
import asyncio
import httpx
async def fetch_all(urls: list[str]) -> list[int]:
async with httpx.AsyncClient() as client:
responses = await asyncio.gather(*[client.get(u) for u in urls])
return [r.status_code for r in responses]When to reach for this:
import asyncio
import httpx
from concurrent.futures import ProcessPoolExecutor
# I/O-bound: asyncio
async def fetch_urls(urls: list[str]) -> list[str]:
async with httpx.AsyncClient(timeout=10.0) as client:
tasks = [client.get(url) for url in urls]
responses = await asyncio.gather(*tasks, return_exceptions=True)
return [r.text[:100] if isinstance(r, httpx.Response) else str(r) for r in responses]
# CPU-bound: process pool
def compute(n: int) -> int:
return sum(i * i for i in range(n))
async def run_cpu_bound(items: list[int]) -> list[int]:
loop = asyncio.get_running_loop()
with ProcessPoolExecutor() as pool:
return list(await asyncio.gather(*[
loop.run_in_executor(pool, compute, n) for n in items
]))
async def main():
urls = ["https://httpbin.org/get"] * 20
io_results = await fetch_urls(urls)
cpu_results = await run_cpu_bound([100_000] * 4)
print(len(io_results), sum(cpu_results))
asyncio.run(main())What this demonstrates:
asyncio.gather runs I/O concurrently on one threadProcessPoolExecutor bypasses GIL for CPU workreturn_exceptions=True prevents one failure from canceling all| Bottleneck | Model | Tool |
|---|---|---|
| I/O (network, disk) | asyncio | httpx, asyncpg |
| CPU | multiprocessing | ProcessPoolExecutor |
| Blocking I/O library | threading | ThreadPoolExecutor |
| Mixed | asyncio + executor | run_in_executor |
asyncio.Semaphore(50) to cap parallel requests.await asyncio.to_thread(blocking_fn) or async library.| Alternative | Use When | Don't Use When |
|---|---|---|
| Sync + threading | Legacy blocking code | Greenfield async |
| Celery/RQ | Distributed task queue | In-process parallelism enough |
| Horizontal scaling | Production throughput | Local optimization phase |
asyncio for native async libraries. Threads for blocking libraries you cannot replace.
Start 10-50. Tune with semaphore. Watch server and client limits.
No. asyncio multiplexes I/O on one thread. CPU parallelism needs multiprocessing.
FastAPI is async-native. Use async database drivers for full benefit.
asyncio.Semaphore or token bucket with asyncio.sleep.
Typically os.cpu_count(). I/O pools can be larger than CPU count.
Python 3.14: prefer asyncio.TaskGroup for structured concurrency and error handling.
Requests per second under load (locust, hey, k6). Not single-request timeit.
Audit shared mutable state. Use locks or immutable data structures.
Low traffic, few concurrent I/O operations, or batch jobs without latency requirements.
Stack versions: This page was written for Python 3.14.0, FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026