Gunicorn & Uvicorn
Gunicorn manages worker processes for WSGI and ASGI apps; Uvicorn implements the ASGI server. Production FastAPI on 0.115+ commonly uses Gunicorn with UvicornWorker for multi-process concurrency plus async I/O.
Search across all documentation pages
Gunicorn manages worker processes for WSGI and ASGI apps; Uvicorn implements the ASGI server. Production FastAPI on 0.115+ commonly uses Gunicorn with UvicornWorker for multi-process concurrency plus async I/O.
gunicorn app.main:app \
-w 4 \
-k uvicorn.workers.UvicornWorker \
-b 0.0.0.0:8000 \
--timeout 60 \
--graceful-timeout 30 \
--access-logfile -When to reach for this:
-k sync default)Tuned Gunicorn config file for FastAPI with logging and worker count from env.
# gunicorn.conf.py
import multiprocessing
import os
bind = f"0.0.0.0:{os.environ.get('PORT', '8000')}"
worker_class = "uvicorn.workers.UvicornWorker"
workers = int(os.environ.get("WEB_CONCURRENCY", multiprocessing.cpu_count()))
timeout = 60
graceful_timeout = 30
keepalive = 5
accesslog = "-"
errorlog = "-"
loglevel = os.environ.get("LOG_LEVEL", "info")gunicorn -c gunicorn.conf.py app.main:appWhat this demonstrates:
WEB_CONCURRENCY pattern matches Heroku/Render conventionsUvicornWorker runs ASGI per worker process| Stack | Worker class |
|---|---|
| Flask/Django sync | sync (default) |
| FastAPI async | uvicorn.workers.UvicornWorker |
| gevent (legacy) | gevent |
2-4 x CPU cores for sync CPU-bound; fewer for memory-heavy ML# Dev only - single worker reload
uvicorn app.main:app --reload --host 127.0.0.1 --port 8000--graceful-timeout + preStop hook delay in K8s.| Alternative | Use When | Don't Use When |
|---|---|---|
| Hypercorn | HTTP/2 ASGI needs | Team standard is Gunicorn |
| uWSGI | Legacy Django shops | Greenfield FastAPI |
| Multiple uvicorn replicas | K8s one-process-per-pod model | Want in-container worker pool |
Gunicorn supervises multiple Uvicorn workers; single uvicorn fine for dev and tiny prod with external scaling.
Start at CPU count for I/O-bound async; half CPU for CPU-bound sync; measure latency and RSS under load.
Django 5.2 async views can; most Django deploys remain sync Gunicorn workers unless async path is explicit.
UvicornWorker supports websockets; sync Gunicorn workers do not.
Prometheus multiprocess dir when using multiple sync workers; ASGI middleware for request metrics.
Deploy new container image - do not rely on HUP reload in immutable infra.
gthread for some WSGI I/O overlap - less common than async ASGI for new FastAPI services.
[::]:8000 when platform requires dual-stack - verify load balancer config.
At load balancer or ingress - not usually on Gunicorn inside container.
One container runs Gunicorn with N workers; HPA scales pods on CPU/RPS - see Kubernetes page.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026