Deep Learning Best Practices
Rules for reproducible, efficient PyTorch training and deployment. Apply during code review and before scaling to multi-GPU.
Search across all documentation pages
Rules for reproducible, efficient PyTorch training and deployment. Apply during code review and before scaling to multi-GPU.
torch.manual_seed(42), np.random.seed(42), and DataLoader generator for reproducible runs.torch.__version__ and CUDA driver in experiment logs.torch.use_deterministic_algorithms(True) for reproducibility (may reduce performance).clip_grad_norm_(model.parameters(), 1.0).torch.amp.autocast + GradScaler for 2x speedup.CrossEntropyLoss for multi-class; BCEWithLogitsLoss for multi-label.model.eval() before TorchScript trace or ONNX export.torch.allclose against ONNX Runtime output.model.eval() during validation.torch.manual_seed(42)
torch.cuda.manual_seed_all(42)
generator = torch.Generator().manual_seed(42)
loader = DataLoader(ds, generator=generator, shuffle=True)torch.cuda.empty_cache() between experiments.fast_dev_run equivalent (few batches) to validate the loop.nvidia-smi).num_workers and pin_memory.def test_model_output_shape():
out = model(torch.randn(1, 3, 224, 224))
assert out.shape == (1, num_classes)Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026