Data & Model Versioning (DVC)
DVC versions large datasets and model artifacts alongside git code. Pipeline stages create reproducible DAGs that rerun only when inputs change.
Search across all documentation pages
DVC versions large datasets and model artifacts alongside git code. Pipeline stages create reproducible DAGs that rerun only when inputs change.
dvc init
dvc add data/train.csv
git add data/train.csv.dvc .gitignore
dvc push # upload to remote storage
dvc pull # download on another machine# dvc.yaml
stages:
prepare:
cmd: python src/prepare.py data/raw data/prepared
deps:
- src/prepare.py
- data/raw
outs:
- data/prepared/train.csv
train:
cmd: python src/train.py data/prepared/train.csv models/model.joblib
deps:
- src/train.py
- data/prepared/train.csv
outs:
- models/model.joblib
metrics:
- metrics.jsondvc repro train # run pipeline, skip unchanged stages
dvc metrics show # compare metrics across git branchesdvc add + remote storage.dvc remote add -d storage s3://bucket/dvc.dvc.yaml.| Alternative | Use When | Don't Use When |
|---|---|---|
| DVC | Git-centric ML teams | Non-git workflows |
| LakeFS | Data lake versioning | Simple CSV projects |
| W&B Artifacts | W&B ecosystem | Self-hosted data control |
| Manual S3 versioning | Minimal tooling | Need pipeline DAGs |
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026