Git for Data/Notebooks
Jupyter notebooks, parquet exports, and model checkpoints need Git policies different from application code - nbstripout, Git LFS, and data versioning tools keep PRs reviewable and repos small.
Search across all documentation pages
Jupyter notebooks, parquet exports, and model checkpoints need Git policies different from application code - nbstripout, Git LFS, and data versioning tools keep PRs reviewable and repos small.
pip install nbstripout
nbstripout --install --attributes
git add .gitattributes notebooks/analysis.ipynb
git commit -m "chore: strip notebook outputs on commit"When to reach for this:
.ipynb exploration# .pre-commit-config.yaml excerpt
repos:
- repo: https://github.com/kynan/nbstripout
rev: 0.7.1
hooks:
- id: nbstripout
# Track large artifacts with LFS
git lfs install
git lfs track "*.parquet"
git lfs track "models/*.pt"
git add .gitattributes
git commit -m "chore: LFS track parquet and model weights"# Prefer referencing data by URI in notebooks
DATA_URI = "s3://acme-analytics/processed/orders/dt=2026-07-08/"
# not: pd.read_csv("data/orders_10gb.csv") # committed to gitWhat this demonstrates:
.gitattributes documents filter rulesoutputs and execution_count..dvc files in git.| Asset | Strategy |
|---|---|
| Notebooks | nbstripout + optional Jupytext .py |
| 10MB-5GB files | Git LFS with quota monitoring |
| Datasets | DVC, cloud partition paths, catalog |
| Secrets in notebooks | Never - use env + pydantic-settings |
# Jupytext paired notebook
jupytext --set-formats ipynb,py:percent analysis.ipynb
git add analysis.ipynb analysis.pyGIT_LFS_SKIP_SMUDGE=1 in CI that does not need data..dvc with code; dvc pull in CI docs.| Alternative | Use When | Don't Use When |
|---|---|---|
| Papermill parameters | Scheduled notebook runs | Library code belongs in .py modules |
| Marimo / percent scripts | Git-friendly source first | Team standardized on classic Jupyter |
| SageMaker / Colab only | No local git for notebooks | Org requires PR review on git |
| Quarto | Report publishing pipeline | Quick internal EDA only |
Automate with pre-commit - manual clear is forgotten on every commit.
LFS for few large binaries in repo; DVC for versioned dataset pipelines and repro experiments.
Use GitHub notebook diff or review paired .py from Jupytext for readable changes.
papermill or nbclient with pinned data fixtures - not live prod data.
Avoid - move logic to src/ modules; keep notebooks in notebooks/ or docs/.
Monitor storage and bandwidth on GitHub - migrate big data to object store when exceeded.
Re-run nbstripout on both sides or choose one author's notebook and re-apply cells manually.
Do not commit HTML/PDF exports - generate in CI artifacts.
Store run artifacts in W&B/S3; git tracks only training commit SHA and config YAML.
Scope nbstripout to notebooks/ path only in hook files regex.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026