Datasets & Splits
Train, validation, and test splits define what your model has seen during training versus what it must generalize to. Correct splitting is the foundation of honest evaluation.
Search across all documentation pages
Train, validation, and test splits define what your model has seen during training versus what it must generalize to. Correct splitting is the foundation of honest evaluation.
Quick-reference recipe card - copy-paste ready.
from sklearn.model_selection import train_test_split, StratifiedKFold, cross_val_score
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
scores = cross_val_score(model, X_train, y_train, cv=cv, scoring="f1_macro")When to reach for this:
"""datasets_splits.py - leak-free split workflow for tabular classification."""
from __future__ import annotations
import pandas as pd
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import classification_report
from sklearn.model_selection import StratifiedKFold, cross_val_score, train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
# Load and inspect
data = load_breast_cancer(as_frame=True)
X: pd.DataFrame = data.data
y: pd.Series = data.target
# 1) Hold out final test set (never used during model selection)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
# 2) Build pipeline - fit only on training folds inside CV
pipe = Pipeline([
("scaler", StandardScaler()),
("clf", RandomForestClassifier(n_estimators=100, random_state=42)),
])
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
cv_scores = cross_val_score(pipe, X_train, y_train, cv=cv, scoring="f1")
print(f"CV F1: {cv_scores.mean():.3f} +/- {cv_scores.std():.3f}")
# 3) Final fit on all training data, evaluate once on test
pipe.fit(X_train, y_train)
print(classification_report(y_test, pipe.predict(X_test)))What this demonstrates:
| Strategy | Use When | sklearn API |
|---|---|---|
| Random split | IID rows, balanced classes | train_test_split |
| Stratified split | Classification with rare classes | stratify=y |
| Time-based split | Forecasting, temporal drift | TimeSeriesSplit |
| Group split | Multiple rows per entity | GroupKFold, GroupShuffleSplit |
| Nested CV | Unbiased hyperparameter tuning | GridSearchCV inside outer CV |
from sklearn.model_selection import TimeSeriesSplit, GroupKFold
# Time series: always train on past, test on future
tscv = TimeSeriesSplit(n_splits=5)
# Groups: same group_id never in both train and test
gkf = GroupKFold(n_splits=5)
for train_idx, val_idx in gkf.split(X, y, groups=group_ids):
...Pipeline.stratify=y or resampling strategies.TimeSeriesSplit or cutoff-based splits.group_id, not by row.
| Alternative | Use When | Don't Use When |
|---|---|---|
| Single train/test split | Large datasets, quick iteration | Small data where split variance is high |
| K-fold cross-validation | Model comparison, limited data | Very large data where one split suffices |
| Holdout + validation set | Deep learning with long training runs | You need many hyperparameter trials (use CV) |
| Bootstrap / repeated CV | Unstable metrics on small data | Training cost per fold is prohibitive |
train_test_split(X, y, stratify=y) # preserves class ratios in both partitionsTimeSeriesSplit.groups array with one ID per row (user_id, patient_id).GroupKFold ensures all rows for a group land in the same fold.random_state=42 (any fixed int) makes splits reproducible.train, test = train_test_split(df, test_size=0.2, stratify=df["label"])
X_train, y_train = train["features"], train["label"]cross_val_predict instead, still only on training data.MultiOutputClassifier or separate models per target.Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026