Pipelines
sklearn Pipeline chains preprocessing steps and estimators into a single object that fits, predicts, and serializes together. Pipelines are the standard way to keep training and inference reproducible and leak-free.
Search across all documentation pages
sklearn Pipeline chains preprocessing steps and estimators into a single object that fits, predicts, and serializes together. Pipelines are the standard way to keep training and inference reproducible and leak-free.
Quick-reference recipe card - copy-paste ready.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
pipe = Pipeline([
("scaler", StandardScaler()),
("clf", LogisticRegression(max_iter=1000)),
])
pipe.fit(X_train, y_train)
preds = pipe.predict(X_test)When to reach for this:
FeatureUnion."""pipelines.py - end-to-end pipeline with ColumnTransformer and grid search."""
from __future__ import annotations
import joblib
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import f1_score
from sklearn.model_selection import GridSearchCV, train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
df = pd.read_csv("customers.csv") # columns: age, tenure, region, plan, churned
X = df.drop(columns=["churned"])
y = df["churned"]
num_cols, cat_cols = ["age", "tenure"], ["region", "plan"]
preprocessor = ColumnTransformer([
("num", StandardScaler(), num_cols),
("cat", OneHotEncoder(handle_unknown="ignore"), cat_cols),
])
pipe = Pipeline([
("prep", preprocessor),
("clf", RandomForestClassifier(random_state=42)),
])
param_grid = {
"clf__n_estimators": [100, 200],
"clf__max_depth": [None, 10],
}
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)
search = GridSearchCV(pipe, param_grid, cv=5, scoring="f1")
search.fit(X_train, y_train)
best = search.best_estimator_
print("best params:", search.best_params_)
print("test F1:", f1_score(y_test, best.predict(X_test)))
joblib.dump(best, "churn_pipeline.joblib")What this demonstrates:
clf__n_estimators).(name, transformer) tuple; the final step must be an estimator.fit calls fit_transform on intermediate steps and fit on the last step.predict runs transform through all steps except the last, then predict.Pipeline implements the sklearn estimator interface - works with CV, grid search, and ensembles.ColumnTransformer) compose cleanly.| Approach | Leakage Risk | Deployment | CV Compatible |
|---|---|---|---|
| Manual fit/transform | High | Error-prone | No |
Pipeline | Low | Single artifact | Yes |
| Notebook cells | High | Not reproducible | No |
from sklearn.pipeline import FeatureUnion
combined = FeatureUnion([
("pca", PCA(n_components=10)),
("raw", "passthrough"),
])
pipe = Pipeline([("features", combined), ("clf", SVC())])
# Access steps by name
pipe.named_steps["clf"].feature_importances_scaler.fit(X_train) then model.fit(X_train_scaled) breaks under CV. Fix: wrap both in Pipeline.n_estimators instead of clf__n_estimators in grid search silently ignores params. Fix: use pipe.get_params().keys() to verify names.joblib.dump(pipe, ...) the full pipeline.sklearn.preprocessing.FunctionTransformer.predict. Fix: end with a classifier/regressor or use TransformedTargetRegressor.X[expected_cols] before predict.| Alternative | Use When | Don't Use When |
|---|---|---|
Pipeline | Standard sklearn workflows | Custom deep learning loops (use PyTorch modules) |
imblearn.pipeline.Pipeline | Resampling inside CV | No class imbalance |
Manual sklearn fit/transform | Teaching/debugging one step | Production or CV workflows |
sklearn.set_output(transform="pandas") | Need DataFrame output from transforms | Pure numpy downstream |
fitted_clf = pipe.named_steps["clf"]
importances = fitted_clf.feature_importances_named_steps is a dict keyed by step name.pipe.fit(X, y, clf__sample_weight=weights).from sklearn.base import BaseEstimator, TransformerMixin
class Log1pTransformer(BaseEstimator, TransformerMixin):
def fit(self, X, y=None):
return self
def transform(self, X):
import numpy as np
return np.log1p(X)BaseEstimator, TransformerMixin for sklearn compatibility.ColumnTransformer which selects columns per branch.Xt = pipe.named_steps["prep"].transform(X_train[:5])
print(Xt.shape, Xt[:2])transform on individual steps after fitting.transform_output="pandas" on ColumnTransformer for named columns.joblib.dump including metadata JSON (training date, data hash, metrics).xgboost.XGBClassifier as the final pipeline step.imblearn pipeline if combining with SMOTE.from sklearn.compose import TransformedTargetRegressor
model = TransformedTargetRegressor(regressor=pipe, func=np.log1p, inverse_func=np.expm1)Pipeline(memory=...) for caching expensive transforms during grid search.skl2onnx when all steps are supported.Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026