Supervised Learning
Supervised learning maps input features to known labels. Regression predicts continuous values; classification assigns discrete categories. sklearn provides a consistent API across algorithms.
Search across all documentation pages
Supervised learning maps input features to known labels. Regression predicts continuous values; classification assigns discrete categories. sklearn provides a consistent API across algorithms.
Quick-reference recipe card - copy-paste ready.
from sklearn.linear_model import Ridge, LogisticRegression
from sklearn.ensemble import RandomForestClassifier, GradientBoostingRegressor
# Classification
clf = LogisticRegression(max_iter=1000)
clf.fit(X_train, y_train)
# Regression
reg = Ridge(alpha=1.0)
reg.fit(X_train, y_train)When to reach for this:
"""supervised_learning.py - regression and classification with sklearn."""
from __future__ import annotations
import pandas as pd
from sklearn.datasets import fetch_california_housing, load_breast_cancer
from sklearn.ensemble import RandomForestClassifier, RandomForestRegressor
from sklearn.linear_model import LogisticRegression, Ridge
from sklearn.metrics import mean_absolute_error, roc_auc_score
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
# --- Classification ---
cancer = load_breast_cancer(as_frame=True)
Xc, yc = cancer.data, cancer.target
Xc_train, Xc_test, yc_train, yc_test = train_test_split(Xc, yc, test_size=0.2, random_state=42)
clf_pipe = Pipeline([
("scale", StandardScaler()),
("clf", LogisticRegression(max_iter=1000, random_state=42)),
])
clf_pipe.fit(Xc_train, yc_train)
proba = clf_pipe.predict_proba(Xc_test)[:, 1]
print("classification AUC:", roc_auc_score(yc_test, proba))
# --- Regression ---
housing = fetch_california_housing(as_frame=True)
Xr, yr = housing.data, housing.target
Xr_train, Xr_test, yr_train, yr_test = train_test_split(Xr, yr, test_size=0.2, random_state=42)
reg = RandomForestRegressor(n_estimators=200, random_state=42, n_jobs=-1)
reg.fit(Xr_train, yr_train)
print("regression MAE:", mean_absolute_error(yr_test, reg.predict(Xr_test)))What this demonstrates:
fit/predict API for regression and classification.| Problem | Start With | Upgrade To |
|---|---|---|
| Binary classification | LogisticRegression | RandomForest, XGBoost |
| Multi-class | LogisticRegression (multinomial) | RandomForestClassifier |
| Regression (linear-ish) | Ridge | GradientBoostingRegressor |
| Regression (non-linear) | RandomForestRegressor | LightGBM |
| Small data, interpretability | Linear models | - |
| Large tabular, accuracy | Gradient boosting | - |
from sklearn.svm import SVC
from sklearn.neighbors import KNeighborsClassifier
# SVM: strong with scaled features, slower on large data
svm = Pipeline([("scale", StandardScaler()), ("clf", SVC(probability=True))])
# k-NN: lazy learner, sensitive to scale and dimensionality
knn = Pipeline([("scale", StandardScaler()), ("clf", KNeighborsClassifier(n_neighbors=5))])StandardScaler in a pipeline.max_depth=None memorizes noise on small data. Fix: limit depth, increase min_samples_leaf, use CV.predict_proba - some models need probability=True (SVC) for probability outputs. Fix: check estimator support before using proba.| Alternative | Use When | Don't Use When |
|---|---|---|
| Linear models | Interpretability, fast baseline | Complex non-linear interactions dominate |
| Random Forest | Robust default, little tuning | Need maximum accuracy on large tabular |
| Gradient boosting | Best tabular accuracy | Need fast training iteration or interpretability |
| Deep learning | Images, text, sequences | Small tabular with <10k rows |
alpha shrinks coefficients toward zero, reducing overfitting.alpha=0 is ordinary least squares.RidgeCV).RandomForest does not natively handle NaN - impute first.joblib.dump the fitted pipeline.pipe.predict(new_row) with the same column names.LogisticRegression(multi_class="multinomial", max_iter=1000)
# or
RandomForestClassifier() # handles multi-class nativelyaverage="macro" F1 when classes are imbalanced.for name, imp in zip(feature_names, reg.feature_importances_):
print(f"{name}: {imp:.3f}")feature_importances_.Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026