Machine Learning Basics
10 examples to get you started with Classical ML - 7 basic and 3 intermediate.
Search across all documentation pages
10 examples to get you started with Classical ML - 7 basic and 3 intermediate.
uv venv && source .venv/bin/activate
uv pip install "pandas>=2.2" "scikit-learn>=1.5" "numpy>=2.0"Start every project by inspecting shape, dtypes, and missing values.
import pandas as pd
from sklearn.datasets import load_iris
iris = load_iris(as_frame=True)
df = iris.frame.assign(species=iris.target_names[iris.target])
print(df.shape, df.dtypes)
print(df.isna().sum())as_frame=True returns a pandas DataFrame directly from sklearn datasets.df.info() and df.describe() before modeling.Related: Datasets & Splits - train/test/validation splits
Define X (inputs) and y (label) explicitly.
X = df.drop(columns=["species"])
y = df["species"]
feature_names = list(X.columns)Related: Feature Engineering - encoding and scaling
Hold out unseen data before fitting anything.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)stratify=y preserves class proportions in classification tasks.X_train only - never on the full dataset.Related: Datasets & Splits - cross-validation
Establish a simple benchmark before tuning complex models.
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
clf = LogisticRegression(max_iter=1000, random_state=42)
clf.fit(X_train, y_train)
preds = clf.predict(X_test)
print("accuracy:", accuracy_score(y_test, preds))max_iter=1000 avoids convergence warnings on some datasets.Related: Supervised Learning - algorithm choices
Many algorithms expect similarly scaled inputs.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test) # fit on train onlyfit learns mean and std from training data; transform applies them.fit_transform on the test set - that leaks statistics.Pipeline to prevent manual mistakes.Related: Feature Engineering -
ColumnTransformer
Chain preprocessing and model into one estimable object.
from sklearn.pipeline import Pipeline
pipe = Pipeline([
("scaler", StandardScaler()),
("clf", LogisticRegression(max_iter=1000, random_state=42)),
])
pipe.fit(X_train, y_train)
print("pipeline accuracy:", accuracy_score(y_test, pipe.predict(X_test)))Pipeline ensures test data never influences preprocessing statistics.pipe.predict(X_test) runs scaling then classification in one call.Related: Pipelines - composing steps
Single train/test splits can be noisy; CV averages multiple folds.
from sklearn.model_selection import cross_val_score
scores = cross_val_score(pipe, X_train, y_train, cv=5, scoring="accuracy")
print(f"CV mean: {scores.mean():.3f} +/- {scores.std():.3f}")scoring appropriate to the problem (F1, ROC-AUC, etc.).Related: Model Evaluation - metrics and curves
Mixed numeric and categorical columns need targeted transforms.
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
num_cols = ["sepal length (cm)", "sepal width (cm)"]
cat_cols = ["petal length (cm)"] # example: treat one column as categorical
preprocessor = ColumnTransformer([
("num", StandardScaler(), num_cols),
("cat", OneHotEncoder(handle_unknown="ignore"), cat_cols),
])ColumnTransformer applies different transforms to different column groups.handle_unknown="ignore" prevents crashes on unseen categories at inference.Related: Feature Engineering - full encoding patterns
Search parameter combinations inside cross-validation.
from sklearn.model_selection import GridSearchCV
param_grid = {"clf__C": [0.1, 1.0, 10.0]}
search = GridSearchCV(pipe, param_grid, cv=5, scoring="accuracy")
search.fit(X_train, y_train)
print("best params:", search.best_params_)
print("test accuracy:", accuracy_score(y_test, search.predict(X_test)))clf__C).GridSearchCV picks the best params by CV score on training data.Related: Hyperparameter Tuning - Optuna and random search
Persist the full artifact for reproducible inference.
import joblib
joblib.dump(pipe, "iris_pipeline.joblib")
loaded = joblib.load("iris_pipeline.joblib")
print(loaded.predict(X_test[:3]))Related: Pipelines - serialization and deployment
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026