Experiment Tracking
Experiment tracking records parameters, metrics, and artifacts for every training run. MLflow and Weights & Biases are the standard tools for comparing runs and reproducing results.
Search across all documentation pages
Experiment tracking records parameters, metrics, and artifacts for every training run. MLflow and Weights & Biases are the standard tools for comparing runs and reproducing results.
import mlflow
with mlflow.start_run(run_name="rf_baseline"):
mlflow.log_params({"n_estimators": 100, "max_depth": 10})
mlflow.log_metric("f1", 0.87)
mlflow.sklearn.log_model(model, "model")"""experiment_tracking.py - MLflow tracking with sklearn."""
from __future__ import annotations
import mlflow
import mlflow.sklearn
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score, train_test_split
mlflow.set_experiment("breast-cancer-classification")
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
for n_est in [50, 100, 200]:
with mlflow.start_run(run_name=f"rf_{n_est}"):
clf = RandomForestClassifier(n_estimators=n_est, random_state=42)
cv_scores = cross_val_score(clf, X_train, y_train, cv=5, scoring="f1")
clf.fit(X_train, y_train)
test_f1 = cross_val_score(clf, X_test, y_test, cv="prefit", scoring="f1")[0]
mlflow.log_param("n_estimators", n_est)
mlflow.log_metric("cv_f1_mean", cv_scores.mean())
mlflow.log_metric("test_f1", test_f1)
mlflow.sklearn.log_model(clf, artifact_path="model",
registered_model_name="breast-cancer-rf")
print(f"n={n_est} cv_f1={cv_scores.mean():.3f} test_f1={test_f1:.3f}")with mlflow.start_run() in scripts.run_name and tags.| Alternative | Use When | Don't Use When |
|---|---|---|
| MLflow | General ML, self-hosted | Need built-in collaboration UI |
| W&B | Team dashboards, deep learning | Minimal local setup |
| TensorBoard | PyTorch training curves | Non-deep-learning experiments |
| CSV log | Quick personal experiments | Team collaboration |
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026