Reference: A Production ML Pipeline
A reference ML pipeline for Python shops: batch training with PyTorch 2.6+ or scikit-learn, artifact registry, FastAPI serving, and monitoring for drift and latency - not notebook-only science projects.
Search across all documentation pages
A reference ML pipeline for Python shops: batch training with PyTorch 2.6+ or scikit-learn, artifact registry, FastAPI serving, and monitoring for drift and latency - not notebook-only science projects.
Quick-reference recipe card - copy-paste ready.
fraud-ml/
pipelines/train.py # scheduled training job
pipelines/features.py # feature store writes
serving/app.py # FastAPI inference
models/ # exported weights (S3 in prod)
eval/metrics.py # offline gates
monitoring/drift.py # prod distribution checksWhen to reach for this:
"""pipelines/train.py - batch training with eval gate."""
from __future__ import annotations
import json
from pathlib import Path
import joblib
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import train_test_split
from pipelines.features import load_feature_frame
def train(version: str, min_auc: float = 0.82) -> Path:
X, y = load_feature_frame()
X_train, X_test, y_train, y_test = train_test_split(X, y, stratify=y, random_state=42)
model = GradientBoostingClassifier(random_state=42)
model.fit(X_train, y_train)
auc = roc_auc_score(y_test, model.predict_proba(X_test)[:, 1])
if auc < min_auc:
raise RuntimeError(f"AUC {auc:.3f} below gate {min_auc}")
out = Path(f"models/{version}/model.joblib")
out.parent.mkdir(parents=True, exist_ok=True)
joblib.dump(model, out)
out.with_suffix(".metrics.json").write_text(json.dumps({"auc": auc}))
return out"""serving/app.py - inference API."""
from __future__ import annotations
import joblib
from fastapi import FastAPI
from pydantic import BaseModel, Field
app = FastAPI()
MODEL = joblib.load("models/prod/model.joblib")
class Features(BaseModel):
amount: float = Field(gt=0)
merchant_risk_score: float
class Score(BaseModel):
fraud_probability: float
model_version: str = "2026.07.09"
@app.post("/v1/score", response_model=Score)
def score(body: Features) -> Score:
prob = float(MODEL.predict_proba([[body.amount, body.merchant_risk_score]])[0, 1])
return Score(fraud_probability=prob)What this demonstrates:
model.joblib + metadata JSON.| Stage | Tooling |
|---|---|
| Ingest | Airflow/Prefect job |
| Train | sklearn / PyTorch script in container |
| Eval | pytest + metric thresholds |
| Deploy | Update serving image env MODEL_VERSION |
| Monitor | Grafana + batch drift job |
# monitoring/drift.py - simplified PSI check
def population_stability_index(expected: list[float], actual: list[float]) -> float:
...features.py package imported both sides.pipelines/ modules.| Alternative | Use When | Don't Use When |
|---|---|---|
| Batch-only scoring | Nightly fraud batch | Real-time decline needed |
| Embedded model in rules | Tiny linear model | Deep network inference |
| Triton server | High RPS GPU | Small sklearn logistic |
| Vendor AutoML | Spike only | Long-term ownership needed |
sklearn for tabular baseline; swap train block for PyTorch when deep model required.
Scheduled K8s Job or managed pipeline; not on serving pods.
Start with warehouse tables; adopt store when online/offline skew hurts.
Separate eval (hallucination, cost) and guardrails - see LLM section.
Data science signs metrics; platform signs serving SLO; PM signs business metric.
Hash or tokenize; document retention in governance policy.
Measure p95 in canary; autoscale CPU serving before GPU unless needed.
MLflow or W&B for params/metrics; link run id to deployed version.
Router in serving app or separate deployments per model family.
Fixture rows with expected scores; contract test train→serve feature parity.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026