Feature Engineering
Feature engineering transforms raw columns into representations models can learn from. Encoding categoricals, scaling numerics, and composing transforms with ColumnTransformer are the core patterns for tabular ML.
Search across all documentation pages
Feature engineering transforms raw columns into representations models can learn from. Encoding categoricals, scaling numerics, and composing transforms with ColumnTransformer are the core patterns for tabular ML.
Quick-reference recipe card - copy-paste ready.
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.ensemble import RandomForestClassifier
preprocessor = ColumnTransformer([
("num", StandardScaler(), numeric_cols),
("cat", OneHotEncoder(handle_unknown="ignore"), categorical_cols),
])
pipe = Pipeline([("prep", preprocessor), ("clf", RandomForestClassifier())])When to reach for this:
"""feature_engineering.py - mixed-type preprocessing with ColumnTransformer."""
from __future__ import annotations
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.metrics import classification_report
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
df = pd.DataFrame({
"age": [25, 34, 45, 29, 52],
"income": [45000, 62000, 88000, 51000, 95000],
"city": ["nyc", "la", "nyc", "chicago", "la"],
"plan": ["basic", "pro", "pro", "basic", "enterprise"],
"churned": [0, 0, 1, 0, 1],
})
X = df.drop(columns=["churned"])
y = df["churned"]
numeric_cols = ["age", "income"]
categorical_cols = ["city", "plan"]
preprocessor = ColumnTransformer(
transformers=[
("num", StandardScaler(), numeric_cols),
("cat", OneHotEncoder(handle_unknown="ignore", sparse_output=False), categorical_cols),
],
remainder="drop",
)
pipe = Pipeline([
("prep", preprocessor),
("clf", GradientBoostingClassifier(random_state=42)),
])
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.4, random_state=42)
pipe.fit(X_train, y_train)
print(classification_report(y_test, pipe.predict(X_test)))What this demonstrates:
OneHotEncoder with handle_unknown="ignore" for production safety.sparse_output=False for dense output compatible with tree and linear models.fit on training data per CV fold, preventing statistic leakage.get_feature_names_out() after fitting.| Transform | Column Type | sklearn Class |
|---|---|---|
| Standard scaling | Numeric | StandardScaler |
| Min-max scaling | Bounded numeric | MinMaxScaler |
| Log transform | Skewed numeric | FunctionTransformer(np.log1p) |
| One-hot | Low-cardinality categorical | OneHotEncoder |
| Ordinal | Ordered categories | OrdinalEncoder |
| Imputation | Missing values | SimpleImputer |
from sklearn.preprocessing import FunctionTransformer
import numpy as np
log_transform = FunctionTransformer(np.log1p, validate=False)
# Chain imputation before scaling in the numeric branch:
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline as SkPipeline
numeric_pipe = SkPipeline([
("impute", SimpleImputer(strategy="median")),
("scale", StandardScaler()),
])OneHotEncoder. Fix: handle_unknown="ignore" or handle_unknown="infrequent_if_exist".TargetEncoder inside a pipeline with proper CV.FeatureHasher), frequency encoding, or embeddings.PolynomialFeatures(degree=3) on 20 columns creates thousands of features. Fix: limit degree, use feature selection, or interaction-only mode.| Alternative | Use When | Don't Use When |
|---|---|---|
ColumnTransformer | Mixed column types in one model | All columns are the same type (use a single transformer) |
pandas.get_dummies | Quick EDA in notebooks | Production pipelines (not serializable, no unknown handling) |
| Target encoding | High-cardinality categoricals with enough data | Small datasets (severe overfitting risk) |
| Feature hashing | Very high cardinality, memory constraints | You need interpretable feature names |
joblib for deployment.Pipeline and cross-validation without leakage.pipe.fit(X_train, y_train)
names = pipe.named_steps["prep"].get_feature_names_out()
print(list(names))cat__city_nyc, num__age.remainder="passthrough" to keep unlisted columns as-is.from sklearn.impute import SimpleImputer
num_pipe = Pipeline([("impute", SimpleImputer(strategy="median")), ("scale", StandardScaler())])strategy="most_frequent" works for categorical imputation with SimpleImputer +p string dtype.FunctionTransformer to wrap any numpy-compatible function.df["order_year"] = df["order_date"].dt.year
df["order_dow"] = df["order_date"].dt.dayofweekTfidfVectorizer in a ColumnTransformer branch.category_encoders.TargetEncoder inside a Pipeline with CV.StandardScaler) when features have different units and scales.MinMaxScaler) when you need bounded [0,1] inputs.Xt = preprocessor.fit_transform(X_train)
print(Xt.shape) # rows x total encoded featuresStack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026