Data Pipeline Skill
Building and reviewing ETL/ELT pipelines - an Agent Skill for Python 3.14, pandas 2.2+, and Polars 1.x.
Search across all documentation pages
Building and reviewing ETL/ELT pipelines - an Agent Skill for Python 3.14, pandas 2.2+, and Polars 1.x.
Scaffolds batch pipelines: source connectors, validation layer, transform functions, partitioned outputs (Parquet/Delta), orchestration hooks (Airflow/dbt/Cron), and data-quality checks with reproducible run IDs.
| Input | Why |
|---|---|
| Source type | API, S3 CSV, Postgres replica, streaming |
| Engine choice | pandas vs Polars ADR |
| Grain | Row-level events vs daily aggregates |
| SLA | Freshness window and allowed downtime |
| PII columns | Masking/hash requirements |
extract.py, transform.py, load.py, validate.py)pyproject.toml deps pinned with uvs3://bucket/dataset/dt=YYYY-MM-DD/uv run python -m pipeline.run --date 2026-07-08 --env staging
uv run pytest tests/test_transform.py -q
uv run python -m pipeline.validate --date 2026-07-08import pandera.pandas as pa
from pandera.typing import Series
class OrdersSchema(pa.DataFrameModel):
order_id: Series[str] = pa.Field(unique=True)
amount: Series[float] = pa.Field(ge=0)
status: Series[str] = pa.Field(isin=["pending", "paid", "refunded"])
def validate_orders(df):
return OrdersSchema.validate(df, lazy=True)Follow team ADR - Polars for large single-node transforms; pandas when ecosystem integrations dominate.
cron + systemd for one job; Airflow when DAG dependencies, backfill, and SLAs matter.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026