pandas Series & DataFrames
pandas adds labeled axes to columnar data: a Series is one typed column; a DataFrame is a table of aligned Series sharing an index.
Search across all documentation pages
pandas adds labeled axes to columnar data: a Series is one typed column; a DataFrame is a table of aligned Series sharing an index.
Quick-reference recipe card - copy-paste ready.
import pandas as pd
df = pd.DataFrame(
{"region": ["East", "West"], "revenue": [120, 340]},
index=pd.Index(["a", "b"], name="row"),
)
east = df.loc[df["region"] == "East", "revenue"]
subset = df.iloc[0:1, [0, 1]]When to reach for this:
import pandas as pd
orders = pd.DataFrame(
{
"order_id": [101, 102, 103, 104],
"region": pd.Categorical(["East", "West", "East", "West"]),
"revenue": [120.0, 340.0, None, 210.0],
"ordered_at": pd.to_datetime(
["2025-01-01", "2025-01-02", "2025-01-03", "2025-01-04"], utc=True
),
}
).set_index("order_id")
# Label-based selection
east = orders.loc[orders["region"] == "East", ["region", "revenue"]]
# Position-based head
head = orders.iloc[:2]
# Add computed column safely
orders = orders.copy()
orders["revenue_filled"] = orders["revenue"].fillna(orders["revenue"].median())
# Series reduction
monthly_total = orders["revenue_filled"].sum()
print(east)
print("monthly_total:", monthly_total)What this demonstrates:
.loc boolean filtering with column subset.copy() before assignment to avoid SettingWithCopyWarning.loc includes both endpoints for labels; .iloc is half-open like Python slices.dtype backend (string[pyarrow]) changes memory and missing-value behavior.| Goal | API |
|---|---|
| Filter rows | df.loc[mask] |
| Pick columns | df[["a", "b"]] |
| Scalar by label | df.loc[row, "col"] |
| Scalar by position | df.iloc[i, j] |
| Parameter | Type | Description |
|---|---|---|
index | Index-like | Row labels; defaults to RangeIndex |
columns | list[str] | Column names for DataFrame |
dtype | str/np.dtype | Force column type on construction |
df[df["x"] > 0]["y"] = 1 triggers SettingWithCopyWarning. Fix: df.loc[df["x"] > 0, "y"] = 1..loc["a"] returns multiple rows; scalar access fails. Fix: reset_index() or enforce unique keys upstream.object columns use lots of RAM. Fix: df["col"] = df["col"].astype("string[pyarrow]") in pandas 2.2+.Int64 when NA is possible.pd.to_datetime(..., utc=True) everywhere at ingest.| Alternative | Use When | Don't Use When |
|---|---|---|
| Polars DataFrame | Lazy scans, very large data, strict typing | Team already standardized on pandas idioms |
| PyArrow Table | Zero-copy interchange between systems | You need rich interactive indexing |
| SQL via DuckDB | Data already in warehouse tables | Tiny in-memory transforms |
| NumPy structured arrays | Fixed schema numeric records | Heterogeneous column types |
groupby when you need the group key as a column..loc access.df["col"] selects one column Series.df.loc[rows, cols] is label-based 2-D selection with alignment rules.inplace=True (discouraged).filter or chunked reads.import pandas as pd
df = pd.read_csv("big.csv", usecols=["region", "revenue"])s.name appears when the Series becomes a DataFrame column.pd.Series(..., name="revenue").import pandas as pd
df.columns.duplicated().any()df["x"] return multiple columns.pd.options.mode.copy_on_write = True, chained ops defer copies until write..loc assignment for clarity.df = df.reset_index()df.rename_axis("order_id").reset_index() to name the former index.dtype or pd.to_numeric(errors="coerce").sample = df.sample(n=100, random_state=42)Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026