Data Analysis Basics
10 examples to get you started with Data Analysis - 7 basic and 3 intermediate.
Search across all documentation pages
10 examples to get you started with Data Analysis - 7 basic and 3 intermediate.
uv pip install "pandas>=2.2" "numpy>=2.0" "pyarrow>=18"Create a homogeneous numeric array and run vectorized math.
import numpy as np
sales = np.array([120, 340, 280, 410], dtype=np.int32)
bonus = sales * 0.1
print(bonus.dtype, bonus.sum())dtype controls memory and precision - pick it deliberately.Related: NumPy Arrays - dtypes and broadcasting
Wrap labeled data in pandas structures.
import pandas as pd
revenue = pd.Series([120, 340, 280], index=["Jan", "Feb", "Mar"], name="revenue")
df = pd.DataFrame({"region": ["East", "West", "East"], "revenue": revenue.values})
print(df.dtypes)Series is a single column with an index.DataFrame is a table of aligned Series sharing the same index..dtypes reveals whether columns are numeric, string, or categorical.Related: pandas Series & DataFrames - construction and selection
Load a file and profile it before transforming.
import pandas as pd
from io import StringIO
csv = "date,region,revenue\n2025-01-01,East,120\n2025-01-02,West,340\n"
df = pd.read_csv(StringIO(csv), parse_dates=["date"])
print(df.info())
print(df.isna().sum())parse_dates stores datetimes as proper dtypes, not strings.info() shows non-null counts and memory usage.Related: Cleaning & Transforming Data - missing values and dtypes
Use label-based and position-based indexing.
import pandas as pd
df = pd.DataFrame({"region": ["East", "West", "East"], "revenue": [120, 340, 280]})
east = df.loc[df["region"] == "East", ["region", "revenue"]]
first_row = df.iloc[0].loc selects by labels; .iloc selects by integer position..copy() if you will mutate.Related: pandas Series & DataFrames - indexing rules
Drop or fill gaps explicitly.
import pandas as pd
import numpy as np
df = pd.DataFrame({"revenue": [120, np.nan, 280]})
filled = df["revenue"].fillna(df["revenue"].median())
dropped = df.dropna(subset=["revenue"])fillna with a statistic is fine for exploration; document the policy for production.dropna removes rows - verify you are not biasing aggregates.Int64) when you need integers with NA.Related: Cleaning & Transforming Data - dtype fixes
Summarize by category with split-apply-combine.
import pandas as pd
df = pd.DataFrame({"region": ["East", "West", "East"], "revenue": [120, 340, 280]})
summary = (
df.groupby("region", observed=True)["revenue"]
.agg(["sum", "mean", "count"])
.reset_index()
)groupby splits rows, applies a function per group, and combines results.observed=True skips unused categorical levels in pandas 2.x..reset_index() turns the group key back into a normal column.Related: GroupBy & Aggregation - pivot tables
Combine datasets on a shared key.
import pandas as pd
orders = pd.DataFrame({"order_id": [1, 2], "region": ["East", "West"]})
revenue = pd.DataFrame({"order_id": [1, 2], "amount": [120, 340]})
merged = orders.merge(revenue, on="order_id", how="inner", validate="one_to_one")how="inner" keeps only matching keys - document what you drop.validate catches accidental many-to-many joins early.Related: Joins & Merges - duplicate detection
Aggregate daily data to monthly totals.
import pandas as pd
idx = pd.date_range("2025-01-01", periods=90, freq="D", tz="UTC")
df = pd.DataFrame({"revenue": range(90)}, index=idx)
monthly = df["revenue"].resample("ME").sum().resample."ME" is month-end frequency in modern pandas.Related: Time Series - rolling windows and timezones
Defer work and let the query planner optimize I/O.
import polars as pl
lf = pl.scan_parquet("sales.parquet")
result = (
lf.filter(pl.col("region") == "East")
.group_by("month")
.agg(pl.col("revenue").sum().alias("total"))
.collect()
)
print(result).collect() executes the plan - keep logic lazy until you need materialized data.Related: Polars - lazy evaluation patterns
Shrink a wide frame before it exhausts RAM.
import pandas as pd
df = pd.read_csv("large.csv", dtype={"region": "category", "units": "int32"})
df["revenue"] = pd.to_numeric(df["revenue"], downcast="float")
print(df.memory_usage(deep=True).sum() / 1e6, "MB")category stores repeated strings once - ideal for low-cardinality columns.downcast picks the smallest numeric dtype that fits your values.memory_usage(deep=True) after every heavy load.Related: Performance & Memory - chunking and Arrow
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026