NumPy Arrays
NumPy ndarray objects store homogeneous numeric data and expose vectorized operations that replace slow Python loops.
Search across all documentation pages
NumPy ndarray objects store homogeneous numeric data and expose vectorized operations that replace slow Python loops.
Quick-reference recipe card - copy-paste ready.
import numpy as np
# Create, cast dtype, vectorize
a = np.array([1, 2, 3], dtype=np.float64)
b = np.arange(3) # [0, 1, 2]
result = a * 2 + b # broadcasting + ufunc
print(result.shape, result.dtype)When to reach for this:
import numpy as np
# Monthly revenue matrix: rows = regions, cols = months
revenue = np.array(
[[120_000, 135_000, 128_000],
[98_000, 102_000, 110_000],
[150_000, 148_000, 155_000]],
dtype=np.int64,
)
# Per-region growth rate vs prior month (column 0 has no prior)
growth = np.diff(revenue, axis=1) / revenue[:, :-1]
# Normalize each region to its own mean (broadcasting)
normalized = (revenue - revenue.mean(axis=1, keepdims=True)) / revenue.std(axis=1, keepdims=True)
# Mask underperforming cells
threshold = revenue.mean()
under = np.where(revenue < threshold, revenue, 0)
print("growth shape:", growth.shape)
print("normalized mean per row:", normalized.mean(axis=1))
print("underperforming total:", under.sum())What this demonstrates:
dtype for stable integer mathnp.diff along axis=1 for column-wise month-over-month changekeepdims=True so row stats align with matrix shapenp.where as a vectorized conditional without Python if per cellndarray is a contiguous buffer plus shape/strides metadata - all elements share one dtype.arr.copy()) or from incompatible assignments.| Function | Typical use |
|---|---|
np.array | From Python sequences |
np.arange / np.linspace | Evenly spaced ranges |
np.zeros / np.ones / np.full | Preallocated buffers |
np.random.default_rng | Reproducible random data |
| dtype | When |
|---|---|
int32 / int64 | Counts and IDs |
float32 | Large ML features where precision trade-off is OK |
float64 | Default analytics precision |
bool | Masks and predicates |
import numpy as np
# Prefer explicit dtype on ingest
ids = np.array(["1", "2", "3"], dtype=np.int64) # ValueError - cast strings first
# Use default_rng, not legacy np.random.seed
rng = np.random.default_rng(42)
samples = rng.normal(0, 1, size=10_000)
# np.ndarray is not a drop-in for list - .append does not existnp.array - np.array([1, 2.5, 3]) becomes float64 and may surprise downstream integer code. Fix: cast explicitly at creation./ always promotes to float; // truncates toward negative infinity. Fix: use np.floor_divide when you need consistent flooring.arr[[0, 2]]) returns a copy; slices often return views. Fix: call .copy() before mutating filtered results.np.mean([1, np.nan, 3]) is NaN unless you pass nanmean. Fix: use np.nanmean or mask first.(3, 4) + (4,) works; (3, 4) + (3,) raises unless you reshape. Fix: align with arr[:, np.newaxis] or reshape.
| Alternative | Use When | Don't Use When |
|---|---|---|
| Python lists | Tiny sequences, heterogeneous objects | Million-element numeric transforms |
| pandas Series | Labeled 1-D data with missing values | Raw linear algebra on unlabeled buffers |
| Polars Series | Lazy columnar pipelines at scale | You only need a quick NumPy ufunc |
array module (stdlib) | Typed numeric arrays without NumPy dep | You need broadcasting or linear algebra |
np.broadcast_to(...).copy().float32 for large tensors where ~7 decimal digits suffice.float64 for financial aggregates unless you have a measured memory win.import numpy as np
rng = np.random.default_rng(42)
rng.integers(0, 100, size=5)default_rng is the modern, statistically safer API.int64 before sum: arr.astype(np.int64).sum().np.sum(arr, dtype=np.int64).np.str_ dtypes exist but pandas/Polars handle text better.(n,) mask on (n, m) needs (n, 1).mask[:, np.newaxis] or np.broadcast_to.safe = arr[mask].copy()
safe *= 2 # does not touch arr@ and np.dot are matrix multiply.@ follows PEP 465 stacking rules - prefer @ for clarity.timeit on array sizes you actually process.n crosses over where Python lists win due to creation overhead.Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026