Search across all documentation pages
13 pages in this section.
Why vectorization and broadcasting make NumPy fast, how pandas builds a labeled Series/DataFrame/Index model on top of raw arrays, and where Polars' lazy query-plan model breaks from both - the mental model behind every other page in this section.
Learn data analysis basics with 10 Python examples using NumPy and pandas. Create arrays, DataFrames, and read CSVs to prepare for analysis.
Learn to create and manipulate NumPy arrays for efficient numerical operations. Understand dtypes, broadcasting, and vectorized math for data analysis.
Learn to select and rearrange NumPy array data using indexing and reshape arrays without changing values. Includes examples and use cases.
Learn to construct, select, and index pandas Series and DataFrames. Explore examples for tabular EDA, time series, and data manipulation.
Clean and transform raw data using pandas. Learn to handle missing values, correct data types, and normalize messy strings for reliable analysis.
Learn to use pandas groupby for split-apply-combine operations and pivot tables. Summarize data, calculate per-group statistics, and transform data for reporting.
Learn to combine pandas DataFrames using merge and concat with explicit join types and validation to prevent data duplication and loss.
Learn to analyze time-indexed data with pandas. Master resampling, rolling statistics, timezone conversions, and grouping for sales and sensor data.
Improve data analysis with best practices for reproducible, vectorized, and well-typed code. Prevent errors in notebooks and production jobs.
A single-page roundup of every highlight bullet from the 12 pages in the Data Analysis section, grouped by source page so you can scan all 48 takeaways without opening each article individually.