Search across all documentation pages
12 pages in this section.
The mental model underneath every data pipeline - batch vs. streaming as a latency/completeness trade-off, DAGs as dependency graphs rather than execution order, and why idempotency is what makes retries and backfills safe - the model behind every other page in this section.
Learn data engineering basics with 9 Python examples. Explore ETL pipelines, batch vs. streaming, and data transformation with pandas.
Learn how dbt integrates Python for warehouse transformations. Discover when to use Python for advanced analytics and ML features in dbt models.
Implement data validation in Python using Pandera for DataFrames and Great Expectations for data contracts. Catch bad data before it impacts your systems.
Learn how to use Parquet and Arrow for data storage in analytics pipelines. Discover benefits like compression, schema embedding, and selective column loading.
Learn data engineering best practices for building robust pipelines. Discover tips for data lake storage, code testing, and operational reliability.
A single-page roundup of every highlight bullet from the 11 pages in the Data Engineering section, grouped by source page so you can scan all 44 takeaways without opening each article individually.