beautifulsoup4 & lxml
BeautifulSoup parses messy HTML/XML into a navigable tree - tags, attributes, text search. Pair it with lxml as the parser backend for speed and XPath when documents grow large or malformed.
Search across all documentation pages
BeautifulSoup parses messy HTML/XML into a navigable tree - tags, attributes, text search. Pair it with lxml as the parser backend for speed and XPath when documents grow large or malformed.
from bs4 import BeautifulSoup
html = '<div class="item"><a href="/p/1">Widget</a><span class="price">$9</span></div>'
soup = BeautifulSoup(html, "lxml")
for item in soup.select(".item"):
name = item.find("a").get_text(strip=True)
price = item.find(class_="price").get_text(strip=True)
print(name, price)When to reach for this:
from __future__ import annotations
from dataclasses import dataclass
import httpx
from bs4 import BeautifulSoup
@dataclass
class Product:
sku: str
title: str
price: str
def parse_catalog(html: str) -> list[Product]:
soup = BeautifulSoup(html, "lxml")
products: list[Product] = []
for card in soup.select("article.product"):
sku = card.get("data-sku")
title_el = card.select_one("h2.title")
price_el = card.select_one("span.price")
if not sku or not title_el or not price_el:
continue
products.append(
Product(
sku=sku,
title=title_el.get_text(strip=True),
price=price_el.get_text(strip=True),
)
)
return products
def fetch_catalog(url: str) -> list[Product]:
with httpx.Client(timeout=10.0, headers={"User-Agent": "catalog-bot/1.0"}) as client:
response = client.get(url)
response.raise_for_status()
return parse_catalog(response.text)
if __name__ == "__main__":
sample = """
<article class="product" data-sku="W1">
<h2 class="title">Widget</h2><span class="price">$9.00</span>
</article>
"""
print(parse_catalog(sample))What this demonstrates:
select / select_oneget_text(strip=True) for normalized texthtml.parser (stdlib), lxml (fast C), html5lib (most lenient)..parent, .next_sibling, .find, .find_all, .select (CSS).tree.xpath("//div[@class='item']") when CSS is awkward.from_encoding when wrong.| Parser | Speed | Lenient | Dependency |
|---|---|---|---|
| lxml | Fast | Moderate | libxml2 |
| html.parser | Slow | Moderate | none |
| html5lib | Slowest | Very | html5lib |
# Parse XML with lxml directly for strict schemas
from lxml import etree
root = etree.fromstring(xml_bytes)
for node in root.xpath("//item[@id]"):
print(node.get("id"), node.text)data-* attributes; contract tests on HTML fixtures.rel=next or API cursor.| Alternative | Use When | Don't Use When |
|---|---|---|
| Official JSON API | Provider offers stable API | No API and low change frequency HTML |
selectolax | Maximum parse speed | Need BeautifulSoup ecosystem examples |
| Playwright/Selenium | JavaScript-rendered SPAs | Static server HTML |
feedparser | RSS/Atom feeds | Arbitrary HTML catalogs |
Use BeautifulSoup for ergonomics; lxml backend for speed. Direct lxml when you need XPath on strict XML.
urllib.parse.urljoin(base_url, href) on extracted a["href"] values.
Keep HTML fixtures in tests/fixtures/ and assert parsed dataclasses - no network in unit tests.
It can prettify or rewrite tags - treat output as derived, not authoritative source.
Use urllib.robotparser or a scraping framework that checks robots before fetch.
Pass session cookies via httpx Client - never commit credentials; use secrets manager.
Rate limit, rotate user agents responsibly, cache ETag responses - do not aggressive crawl.
pandas.read_html uses lxml/html5lib under the hood - good for one-off table extraction.
Register namespaces in lxml xpath: namespaces={"ns": "http://..."}.
Parsing is CPU-bound - run parse_catalog(html) in asyncio.to_thread after async fetch.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026