Chunking & Ingestion
Ingestion transforms raw documents into searchable chunks with embeddings. Chunk size, overlap, and metadata directly affect RAG retrieval quality.
Search across all documentation pages
Ingestion transforms raw documents into searchable chunks with embeddings. Chunk size, overlap, and metadata directly affect RAG retrieval quality.
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
chunks = splitter.split_documents(documents)When to reach for this: indexing PDFs, wikis, code repos, or any document corpus for RAG.
"""chunking_ingestion.py - clean, chunk, embed, and index documents."""
from __future__ import annotations
import re
import chromadb
from openai import OpenAI
client = OpenAI()
RAW_DOCS = [
{"id": "doc1", "source": "handbook", "text": "# Python Testing\n\nUse pytest for unit tests. " * 20},
{"id": "doc2", "source": "handbook", "text": "# FastAPI\n\nBuild APIs with type hints. " * 20},
]
def clean_text(text: str) -> str:
text = re.sub(r"\s+", " ", text).strip()
return text
def chunk_text(text: str, size: int = 400, overlap: int = 80) -> list[str]:
chunks = []
start = 0
while start < len(text):
chunks.append(text[start:start + size])
start += size - overlap
return chunks
chroma = chromadb.Client()
collection = chroma.get_or_create_collection("handbook")
for doc in RAW_DOCS:
cleaned = clean_text(doc["text"])
for i, chunk in enumerate(chunk_text(cleaned)):
chunk_id = f"{doc['id']}_chunk_{i}"
emb = client.embeddings.create(model="text-embedding-3-small", input=[chunk])
collection.add(
ids=[chunk_id],
documents=[chunk],
embeddings=[emb.data[0].embedding],
metadatas=[{"source": doc["source"], "doc_id": doc["id"], "chunk_index": i}],
)
print(f"indexed {collection.count()} chunks")What this demonstrates: text cleaning, fixed-size chunking with overlap, metadata per chunk, and embedding at ingest time.
| Content Type | Chunk Size | Overlap |
|---|---|---|
| Technical prose | 500-1000 chars | 10-20% |
| Code | Function/class level | Minimal |
| Legal/contracts | 200-500 tokens | 20% |
| Chat logs | Full turn or paragraph | 0-10% |
| Alternative | Use When | Don't Use When |
|---|---|---|
| Fixed-size chunks | General prose | Structured docs with clear sections |
| Recursive splitter | Mixed content | Code with strict boundaries |
| Semantic chunking | Variable section lengths | Simple, uniform documents |
| Parent-child chunks | Need both precision and context | Small corpora |
pymupdf or pdfplumber.#, ##) for semantic sections.Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026