LLM App / RAG Skill
Retrieval and agent app patterns - an Agent Skill for Python 3.14 RAG/LLM services.
Search across all documentation pages
Retrieval and agent app patterns - an Agent Skill for Python 3.14 RAG/LLM services.
Produces RAG application checklist: document loaders, chunking strategy, embedding pipeline, vector store adapter, retrieval API, LLM orchestration (LangChain/LlamaIndex/LangGraph stubs), evaluation harness, and observability (latency, token usage, retrieval hit rate).
| Input | Why |
|---|---|
| Corpus source | PDF, HTML, tickets, code repos |
| Latency budget | p95 target for user-facing chat |
| Model provider | OpenAI, Anthropic, local vLLM |
| Vector DB | pgvector, Pinecone, Chroma |
| Compliance | PII redaction, retention, audit log |
ingest/ and retrieval/ modules with typed configuv run pytest tests/test_retrieval.pyuv run python -m rag.ingest --source data/docs/
uv run python -m rag.index --rebuild
uv run pytest tests/test_golden.py -q
uv run python -m rag.chat --question "What is the refund policy?"from dataclasses import dataclass
@dataclass
class Chunk:
doc_id: str
text: str
score: float
def retrieve(query: str, store, k: int = 5) -> list[Chunk]:
hits = store.similarity_search_with_score(query, k=k)
return [Chunk(doc_id=h[0].metadata["doc_id"], text=h[0].page_content, score=h[1]) for h in hits]Follow team ADR - skill outputs interface stubs both can implement; human docs in ai-agents-rag are authoritative.
Agents when multi-step tools needed - start RAG-only until retrieval quality plateaus.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026