Embeddings & Similarity
Embeddings convert text into fixed-size vectors where semantic similarity corresponds to geometric closeness. They power semantic search, clustering, and RAG retrieval.
Search across all documentation pages
Embeddings convert text into fixed-size vectors where semantic similarity corresponds to geometric closeness. They power semantic search, clustering, and RAG retrieval.
Quick-reference recipe card - copy-paste ready.
from openai import OpenAI
import numpy as np
client = OpenAI()
emb = client.embeddings.create(model="text-embedding-3-small", input=["query text"])
vec = np.array(emb.data[0].embedding)
# cosine similarity
def cosine(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))When to reach for this:
"""embeddings_similarity.py - embed, index, and search documents."""
from __future__ import annotations
import numpy as np
from openai import OpenAI
client = OpenAI()
MODEL = "text-embedding-3-small"
documents = [
"FastAPI is a modern Python web framework for building APIs.",
"PyTorch is a deep learning framework with dynamic computation graphs.",
"pandas 2.2 provides DataFrame operations for tabular data analysis.",
"pytest is the standard testing framework for Python projects.",
]
def embed_texts(texts: list[str]) -> np.ndarray:
response = client.embeddings.create(model=MODEL, input=texts)
vectors = [item.embedding for item in response.data]
arr = np.array(vectors, dtype=np.float32)
norms = np.linalg.norm(arr, axis=1, keepdims=True)
return arr / norms # L2 normalize for cosine via dot product
doc_embeddings = embed_texts(documents)
query = "How do I test my Python code?"
query_vec = embed_texts([query])[0]
scores = doc_embeddings @ query_vec # cosine similarity (normalized)
ranked = sorted(zip(scores, documents), reverse=True)
for score, doc in ranked:
print(f"{score:.3f} {doc}")What this demonstrates:
text-embedding-3-small.| Model | Dims | Provider | Use |
|---|---|---|---|
| text-embedding-3-small | 1536 | OpenAI | General purpose, cost-effective |
| text-embedding-3-large | 3072 | OpenAI | Higher quality |
| all-MiniLM-L6-v2 | 384 | Local (sentence-transformers) | Offline, fast |
| voyage-3 | 1024 | Voyage AI | Retrieval-optimized |
# Local embeddings with sentence-transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
vectors = model.encode(documents, normalize_embeddings=True)| Alternative | Use When | Don't Use When |
|---|---|---|
| API embeddings (OpenAI) | No GPU, high quality | Data cannot leave your network |
| sentence-transformers | Local/offline, free | Need highest retrieval quality |
| BM25 keyword search | Exact term matching | Semantic/paraphrase matching needed |
| Hybrid (BM25 + dense) | Production RAG | Simple prototype with few docs |
client.embeddings.create(model="text-embedding-3-small", input=texts, dimensions=512)Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026