Hugging Face transformers
The transformers library loads and runs open-source models from the Hugging Face Hub. Use pipeline for quick inference or AutoModel for custom generation loops.
Search across all documentation pages
The transformers library loads and runs open-source models from the Hugging Face Hub. Use pipeline for quick inference or AutoModel for custom generation loops.
Quick-reference recipe card - copy-paste ready.
from transformers import pipeline
generator = pipeline("text-generation", model="meta-llama/Llama-3.2-1B-Instruct")
output = generator("Explain Python decorators:", max_new_tokens=100)
print(output[0]["generated_text"])When to reach for this:
"""hugging_face_transformers.py - pipeline and manual generation."""
from __future__ import annotations
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
model_id = "meta-llama/Llama-3.2-1B-Instruct"
# Quick inference with pipeline
gen = pipeline("text-generation", model=model_id, torch_dtype=torch.float16, device_map="auto")
result = gen("What is pytest?", max_new_tokens=80, do_sample=True, temperature=0.7)
print("pipeline:", result[0]["generated_text"][-200:])
# Manual generation loop
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.float16, device_map="auto")
messages = [{"role": "user", "content": "Write a one-line Python hello world."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.inference_mode():
outputs = model.generate(**inputs, max_new_tokens=60, temperature=0.7, do_sample=True)
response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print("manual:", response)What this demonstrates:
pipeline for zero-config text generation.AutoTokenizer and AutoModelForCausalLM for manual control.apply_chat_template for instruction-tuned models.device_map="auto" for automatic GPU placement.config.json.| API | Use When | Complexity |
|---|---|---|
pipeline() | Quick prototyping | Low |
model.generate() | Custom generation params | Medium |
Trainer | Fine-tuning | Medium |
trl (SFT, DPO) | RLHF-style alignment | High |
# 4-bit quantization for limited GPU memory
from transformers import BitsAndBytesConfig
bnb_config = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.float16)
model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config=bnb_config, device_map="auto")tokenizer.apply_chat_template.device_map="auto", or smaller models.huggingface-cli login and accept license on model page.tokenizer.pad_token = tokenizer.eos_token.HF_HOME cache directory; use local_files_only=True in production.torch_dtype=torch.float16 or bfloat16.
| Alternative | Use When | Don't Use When |
|---|---|---|
pipeline | Quick inference | Custom decoding logic |
| vLLM | High-throughput serving | Simple local experiments |
| Ollama | Easy local model management | Fine-tuning |
| API providers | No GPU infrastructure | Data privacy requires local |
pip install huggingface_hub
huggingface-cli loginfrom transformers import TrainingArguments, Trainer
args = TrainingArguments(output_dir="./out", num_train_epochs=3, per_device_train_batch_size=4)
trainer = Trainer(model=model, args=args, train_dataset=dataset)
trainer.train()max_new_tokens limits new tokens (preferred over max_length).min_new_tokens ensures minimum response length.do_sample=False for greedy (deterministic) decoding.generate.torch.compile can speed up generation on supported models.export HF_HOME=/data/hf_cachefrom transformers import AutoModel
embedder = pipeline("feature-extraction", model="sentence-transformers/all-MiniLM-L6-v2")sentence-transformers library directly.Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026