Encoding & Unicode Bugs
Text bugs appear when bytes are decoded with the wrong codec, str and bytes mix without an explicit boundary, or visually identical Unicode strings compare unequal due to missing normalization.
Search across all documentation pages
Text bugs appear when bytes are decoded with the wrong codec, str and bytes mix without an explicit boundary, or visually identical Unicode strings compare unequal due to missing normalization.
Quick-reference recipe card - copy-paste ready.
# Always label bytes; decode at the boundary
raw: bytes = b"\xc3\xa9" # UTF-8 for é
text: str = raw.decode("utf-8")
# Writing text
path.write_text(text, encoding="utf-8")
path.write_bytes(raw)When to reach for this:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xffstr + bytes raises TypeErrormojibake like é instead of éimport unicodedata
from pathlib import Path
# --- bug: implicit decoding ---
def read_config_legacy(path: str) -> str:
with open(path) as f: # text mode without encoding= on Windows/locale edge cases
return f.read()
# --- fix: explicit UTF-8 ---
def read_config(path: str) -> str:
return Path(path).read_text(encoding="utf-8")
# --- bug: mixing str and bytes ---
def broken_join(prefix: str, payload: bytes) -> bytes:
return prefix + payload # TypeError
# --- fix: encode at boundary ---
def fixed_join(prefix: str, payload: bytes) -> bytes:
return prefix.encode("utf-8") + payload
# --- bug: comparison without normalization ---
def usernames_equal(a: str, b: str) -> bool:
return a == b
# --- fix: NFC normalization for user identifiers ---
def usernames_equal_normalized(a: str, b: str) -> bool:
return unicodedata.normalize("NFC", a.casefold()) == unicodedata.normalize("NFC", b.casefold())
# --- handling unknown bytes safely ---
def decode_best_effort(raw: bytes) -> str:
return raw.decode("utf-8", errors="replace")
# demo
assert decode_best_effort(b"ok\xc3\xa9") == "oké"
assert usernames_equal_normalized("café", "café") # visually equal formsWhat this demonstrates:
encoding="utf-8"str to bytes before binary concatenationunicodedata.normalize plus casefold for robust string equalityerrors="replace" for logging dirty data without crashingstr is Unicode code points; bytes is raw octets| Location | Mistake | Fix |
|---|---|---|
open() | default locale encoding | encoding="utf-8" |
| HTTP | wrong Content-Type charset | decode per header |
| JSON | json.dumps to bytes manually | use str JSON, encode once |
| Filenames | undecoded bytes paths | os.fsdecode / pathlib |
| DB | legacy Latin-1 columns | migrate to UTF-8 or decode explicitly |
# subprocess: text mode with explicit encoding (3.14)
import subprocess
subprocess.run(["echo", "hi"], capture_output=True, text=True, encoding="utf-8")
# sqlite stores TEXT as Unicode str when using text factory correctlychardet guessing in production - wrong guess corrupts data silently. Fix: fix upstream encoding; use errors="strict" on ingest.hash(s) not for crypto; hashlib needs .encode("utf-8").| Alternative | Use When | Don't Use When |
|---|---|---|
| UTF-8 everywhere | modern services | mandated legacy encodings |
errors="replace" | log ingestion | financial/legal raw records |
| IDNA/punycode | domain names | general text storage |
bytes end-to-end | crypto/protocol work | user-facing text without decode plan |
Use for Excel/Windows files with BOM. Most APIs and JSON use plain UTF-8 without BOM.
NFC composes characters (é as one code point); NFD decomposes (e + combining accent). Pick one form at storage and normalize on input.
Identify columns, migrate schema to UTF-8, convert bytes with known source encoding in a maintenance window, verify with checksum samples.
response.text guesses from headers; set response.encoding or use response.content.decode("utf-8") explicitly when headers lie.
casefold handles more Unicode special cases (e.g. German ß) for case-insensitive comparison of user-visible strings.
Use re.UNICODE default in Python 3; add (?u) or character classes like \w knowing they match Unicode letters.
Path uses str on Unix and Unicode APIs on Windows. Avoid passing undecoded bytes to open on Windows.
Base64 encodes bytes. Decode base64 to bytes, then decode bytes to str with the correct text codec.
JSON is Unicode; json.dumps returns str. Encode to UTF-8 bytes only at wire/storage boundary.
Inspect raw bytes with repr(raw) and try correct decode. Compare expected UTF-8 hex (e.g. c3 a9 for é).
UTF-8 mode (PYTHONUTF8=1) and explicit encoding= remain best practice on all supported versions including 3.14.0.
Rarely. Acceptable for lossy display of third-party logs, never for data you persist or bill on.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 19, 2026