Pipes & Text Processing
Compose Linux tools with pipes to filter logs, summarize CSV exports, and debug production - |, jq, awk, sort, and uniq before reaching for a Python script.
Search across all documentation pages
Compose Linux tools with pipes to filter logs, summarize CSV exports, and debug production - |, jq, awk, sort, and uniq before reaching for a Python script.
journalctl -u billing-api --since today | jq -r 'select(.level=="error") | .message' | sort | uniq -c | sort -nr | headWhen to reach for this:
#!/usr/bin/env bash
set -euo pipefail
# API latency p95 from structured access logs
curl -sf http://localhost:8000/metrics 2>/dev/null | head -1 # skip if not prometheus
# JSON application logs
journalctl -u billing-api --since "30 min ago" -o cat \
| jq -c 'select(.event=="request_complete")' \
| jq -s 'map(.duration_ms) | sort | .[length * 0.95 | floor]'
# CSV: top customers by spend
cut -d, -f2,5 orders.csv | tail -n +2 | sort -t, -k2 -nr | head -10What this demonstrates:
pipefail fails script if jq errors on bad JSON linejq -s slurps stream into array for percentile mathcut + sort for simple CSV analytics without pandasevent field filterstdbuf if needed.-c compact one object per line in stream.<(cmd) feeds command output as file to diff/compare.| Goal | Pipeline |
|---|---|
| Error histogram | ... | jq select(.level) | sort | uniq -c |
| Top IPs | awk '{print $1}' access.log | sort | uniq -c | sort -nr |
| Live tail filter | tail -f app.log | rg --line-buffered ERROR |
| Merge sorted files | sort -m file1 file2 |
# When pipeline gets unwieldy, one-liner Python is OK
journalctl -u billing-api -n 500 | python -c "
import sys, json
for line in sys.stdin:
try:
o=json.loads(line)
if o.get('level')=='error': print(o['message'])
except json.JSONDecodeError: pass
"set -o pipefail in bash scripts.jq -R 'fromjson? | select(.)' or filter with rg first.LC_ALL=C sort.bin/analyze_logs.sh in repo.LANG. Fix: export LANG=en_US.UTF-8.| Alternative | Use When | Don't Use When |
|---|---|---|
| pandas one-liner | Complex CSV in notebook | SSH quick triage |
| DuckDB CLI | SQL on CSV/parquet files | JSON logs already in journald |
| Log aggregator UI | Datadog/Splunk/Loki | No SaaS on air-gapped prod |
| Python structlog processor | Repeatable transform | One-off midnight incident |
jq understands structure; grep misses nested fields and breaks on reorder.
Insert tee /tmp/debug.jsonl between stages to inspect intermediate output.
find logs -name '*.jsonl' -print0 | xargs -0 -P4 jq ... for parallel chunk processing.
awk -F'\t' '{print $1}' - specify delimiter explicitly.
cmd | sponge file from moreutils - safer than redirect to same file.
echo "a b c" | column -t for human-readable aligned output.
-F retries when log file rotated - better for logrotate environments.
zcat file.gz | wc -l or rg -c '' per file.
Promote to scripts/ module with tests when used twice.
This doc targets Linux servers; dev on Windows use WSL for same pipelines.
Stack versions: This page was written for Python 3.14.0 (stable 3.14, maintenance 3.13), FastAPI 0.115+, Django 5.2, Flask 3.1, Pydantic 2, PyTorch 2.6+, pandas 2.2+, Polars 1.x, ruff 0.9+, and uv 0.6+.
Reviewed by Chris St. John·Last updated Jul 16, 2026