## Description Adds `headroom wrap omp` / `headroom unwrap omp` — a one-command wrap for [Oh My Pi](https://www.npmjs.com/package/@oh-my-pi/pi-coding-agent) (`omp`), the pi-mono-lineage coding agent, as proposed in #1149. One honest correction to the issue: #1149 proposed reusing the `ANTHROPIC_BASE_URL` redirect from `wrap claude`. During implementation I probed that empirically and it turned out to be wrong — omp only reads `ANTHROPIC_BASE_URL` in its web-search helper; its **chat** endpoint comes from the model registry (`providers.anthropic.baseUrl` in `~/.omp/agent/models.yml`). With the env var pointed at a local probe server, omp's chat traffic still went straight to the real endpoint (0 probe hits); with a `models.yml` same-ID override, every request arrived at the probe (9/9 hits on `/v1/messages`). A same-ID override keeps omp's bundled Anthropic model catalog and stored credentials (both keyed by provider id `anthropic`), so only the endpoint moves. The wrap therefore injects a marker-fenced `providers.anthropic.baseUrl` override into `models.yml`, snapshotting the pre-wrap file **byte-for-byte** first, and `headroom unwrap omp` restores it exactly (or removes the file when the wrap created it) — the same durable-wrap + backup + unwrap contract `wrap codex` uses for `config.toml`. Closes #1149 ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `headroom/providers/omp/` (new provider slice): `models_yml_path()` (honors `PI_CODING_AGENT_DIR`), `inject_models_override()` (yaml-merge preserving user providers; pristine byte-for-byte backup, never re-snapshotted while managed), `restore_models_override()` (`restored` / `removed` / `noop`; never touches an unmanaged file), `build_launch_env()` - `headroom/cli/wrap.py`: `wrap omp` (mirrors the aider/vibe `_launch_tool` shape; rtk instructions into the project's `AGENTS.md`, which omp reads natively) and `unwrap omp` (restore models.yml + scrub rtk block + stop proxy) - `headroom/telemetry/context.py`: `omp` added to `_KNOWN_WRAP_AGENTS` so the stack slug reports `wrap_omp` instead of `unknown` - `README.md` (agent matrix row + unwrap list), `llms.txt`, `CHANGELOG.md` - `tests/test_cli/test_wrap_omp.py`: 16 tests (injection fresh/merge/re-inject, restore statuses incl. unmanaged-file safety, env passthrough, CLI wiring, unwrap flows) ## Testing - [ ] Unit tests pass (`pytest`) — all new + `test_cli` tests pass; the full suite carries **3 pre-existing failures** that reproduce identically on unmodified `origin/main` (same set, same asserts — see Test Output and the rebase-validation comment) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ uv run pytest -q # post-rebase, base4f22cbb03 failed, 7723 passed, 515 skipped in 262.64s FAILED tests/test_cli/test_wrap_claude_base_url.py::test_wrap_marker_is_stale_when_pid_reused FAILED tests/test_rtk_session_savings.py::test_rtk_reader_returns_none_on_nonzero_exit FAILED tests/test_rtk_session_savings.py::test_lean_ctx_reader_returns_none_on_failure_and_logs → all three reproduce identically on unmodified origin/main (4f22cbb0), run the same way (same worktree + venv, sources switched): 3 failed, 7707 passed — this branch = baseline + the 16 new tests, nothing else changes. (The pre-rebase run againste8151f05showed the same shape: one order-dependent flake that also reproduced on its baseline; these are env/order-dependent.) $ uv run pytest tests/test_cli/ -q # post-rebase 542 passed + 1 of the pre-existing failures above # includes the 16 new test_wrap_omp.py tests $ uv run ruff check . ; echo ruff-check-exit:$? All checks passed! ruff-check-exit:0 $ uv run ruff format --check . # post-rebase 1 pre-existing violation: headroom/proxy/handlers/anthropic.py — flagged identically on unmodified origin/main (not touched by this PR); every file this PR touches is clean $ uv run mypy headroom # post-rebase; output redirected to file; exit captured Success: no issues found in 409 source files mypy-exit:0 ``` ## Real Behavior Proof - Environment: macOS 15 (arm64, M1 Pro), Python 3.12.13 (uv venv, editable install incl. Rust `_core`), headroom @ this branch, base extras only (no `[ml]`), Anthropic account signed into omp. Initial proof ran on basee8151f05with omp 16.3.6 (`@oh-my-pi/pi-coding-agent` via bun); re-validated after the rebase onto4f22cbb0with omp 16.3.11 — fresh numbers in the rebase-validation comment. - Exact command / steps: four scenarios, run in this order — 1. Mechanism probe (why models.yml, not env): local HTTP probe server on `127.0.0.1:18999`; ran `omp -p "say ok" --model claude-fable-5 --no-session --no-tools` once with `ANTHROPIC_BASE_URL=http://127.0.0.1:18999`, once with `~/.omp/agent/models.yml` containing `providers.anthropic.baseUrl: http://127.0.0.1:18999`. 2. One-command path: `headroom wrap omp --no-rtk --port 8790 -- -p "Read CHANGELOG.md and count how many '### Fixed' headings it contains. Answer with just the number." --model claude-fable-5 --no-session --max-time 180` 3. Routing stats: separate proxy on :8788, wrap with `--no-proxy`, then `GET /stats`. 4. Restore: `headroom unwrap omp`, plus an isolated `PI_CODING_AGENT_DIR=/tmp/omp-agent-test` run with a pre-existing user `models.yml`, then `cmp` against the original. - Observed result: end-to-end routing through the proxy proven for every scenario — - Probe: env-var run → **0 probe hits**, omp answered normally (bypassed). models.yml run → **9 hits on `/v1/messages?beta=true`** with real Messages bodies. This is the routing mechanism the wrap uses. - One-command run: wrap started the proxy ("Proxy ready on http://127.0.0.1:8790"), wrote the override (`models.yml: providers.anthropic.baseUrl=http://127.0.0.1:8790/p/headroom-wrap-omp`), launched omp, and omp answered **"7"** (correct — real `read` tool work through the proxy). Proxy log for the session (3 requests, `anthropic_messages` path): ``` PERF model=claude-fable-5 msgs=1 tok_before=36 cache_read=0 cache_write=61939 cache_hit_pct=0 PERF model=claude-fable-5 msgs=3 tok_before=796 cache_read=0 cache_write=63308 cache_hit_pct=0 PERF model=claude-fable-5 msgs=5 tok_before=935 cache_read=63308 cache_write=215 cache_hit_pct=100 ``` Prompt caching survives the proxy (100% hit on the follow-up turn). - Routing stats (:8788 session): `requests.total: 2, by_provider: {"anthropic": 2}, by_model: {"claude-fable-5": 2}`, per-project prefix `/p/headroom-wrap-omp` attributed. - Unwrap: `Removed wrap-created models.yml` (file gone); isolated pre-existing-file run: backup created, user's `my-gw` provider preserved in the managed file, and after `unwrap omp` the restored file is **byte-identical** (`cmp` clean). - Compression: **not observed in this environment** — `tok_saved=0`, `transforms=router:noop` / `too_small`. Honest reading: omp minimizes its own tool outputs client-side (a 300-item JSON tool result reached the proxy at only ~657 tokens) and the `[ml]` text compressor wasn't installed; small print-mode payloads sit below crush thresholds, and passthrough-by-default is the documented safety contract. The wrap's value here is proven at the routing/lifecycle/cache layer; compression numbers will match whatever the proxy does for a given content mix. - Not tested: Windows / Linux; lean-ctx mode with omp (`HEADROOM_CONTEXT_TOOL=lean-ctx` — `lean-ctx init --agent omp` depends on lean-ctx recognizing the agent; failure degrades with a warning by design); long interactive (non `-p`) sessions; `--memory` / `--learn` / `--code-graph` flags combined with omp; OAuth-vs-API-key matrix beyond my local account. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [ ] New and existing unit tests pass locally with my changes — all except the 3 documented pre-existing failures, which fail identically on unmodified origin/main - [x] I have updated the CHANGELOG.md if applicable ## Screenshots (if applicable) N/A — terminal evidence inline above. ## Additional Notes - The models.yml override is regenerated from the pristine backup on every wrap, so re-running with a different `--port` updates the endpoint idempotently and the backup is never clobbered. - Scope note from #1149 stands: this routes omp's **Anthropic** provider family. omp's other providers (OpenAI-direct, Gemini, ...) resolve their endpoints from their own registry entries; users can already point those at Headroom with their own custom provider in `models.yml`. - `headroom/providers/omp/` deliberately contains no install-time / MCP pieces — this is the thin wrap + unwrap slice only. --------- Co-authored-by: JerrettDavis <mxjerrett@gmail.com> Co-authored-by: Tejas Chopra <chopratejas@gmail.com>
32 KiB
██╗ ██╗███████╗ █████╗ ██████╗ ██████╗ ██████╗ ██████╗ ███╗ ███╗
██║ ██║██╔════╝██╔══██╗██╔══██╗██╔══██╗██╔═══██╗██╔═══██╗████╗ ████║
███████║█████╗ ███████║██║ ██║██████╔╝██║ ██║██║ ██║██╔████╔██║
██╔══██║██╔══╝ ██╔══██║██║ ██║██╔══██╗██║ ██║██║ ██║██║╚██╔╝██║
██║ ██║███████╗██║ ██║██████╔╝██║ ██║╚██████╔╝╚██████╔╝██║ ╚═╝ ██║
╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝╚═════╝ ╚═╝ ╚═╝ ╚═════╝ ╚═════╝ ╚═╝ ╚═╝
The context compression layer for AI agents
60–95% fewer tokens (for JSON data), 15-20% fewer tokens (for coding agents) · library · proxy · MCP · content-aware compressors · local-first · reversible
Docs · Install · Proof · Agents · Discord · llms.txt
AI agents / LLMs: read /llms.txt here, or fetch the live index / full docs blob.
Headroom compresses everything your AI agent reads — tool outputs, logs, RAG chunks, files, and conversation history — before it reaches the LLM. Same answers, fraction of the tokens.
Live: 10,144 → 1,260 tokens — same FATAL found.
What it does
- Library —
compress(messages)in Python or TypeScript, inline in any app - Proxy —
headroom proxy --port 8787, zero code changes, any language - Agent wrap —
headroom wrap claude|codex|grok|copilot|cursor|aider|opencode|cline|continue|goose|openhands|openclaw|vibe|omp|zcodein one command; undo withheadroom unwrap <tool> - MCP server —
headroom_compress,headroom_retrieve,headroom_statsfor any MCP client - Cross-agent memory — shared store across Claude, Codex, Gemini, Grok, auto-dedup
headroom learn— mines failed sessions, writes corrections toCLAUDE.local.md(default, gitignored) orCLAUDE.md/AGENTS.md/GEMINI.md/GROK.md- Output token reduction — trims what the model writes back (not just what you send): drops ceremony/restated code and skips deep "thinking" on routine steps. See Output token reduction.
- Reversible (CCR) — originals are cached for retrieval on demand
How it works (30 seconds)
Your agent / app
(Claude Code, Cursor, Codex, LangChain, Agno, Strands, your own code…)
│ prompts · tool outputs · logs · RAG results · files
▼
┌────────────────────────────────────────────────────┐
│ Headroom (runs locally — your data stays here) │
│ ──────────────────────────────────────────────── │
│ CacheAligner → ContentRouter → CCR │
│ ├─ SmartCrusher (JSON) │
│ ├─ CodeCompressor (AST) │
│ └─ Kompress-v2-base (text, HF) │
│ │
│ Cross-agent memory · headroom learn · MCP │
└────────────────────────────────────────────────────┘
│ compressed prompt + retrieval tool
▼
LLM provider (Anthropic · OpenAI · Bedrock · …)
- ContentRouter — detects content type, selects the right compressor
- SmartCrusher / CodeCompressor / Kompress-v2-base — compress JSON, AST, or prose
- CacheAligner — stabilizes prefixes so provider KV caches actually hit
- CCR — stores originals locally; LLM calls
headroom_retrieveif it needs them
→ Architecture · CCR reversible compression · Kompress-v2-base model card
Get started (60 seconds)
# 1 — Install
uv tool install "headroom-ai[all]" # Install `headroom` CLI as a global tool in self-contained virtual env
pip install "headroom-ai[all]" # Python — ships the `headroom` CLI
npm install headroom-ai # TypeScript SDK only — no `headroom` CLI
# 2 — Pick your mode (the `headroom` commands below come from the uv or pip install)
headroom deploy # turnkey local deployment + agent config
headroom wrap claude # wrap a coding agent
headroom proxy --port 8787 # drop-in proxy, zero code changes
# or: from headroom import compress # inline library
# 3 — Verify setup and see the savings
headroom doctor # health check — confirms routing is working
headroom perf
headroom dashboard # live savings dashboard (proxy must be running)
To use headroom, it is recommended you launch a wrapped agent session each time so that all necessary setup is completed. When wrapping a coding agent, headroom starts a local proxy, sets up an MCP server that provides tools such as rtk and tokensave, and launches a coding agent session configured to proxy requests to headroom.
The headroom CLI ships only via the PyPI package. The npm headroom-ai is the TypeScript SDK — a library you import (import { compress } from 'headroom-ai'), not a CLI, so it provides no headroom command.
Granular extras: [proxy], [mcp], [ml], [code], [memory], [vector] (optional HNSW backend — needs a C++ toolchain, not in [all]), [relevance], [image], [agno], [langchain], [evals], [pytorch-mps] (Apple-GPU memory-embedder offload — set HEADROOM_EMBEDDER_RUNTIME=pytorch_mps). Requires Python 3.10+.
Codex / global install
If Codex or another MCP client cannot inherit a shell PATH reliably, install Headroom as a persistent uv tool and point the client at the absolute binary path:
uv tool install "headroom-ai[all]"
command -v headroom
Then use the returned path in MCP config:
[mcp_servers.headroom]
command = "/absolute/path/from/command-v/headroom"
args = ["mcp", "serve"]
command = "headroom" only works when the client starts with a PATH that already includes the uv tool directory.
Proof
Savings on real agent workloads:
| Workload | Before | After | Savings |
|---|---|---|---|
| Code search (100 results) | 17,765 | 1,408 | 92% |
| SRE incident debugging | 65,694 | 5,118 | 92% |
| GitHub issue triage | 54,174 | 14,761 | 73% |
| Codebase exploration | 78,502 | 41,254 | 47% |
Accuracy preserved on standard benchmarks:
| Benchmark | Category | N | Baseline | Headroom | Delta |
|---|---|---|---|---|---|
| GSM8K | Math | 100 | 0.870 | 0.870 | ±0.000 |
| TruthfulQA | Factual | 100 | 0.530 | 0.560 | +0.030 |
| SQuAD v2 | QA | 100 | — | 97% | 19% compression |
| BFCL | Tools | 100 | — | 97% | 32% compression |
Reproduce: python -m headroom.evals suite --tier 1 · Full benchmarks & methodology
Output token reduction (cut what the model writes back)
Everything above shrinks the prompt you send. But you also pay for every token the model writes back — and on Opus-class models output costs 5× input. A lot of that output is waste: "Great, let me…" preambles, re-printing code you just showed it, and deep "thinking" on routine steps like reading a file.
Headroom can trim that too, from the proxy, without you changing any code:
- Verbosity steering — appends a short "be terse, don't restate context" note to the end of the system prompt (so your prompt cache still hits).
- Effort routing — when a turn is just the model resuming after a tool result (a file read, a passing test), it dials the model's thinking effort down. New questions and errors keep full effort.
Applies to Anthropic /v1/messages and OpenAI-compatible endpoints
(/v1/chat/completions, /v1/responses). Effort routing uses
reasoning_effort on OpenAI, thinking.budget_tokens /
output_config.effort on Anthropic — same clamp-only invariant on both
paths, same output_shaper:* label vocabulary.
Turn it on:
export HEADROOM_OUTPUT_SHAPER=1 # off by default
headroom proxy --port 8787
Already running a proxy? These switches are read live on every request, so a proxy that
headroom wrapreused (rather than started) would not see a value you export afterwards — its environment was snapshotted at launch.headroom wrapnow hot-syncs your current settings to the running proxy via a loopbackPOST /admin/runtime-env, so they take effect immediately with no restart (no cold start, no dropped requests, no lost caches). Set them before youwrap. On a shared proxy these overrides are global — the last explicit setting wins.
Learn the right terseness for you. People don't say how terse they want
answers — they show it (they interrupt long replies, or move on before they
could have read them). headroom learn --verbosity reads your past sessions and
picks the level automatically:
headroom learn --verbosity # preview what it found (dry run)
headroom learn --verbosity --apply # save it; the proxy uses it from now on
See how many output tokens you saved. Output savings are counterfactual — we never see what the model would have written — so Headroom reports an honest estimate with a confidence range, never a made-up number:
headroom output-savings
# Reduction: 31.7% (95% CI 27.7% … 35.7%) [estimated]
Want a measured number instead of an estimate? Leave 10% of conversations
unshaped as a control group: export HEADROOM_OUTPUT_HOLDOUT=0.1. The dashboard
shows an Output Tokens Saved card next to input compression, labelled
measured or estimated with the confidence band.
→ Full write-up incl. the measurement methodology: Output token reduction
Agent compatibility matrix
| Agent | headroom wrap |
Notes |
|---|---|---|
| Claude Code | ✅ | --memory · --code-graph · --1m · --tool-search |
| Codex | ✅ | shares memory with Claude |
| Grok CLI | ✅ | routes via GROK_CLI_CHAT_PROXY_BASE_URL |
| Cursor | Manual setup | starts proxy and prints base URLs for Cursor settings |
| Aider | ✅ | starts proxy + launches |
| Copilot CLI | ✅ | starts proxy + launches |
| OpenClaw | ✅ | installs as ContextEngine plugin |
| OpenCode | ✅ | injects config · starts proxy + launches |
| Cline | ✅ | starts proxy + injects config |
| Continue | ✅ | starts proxy + injects config |
| Goose | ✅ | starts proxy + launches |
| OpenHands | ✅ | starts proxy + launches |
| Mistral Vibe | ✅ | starts proxy + launches |
| Oh My Pi | ✅ | injects config · starts proxy + launches |
| Cortex Code | Library only | 60–65% savings (library mode; no wrap) |
| ZCode | ✅ | starts proxy and prints base URLs for ZCode settings |
Any OpenAI-compatible client works via headroom proxy. MCP-native: headroom mcp install.
Undo durable wrapping with headroom unwrap <tool> (supports: claude, copilot, codex, grok, omp, opencode, openclaw, zcode).
Registry authors can use the canonical server.json in the repo root instead of reconstructing the headroom mcp serve contract from prose.
GitHub Copilot CLI subscription mode
Headroom can route GitHub Copilot CLI subscription traffic through the local proxy:
headroom copilot-auth login
headroom wrap copilot --subscription -- --model gpt-4o
This lets Headroom intercept OpenAI-compatible Copilot CLI requests and apply the same proxy compression pipeline before forwarding to GitHub Copilot's hosted API. The wrapper exchanges Headroom's reusable GitHub OAuth token for Copilot's short-lived API token and prints the upstream endpoint as COPILOT_PROVIDER_API_URL=... during launch.
headroom copilot-auth login stores a Headroom-specific Copilot OAuth token.
This avoids relying on generic GitHub or Copilot CLI tokens that can read
Copilot account metadata but may still be rejected by Copilot's token-exchange
endpoint.
For GitHub Enterprise Server or custom-domain Copilot deployments, set one of these before launching:
export GITHUB_COPILOT_ENTERPRISE_DOMAIN=ghe.example.com
# or
export GITHUB_COPILOT_ENTERPRISE_URL=https://ghe.example.com
Both variables are supported. If both are set,
GITHUB_COPILOT_ENTERPRISE_URL takes precedence.
For GitHub.com Enterprise Cloud URLs such as
github.com/enterprises/your-enterprise, do not set an enterprise-domain
override. Headroom uses GitHub's normal token-exchange endpoint and the Copilot
API endpoint advertised for the signed-in account.
Platform support note: macOS auth reuse via Copilot CLI Keychain storage has been smoke-tested. Windows Credential Manager, Linux Secret Service / secret-tool, and Docker/CI token-injection paths are implemented or planned as auth-discovery paths, but still need real OS validation before they should be considered fully vetted. For Docker and CI, prefer passing an explicit GITHUB_COPILOT_TOKEN or GITHUB_COPILOT_GITHUB_TOKEN rather than relying on host keychain access.
When to use · When to skip
Great fit if you…
- run AI coding agents daily and want savings without changing your code
- work across multiple agents and want shared memory
- need reversible compression — originals are retrievable via CCR within the configured TTL
Skip it if you…
- only use a single provider's native compaction and don't need cross-agent memory
- work in a sandboxed environment where local processes can't run
Integrations — drop Headroom into any stack
| Your setup | Hook in with |
|---|---|
| Any Python app | compress(messages, model=…) |
| Any TypeScript app | await compress(messages, { model }) |
| Anthropic / OpenAI SDK | withHeadroom(new Anthropic()) · withHeadroom(new OpenAI()) |
| Vercel AI SDK | wrapLanguageModel({ model, middleware: headroomMiddleware() }) |
| LiteLLM | litellm.callbacks = [HeadroomCallback()] |
| LangChain | HeadroomChatModel(your_llm) |
| Agno | HeadroomAgnoModel(your_model) |
| Strands | Strands guide |
| ASGI apps | app.add_middleware(CompressionMiddleware) |
| Multi-agent | SharedContext().put / .get |
| MCP clients | headroom mcp install |
What's inside
- SmartCrusher — universal JSON: arrays of dicts, nested objects, mixed types.
- CodeCompressor — AST-aware for Python, JS/TS, Go, Rust, Java, C/C++, Perl.
- Kompress-v2-base — our HuggingFace model, trained on agentic traces.
- Image compression — 40–90% reduction via trained ML router.
- CacheAligner — stabilizes prefixes so Anthropic/OpenAI KV caches actually hit.
- Live-zone compression — compresses only new bytes (fresh tool output, latest turn); frozen prefix stays byte-identical so provider cache is not busted. History is never dropped.
- CCR — reversible compression; LLM retrieves originals on demand.
- Cross-agent memory — shared store, agent provenance, auto-dedup.
- SharedContext — compressed context passing across multi-agent workflows.
headroom learn— plugin-based failure mining for Claude, Codex, Gemini.
Pipeline internals
Headroom exposes one stable request lifecycle across compress(), the SDK, and the proxy:
Setup → Pre-Start → Post-Start → Input Received → Input Cached → Input Routed → Input Compressed → Input Remembered → Pre-Send → Post-Send → Response Received
- Transforms do the work: CacheAligner → ContentRouter → SmartCrusher / CodeCompressor / Kompress-base (live-zone only; IntelligentContext and RollingWindow were retired in PR-B1).
- Pipeline extensions observe or customize lifecycle stages via
on_pipeline_event(...). - Compression hooks sit alongside the canonical lifecycle as an additional extension seam.
- Proxy extensions remain the server/app integration seam for ASGI middleware, routes, and startup policy.
Provider and tool-specific behavior lives under headroom/providers/ so core orchestration stays focused on lifecycle, sequencing, and policy.
- CLI/tool slices:
headroom/providers/claude,copilot,codex,grok,openclaw - Provider runtime slices:
headroom/providers/claude,gemini, plus shared backend/runtime dispatch inheadroom/providers/registry.py - Core files stay orchestration-first:
wrap.py,client.py,cli/proxy.py, andproxy/server.pydelegate provider-specific env shaping, API target normalization, backend selection, and transport dispatch.
Headroom for teams
Headroom OSS is built for individual developers: run headroom proxy or headroom wrap on your laptop and start cutting tokens in minutes — free, local-first, your data never leaves your machine.
Running it across a whole engineering org is a different job: a shared, always-on deployment; centralized config and version rollout; org-wide savings dashboards; SSO and access controls; air-gapped / VPC installs; and someone to call when it matters. That's what we help companies with — self-hosted with support, or fully managed.
If your team is spending real money on LLM tokens — Claude Code, Codex, Cursor, or agents running in CI — and you want those savings across everyone, not just one laptop:
→ Email hello@headroomlabs.ai with your stack and rough monthly LLM spend, and we'll help you roll Headroom out across your organization.
Everything in this repo stays open source (Apache 2.0). The managed offering is simply for teams that would rather have it deployed, supported, and scaled for them.
Install
pip install "headroom-ai[all]" # Python, everything — includes the `headroom` CLI
npm install headroom-ai # TypeScript SDK (library only — no `headroom` CLI)
docker pull ghcr.io/chopratejas/headroom:latest
Granular extras: [proxy], [mcp], [ml] (Kompress-v2-base), [code], [memory], [vector] (optional HNSW backend — needs a C++ toolchain, not in [all]), [relevance], [image], [agno], [langchain], [evals], [pytorch-mps] (Apple-GPU memory-embedder offload — set HEADROOM_EMBEDDER_RUNTIME=pytorch_mps). Requires Python 3.10+.
Note
:
[all]covers the core stack but excludes framework adapters. Install them separately:pip install "headroom-ai[langchain]"(also[agno],[strands],[anyllm],[bedrock]).
Using pipx? Choose a supported interpreter explicitly:
pipx install --python python3.13 "headroom-ai[all]"
Pick 3.13 if you want dollar savings. The dashboard's Proxy $ Saved tile prices compression with LiteLLM, and LiteLLM can't be installed on Python 3.14+. On 3.14 token savings still track, but the dollar figure stays
$0.00. If you already installed on 3.14, switch withpipx reinstall headroom-ai --python python3.13and restart the proxy.
→ Installation guide — Docker tags, persistent service, PowerShell, devcontainers.
CPU requirement (x86/x86_64): the ONNX-backed features — Magika content detection and embedding relevance — use a precompiled ONNX Runtime that needs AVX2. On x86 hosts without AVX2 (some Docker/QEMU setups and older cloud VMs) Headroom automatically falls back to its non-ONNX paths (BM25 relevance, heuristic detection) rather than crashing.
arm64/Apple Silicon needs no AVX2.
Updating
headroom update # detects pip / pipx / uv tool and upgrades in place
headroom update --check # report the latest release without upgrading
headroom update --pre # include pre-releases
headroom update figures out how Headroom was installed (pip/venv, pip --user,
pipx, uv tool) and runs the matching upgrade across macOS, Linux, and Windows.
For git checkouts, editable installs, Docker images, and externally-managed
system Pythons (PEP 668) it prints the correct manual step instead of guessing.
The proxy also shows a one-line "update available" notice on startup. It checks
PyPI at most once a day, in the background, and never blocks. Opt out with
HEADROOM_UPDATE_CHECK=off (also skipped in --stateless mode and CI).
Corporate / SSL-inspection environments
If pip install "headroom-ai[all]" fails with CERTIFICATE_VERIFY_FAILED
(unable to get local issuer certificate), your network uses SSL inspection — a MITM
proxy presenting a company-issued CA. The build backend (maturin) downloads rustup over a
connection your TLS stack doesn't trust. Install Rust first so the build doesn't fetch it:
# macOS / Linux
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh && rustup default stable
# Windows
winget install Rustlang.Rustup && rustup default stable
Restart your shell, then pip install "headroom-ai[all]". A prebuilt wheel avoids the Rust
build entirely where available: pip install --only-binary headroom-ai headroom-ai. Prebuilt
wheels are published for Windows (win_amd64), Linux (x86_64 / aarch64), and macOS
(Apple Silicon and Intel), so installs on those platforms never need a local Rust toolchain — the
Rust-first dance above is only for the platform-independent sdist fallback when no wheel matches.
Two runtime assets are fetched over TLS; if they are blocked, trust your corporate CA via
REQUESTS_CA_BUNDLE / SSL_CERT_FILE / CURL_CA_BUNDLE:
cdn.pyke.io— the ONNX Runtime for the Rust core. Alternatively pre-provide it withORT_STRATEGY=systemandORT_LIB_LOCATION=/path/to/onnxruntime.huggingface.co— thekompress-basecompression model. Pre-download it and run withHF_HUB_OFFLINE=1, or setHF_ENDPOINTto a trusted mirror.
Running with compression disabled (pure gateway) requires neither asset.
Intel macOS (x86_64-apple-darwin): no prebuilt ONNX Runtime binary (#941)
ort-sys ships no prebuilt ONNX Runtime binary for Intel macOS, so a source
build fails by default even outside a corporate-proxy environment. The same
ORT_STRATEGY=system mechanism above fixes it — point it at a system ONNX
Runtime instead:
brew install onnxruntime
ORT_STRATEGY=system \
ORT_LIB_LOCATION="$(brew --prefix onnxruntime)/lib" \
ORT_PREFER_DYNAMIC_LINK=1 \
pip install "headroom-ai[all]"
# ORT is dlopen'd at runtime too:
export ORT_DYLIB_PATH="$(brew --prefix onnxruntime)/lib/libonnxruntime.dylib"
ORT_LIB_LOCATION must point at lib/ (not the bare prefix) and
ORT_PREFER_DYNAMIC_LINK=1 is required, or ORT_STRATEGY=system still
attempts static linking, which the Homebrew keg doesn't provide.
"Basic Constraints of CA cert not marked critical" (Python 3.13+ strict mode)
A different failure from the one above. If TLS fails with:
[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed:
Basic Constraints of CA cert not marked critical
then the corporate CA is found and trusted — adding it to a CA bundle changes nothing.
Python 3.13 + OpenSSL 3.x enable VERIFY_X509_STRICT by default, which enforces RFC 5280
§4.2.1.9: a CA cert's basicConstraints must be marked critical. Inspection roots like
Zscaler set CA:TRUE without the critical bit, so the chain is rejected.
Set HEADROOM_TLS_STRICT=0 to clear only the strict flag from every TLS context
Headroom controls — the proxy's httpx upstream client and the urllib3/huggingface_hub
path used for model downloads. Chain validation, signature, expiry, and hostname checks all
stay on; this is strictly narrower than disabling verification.
HEADROOM_TLS_STRICT=0 headroom proxy --port 8787
The Rust core's ONNX download (cdn.pyke.io) uses a separate TLS stack (rustls / OS trust
store), unaffected by HEADROOM_TLS_STRICT. On Windows the corporate root must be in the
machine certificate store (browsers already trust it there); or pre-provision ONNX
Runtime with ORT_STRATEGY=system + ORT_LIB_LOCATION=/path/to/onnxruntime to skip the
download entirely.
headroom learn
headroom learn — mines failed sessions, writes corrections to CLAUDE.local.md (default, gitignored; use --target CLAUDE.md for the shared team file) / AGENTS.md / GEMINI.md.
Documentation
| Start here | Go deeper |
|---|---|
| Quickstart | Architecture |
| Proxy | How compression works |
| MCP tools | CCR — reversible compression |
| Memory | Cache optimization |
| Failure learning | Benchmarks |
| Configuration | Limitations |
Persistent installs (headroom init / headroom install apply) |
Savings analytics (headroom savings / headroom perf / headroom doctor) |
Compared to
Headroom runs locally, covers every content type, works with every major framework, and is reversible.
| Scope | Deploy | Local | Reversible | |
|---|---|---|---|---|
| Headroom | All context — tools, RAG, logs, files, history | Proxy · library · middleware · MCP | Yes | Yes |
| RTK | CLI command outputs | CLI wrapper | Yes | No |
| lean-ctx | Tool output, files, shell, history | Proxy · library · middleware · MCP · CLI | Yes | Yes |
| Compresr, Token Co. | Text sent to their API | Hosted API call | No | No |
| OpenAI Compaction | Conversation history | Provider-native | No | No |
Attribution. Headroom ships with the excellent RTK binary for shell-output rewriting —
git show --short, scopedls, summarized installers. Huge thanks to the RTK team; their tool is a first-class part of our stack, and Headroom compresses everything downstream of it. Headroom can also use lean-ctx as the selected CLI context tool; setHEADROOM_CONTEXT_TOOL=lean-ctxbefore runningheadroom wrap ....
Contributing
git clone https://github.com/chopratejas/headroom.git && cd headroom
uv sync --extra dev && uv run pytest
Devcontainers in .devcontainer/ (default + memory-stack with Qdrant & Neo4j). See CONTRIBUTING.md.
Community
- Discord — questions, feedback, war stories.
- Kompress-v2-base on HuggingFace — the model behind our text compression.
Community projects
- Claude Code status-line indicator — a Claude Code plugin that shows live Headroom usage in your status line: idle until
headroom_compressfires, then the running total of tokens saved.
License
Apache 2.0 — see LICENSE.