mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
## Description Adds `headroom wrap omp` / `headroom unwrap omp` — a one-command wrap for [Oh My Pi](https://www.npmjs.com/package/@oh-my-pi/pi-coding-agent) (`omp`), the pi-mono-lineage coding agent, as proposed in #1149. One honest correction to the issue: #1149 proposed reusing the `ANTHROPIC_BASE_URL` redirect from `wrap claude`. During implementation I probed that empirically and it turned out to be wrong — omp only reads `ANTHROPIC_BASE_URL` in its web-search helper; its **chat** endpoint comes from the model registry (`providers.anthropic.baseUrl` in `~/.omp/agent/models.yml`). With the env var pointed at a local probe server, omp's chat traffic still went straight to the real endpoint (0 probe hits); with a `models.yml` same-ID override, every request arrived at the probe (9/9 hits on `/v1/messages`). A same-ID override keeps omp's bundled Anthropic model catalog and stored credentials (both keyed by provider id `anthropic`), so only the endpoint moves. The wrap therefore injects a marker-fenced `providers.anthropic.baseUrl` override into `models.yml`, snapshotting the pre-wrap file **byte-for-byte** first, and `headroom unwrap omp` restores it exactly (or removes the file when the wrap created it) — the same durable-wrap + backup + unwrap contract `wrap codex` uses for `config.toml`. Closes #1149 ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `headroom/providers/omp/` (new provider slice): `models_yml_path()` (honors `PI_CODING_AGENT_DIR`), `inject_models_override()` (yaml-merge preserving user providers; pristine byte-for-byte backup, never re-snapshotted while managed), `restore_models_override()` (`restored` / `removed` / `noop`; never touches an unmanaged file), `build_launch_env()` - `headroom/cli/wrap.py`: `wrap omp` (mirrors the aider/vibe `_launch_tool` shape; rtk instructions into the project's `AGENTS.md`, which omp reads natively) and `unwrap omp` (restore models.yml + scrub rtk block + stop proxy) - `headroom/telemetry/context.py`: `omp` added to `_KNOWN_WRAP_AGENTS` so the stack slug reports `wrap_omp` instead of `unknown` - `README.md` (agent matrix row + unwrap list), `llms.txt`, `CHANGELOG.md` - `tests/test_cli/test_wrap_omp.py`: 16 tests (injection fresh/merge/re-inject, restore statuses incl. unmanaged-file safety, env passthrough, CLI wiring, unwrap flows) ## Testing - [ ] Unit tests pass (`pytest`) — all new + `test_cli` tests pass; the full suite carries **3 pre-existing failures** that reproduce identically on unmodified `origin/main` (same set, same asserts — see Test Output and the rebase-validation comment) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ uv run pytest -q # post-rebase, base4f22cbb03 failed, 7723 passed, 515 skipped in 262.64s FAILED tests/test_cli/test_wrap_claude_base_url.py::test_wrap_marker_is_stale_when_pid_reused FAILED tests/test_rtk_session_savings.py::test_rtk_reader_returns_none_on_nonzero_exit FAILED tests/test_rtk_session_savings.py::test_lean_ctx_reader_returns_none_on_failure_and_logs → all three reproduce identically on unmodified origin/main (4f22cbb0), run the same way (same worktree + venv, sources switched): 3 failed, 7707 passed — this branch = baseline + the 16 new tests, nothing else changes. (The pre-rebase run againste8151f05showed the same shape: one order-dependent flake that also reproduced on its baseline; these are env/order-dependent.) $ uv run pytest tests/test_cli/ -q # post-rebase 542 passed + 1 of the pre-existing failures above # includes the 16 new test_wrap_omp.py tests $ uv run ruff check . ; echo ruff-check-exit:$? All checks passed! ruff-check-exit:0 $ uv run ruff format --check . # post-rebase 1 pre-existing violation: headroom/proxy/handlers/anthropic.py — flagged identically on unmodified origin/main (not touched by this PR); every file this PR touches is clean $ uv run mypy headroom # post-rebase; output redirected to file; exit captured Success: no issues found in 409 source files mypy-exit:0 ``` ## Real Behavior Proof - Environment: macOS 15 (arm64, M1 Pro), Python 3.12.13 (uv venv, editable install incl. Rust `_core`), headroom @ this branch, base extras only (no `[ml]`), Anthropic account signed into omp. Initial proof ran on basee8151f05with omp 16.3.6 (`@oh-my-pi/pi-coding-agent` via bun); re-validated after the rebase onto4f22cbb0with omp 16.3.11 — fresh numbers in the rebase-validation comment. - Exact command / steps: four scenarios, run in this order — 1. Mechanism probe (why models.yml, not env): local HTTP probe server on `127.0.0.1:18999`; ran `omp -p "say ok" --model claude-fable-5 --no-session --no-tools` once with `ANTHROPIC_BASE_URL=http://127.0.0.1:18999`, once with `~/.omp/agent/models.yml` containing `providers.anthropic.baseUrl: http://127.0.0.1:18999`. 2. One-command path: `headroom wrap omp --no-rtk --port 8790 -- -p "Read CHANGELOG.md and count how many '### Fixed' headings it contains. Answer with just the number." --model claude-fable-5 --no-session --max-time 180` 3. Routing stats: separate proxy on :8788, wrap with `--no-proxy`, then `GET /stats`. 4. Restore: `headroom unwrap omp`, plus an isolated `PI_CODING_AGENT_DIR=/tmp/omp-agent-test` run with a pre-existing user `models.yml`, then `cmp` against the original. - Observed result: end-to-end routing through the proxy proven for every scenario — - Probe: env-var run → **0 probe hits**, omp answered normally (bypassed). models.yml run → **9 hits on `/v1/messages?beta=true`** with real Messages bodies. This is the routing mechanism the wrap uses. - One-command run: wrap started the proxy ("Proxy ready on http://127.0.0.1:8790"), wrote the override (`models.yml: providers.anthropic.baseUrl=http://127.0.0.1:8790/p/headroom-wrap-omp`), launched omp, and omp answered **"7"** (correct — real `read` tool work through the proxy). Proxy log for the session (3 requests, `anthropic_messages` path): ``` PERF model=claude-fable-5 msgs=1 tok_before=36 cache_read=0 cache_write=61939 cache_hit_pct=0 PERF model=claude-fable-5 msgs=3 tok_before=796 cache_read=0 cache_write=63308 cache_hit_pct=0 PERF model=claude-fable-5 msgs=5 tok_before=935 cache_read=63308 cache_write=215 cache_hit_pct=100 ``` Prompt caching survives the proxy (100% hit on the follow-up turn). - Routing stats (:8788 session): `requests.total: 2, by_provider: {"anthropic": 2}, by_model: {"claude-fable-5": 2}`, per-project prefix `/p/headroom-wrap-omp` attributed. - Unwrap: `Removed wrap-created models.yml` (file gone); isolated pre-existing-file run: backup created, user's `my-gw` provider preserved in the managed file, and after `unwrap omp` the restored file is **byte-identical** (`cmp` clean). - Compression: **not observed in this environment** — `tok_saved=0`, `transforms=router:noop` / `too_small`. Honest reading: omp minimizes its own tool outputs client-side (a 300-item JSON tool result reached the proxy at only ~657 tokens) and the `[ml]` text compressor wasn't installed; small print-mode payloads sit below crush thresholds, and passthrough-by-default is the documented safety contract. The wrap's value here is proven at the routing/lifecycle/cache layer; compression numbers will match whatever the proxy does for a given content mix. - Not tested: Windows / Linux; lean-ctx mode with omp (`HEADROOM_CONTEXT_TOOL=lean-ctx` — `lean-ctx init --agent omp` depends on lean-ctx recognizing the agent; failure degrades with a warning by design); long interactive (non `-p`) sessions; `--memory` / `--learn` / `--code-graph` flags combined with omp; OAuth-vs-API-key matrix beyond my local account. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [ ] New and existing unit tests pass locally with my changes — all except the 3 documented pre-existing failures, which fail identically on unmodified origin/main - [x] I have updated the CHANGELOG.md if applicable ## Screenshots (if applicable) N/A — terminal evidence inline above. ## Additional Notes - The models.yml override is regenerated from the pristine backup on every wrap, so re-running with a different `--port` updates the endpoint idempotently and the backup is never clobbered. - Scope note from #1149 stands: this routes omp's **Anthropic** provider family. omp's other providers (OpenAI-direct, Gemini, ...) resolve their endpoints from their own registry entries; users can already point those at Headroom with their own custom provider in `models.yml`. - `headroom/providers/omp/` deliberately contains no install-time / MCP pieces — this is the thin wrap + unwrap slice only. --------- Co-authored-by: JerrettDavis <mxjerrett@gmail.com> Co-authored-by: Tejas Chopra <chopratejas@gmail.com>
67 lines
5.4 KiB
Text
67 lines
5.4 KiB
Text
# Headroom
|
||
|
||
> Context optimization layer for LLM applications. Compress tool outputs, logs, files, and RAG chunks before they reach the model. Same answers, 60–95% fewer tokens. Library, proxy, and MCP server. Apache 2.0, local-first.
|
||
|
||
Headroom is shipped as a Python package (`headroom-ai`), a TypeScript package (`headroom-ai`), an OpenAI + Anthropic-compatible HTTP proxy (`headroom proxy`), and an MCP server (`headroom_compress`, `headroom_retrieve`, `headroom_stats` tools). All four modes use the same compression pipeline: per-content-type compressors (JSON, code, logs, diffs, text) feed into a Compress-Cache-Retrieve (CCR) store so compression stays reversible — the LLM can ask for the original whenever it wants.
|
||
|
||
The canonical, always-current documentation index lives at the docs site below. If you can fetch one URL, fetch that one; the entries here are a hand-curated subset.
|
||
|
||
## Canonical docs (start here)
|
||
|
||
- [Live llms.txt (full doc index)](https://headroom-docs.vercel.app/llms.txt): Auto-generated index of every doc page with descriptions.
|
||
- [Live llms-full.txt (every doc page concatenated)](https://headroom-docs.vercel.app/llms-full.txt): One Markdown blob containing every doc page. Use when you can spend the tokens for full context.
|
||
- [Docs site](https://headroom-docs.vercel.app/docs): Human-browsable docs with search.
|
||
- [GitHub repo](https://github.com/chopratejas/headroom): Source, issues, releases.
|
||
- [PyPI package](https://pypi.org/project/headroom-ai/): Python install.
|
||
- [npm package](https://www.npmjs.com/package/headroom-ai): TypeScript install.
|
||
|
||
## Install (copy-paste-runnable)
|
||
|
||
- Python: `pip install headroom-ai` (add `[all]` for every optional extra)
|
||
- TypeScript / Node: `npm install headroom-ai` (or `pnpm add headroom-ai`, `bun add headroom-ai`)
|
||
- Docker: `docker run -p 8787:8787 ghcr.io/chopratejas/headroom:latest`
|
||
- Run the proxy: `headroom proxy --port 8787` then point any client at `http://127.0.0.1:8787`
|
||
- Wrap an agent in one command: `headroom wrap claude` (also: `codex`, `copilot`, `cursor`, `aider`, `opencode`, `cline`, `continue`, `goose`, `openhands`, `openclaw`, `vibe`, `omp`)
|
||
|
||
## Entry points
|
||
|
||
- [Quickstart](https://headroom-docs.vercel.app/docs/quickstart): 5-minute end-to-end (install → compress → call the model).
|
||
- [Installation](https://headroom-docs.vercel.app/docs/installation): All install paths, extras, Docker tags, env vars.
|
||
- [Proxy server](https://headroom-docs.vercel.app/docs/proxy): Run as a local HTTP proxy in front of OpenAI / Anthropic / Gemini.
|
||
- [MCP server](https://headroom-docs.vercel.app/docs/mcp): `headroom_compress`, `headroom_retrieve`, `headroom_stats` for Claude Code / Cursor / any MCP host.
|
||
- [API reference](https://headroom-docs.vercel.app/docs/api-reference): Python + TypeScript `compress()` API.
|
||
|
||
## How it works
|
||
|
||
- [How compression works](https://headroom-docs.vercel.app/docs/how-compression-works): Three-stage pipeline + automatic content routing.
|
||
- [SmartCrusher](https://headroom-docs.vercel.app/docs/smart-crusher): Statistical JSON / array compression (70–90% on tool outputs).
|
||
- [Code compression](https://headroom-docs.vercel.app/docs/code-compression): AST-aware via tree-sitter (preserves imports, signatures, types).
|
||
- [Text & log compression](https://headroom-docs.vercel.app/docs/text-and-logs): Search results, build logs, diffs.
|
||
- [CCR (reversible)](https://headroom-docs.vercel.app/docs/ccr): Compress-Cache-Retrieve — originals never deleted; LLM retrieves on demand.
|
||
|
||
## SDK / framework integrations
|
||
|
||
- [Anthropic SDK](https://headroom-docs.vercel.app/docs/anthropic-sdk): `withHeadroom(anthropic)` wrapper.
|
||
- [OpenAI SDK](https://headroom-docs.vercel.app/docs/openai-sdk): `withHeadroom(openai)` wrapper.
|
||
- [Vercel AI SDK](https://headroom-docs.vercel.app/docs/vercel-ai-sdk): Middleware + `withHeadroom()`.
|
||
- [LangChain](https://headroom-docs.vercel.app/docs/langchain): Chat models, memory, retrievers, agents.
|
||
- [Agno](https://headroom-docs.vercel.app/docs/agno): Model wrapping + observability hooks.
|
||
- [Strands](https://headroom-docs.vercel.app/docs/strands): Model wrapping + hook-based tool output compression.
|
||
- [LiteLLM](https://headroom-docs.vercel.app/docs/litellm): Single callback; works with all 100+ LiteLLM providers.
|
||
|
||
## Memory & cross-agent state
|
||
|
||
- [Persistent memory](https://headroom-docs.vercel.app/docs/memory): Per-project SQLite + HNSW vector store. No cross-project bleed (GH #462).
|
||
- [SharedContext](https://headroom-docs.vercel.app/docs/shared-context): Compressed inter-agent context handoffs.
|
||
- [Failure learning](https://headroom-docs.vercel.app/docs/failure-learning): Offline analysis writes corrections to `CLAUDE.local.md` (default, gitignored) or `CLAUDE.md` (shared) / `AGENTS.md` / `GEMINI.md`.
|
||
|
||
## Operations
|
||
|
||
- [Configuration](https://headroom-docs.vercel.app/docs/configuration): Env vars, config file, per-call overrides.
|
||
- [Benchmarks](https://headroom-docs.vercel.app/docs/benchmarks): Token-savings numbers across content types.
|
||
- [Troubleshooting](https://headroom-docs.vercel.app/docs/troubleshooting): Common failure modes and fixes.
|
||
- [Limitations](https://headroom-docs.vercel.app/docs/limitations): What Headroom won't do well today.
|
||
|
||
## Licensing
|
||
|
||
Apache 2.0. Use commercially, modify, redistribute. Data stays on the user's machine when running the library, proxy, or MCP server locally. Anonymous telemetry is **off by default** (opt-in); enable with `HEADROOM_TELEMETRY=on` or `headroom proxy --telemetry`.
|