## Description Sync the docs with the code after the live-zone realignment. The `IntelligentContextManager` (ICM), `RollingWindow`, and scoring modules were deleted in PR #350 (May 2026), but the README and benchmark docstrings still advertised them as live, and an example still imported the deleted module (broken on run). This fixes the README + benchmarks and removes the dead example. I validated the README against the code with three parallel static-analysis sub-agents (features/architecture, CLI/extras/wrap-matrix, public API/integrations). Most of the README checked out accurate; only the items below were stale/wrong. Closes # ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [x] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - README: removed the `IntelligentContext` bullet and `IntelligentContext / RollingWindow` from the transforms list (both deleted in PR #350). - README: standardized `Kompress-base` -> `Kompress-v2-base` to match the HF model id `chopratejas/kompress-v2-base` and the existing badges (diagram re-aligned). - README: corrected the CodeCompressor language list to match the `CodeLanguage` enum (added TS, C, Perl). - README: softened the unanchored "6 algorithms" tagline to "content-aware compressors". - README: Cortex Code is library-mode only — there is no `headroom wrap cortex`, so the compatibility-matrix row no longer shows a wrap checkmark. - Deleted `examples/test_intelligent_context_toin_ccr.py` — it imported the deleted `IntelligentContextManager` (ImportError on run) and is unreferenced. - Removed stale `RollingWindow` mentions from benchmark docstrings/comments (`benchmarks/__init__.py`, `bench_transforms.py`, `bench_latency.py`, `scenarios/conversations.py`); the accurate PR-B1 retirement comment is kept. ## Testing - [ ] Unit tests pass (`pytest`) — N/A, docs/docstring + example deletion only - [x] Linting passes — `ruff check` clean on all changed benchmark files - [ ] Type checking passes — N/A (no type-relevant changes) - [ ] New tests added — N/A - [x] Manual testing performed — see Real Behavior Proof ### Test Output ```text $ ruff check benchmarks/__init__.py benchmarks/bench_transforms.py benchmarks/bench_latency.py benchmarks/scenarios/conversations.py All checks passed! # stale refs remaining in README/benchmarks (excluding accurate retirement notes): $ grep -rn "IntelligentContext|RollingWindow|Kompress-base" README.md benchmarks/ | grep -v retire (only benchmarks/bench_transforms.py:362 — the accurate PR-B1 retirement comment) # deleted example is unreferenced anywhere: $ grep -rn "test_intelligent_context_toin_ccr" --include=*.md --include=*.yml --include=*.py . (no hits) ``` ## Real Behavior Proof - Environment: macOS (darwin, arm64), Python 3.12 `.venv`, ruff 0.14.x, repo at branch `docs/sync-readme-with-code` off latest `main`. - Exact command / steps: (1) three parallel sub-agents grep/Read-validated README claims vs `headroom/`, `pyproject.toml`, `sdk/typescript/`; (2) directly verified each flagged mismatch (`CodeLanguage` enum, `HF_MODEL_ID`, absence of `IntelligentContext`/`RollingWindow` classes); (3) confirmed the example imports a deleted module and is unreferenced; (4) `ruff check` on changed benchmark files; (5) re-grepped README + benchmarks for any remaining stale refs. - Observed result: README and benchmark docstrings now match the code; the only surviving `RollingWindow` string is the accurate retirement comment; the broken example is removed; ruff passes; the ASCII architecture diagram still aligns after the `Kompress-v2-base` rename. - Not tested: rendering of the README on GitHub/PyPI (text-only change); the separate `docs/content/` and `wiki/` doc sets (see Additional Notes — out of scope for this PR). ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [ ] I have added tests that prove my fix is effective — N/A (docs/example cleanup) - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md — N/A (Release Please auto-generates from the conventional commit) ## Additional Notes **Larger related finding (NOT in this PR):** the published docs site (`docs/content/docs/*.mdx`) and the `wiki/*.md` set still document `IntelligentContextManager`, `RollingWindow`, `RollingWindowConfig`, `IntelligentContextConfig`, and `ScoringWeights` as live API — with `from headroom import RollingWindow` / `from headroom.transforms import IntelligentContextManager` code examples that would `ImportError`. It is half-migrated (a couple of `.mdx` files already note "removed in 0.9.x" while neighbors still teach it as current). This is ~15 files and the fixes require rewriting examples to the live-zone model, not just deletions — recommended as a focused follow-up PR rather than bundling it here.
12 KiB
Headroom
The Context Optimization Layer for LLM Applications
Compress everything your AI agent reads. Same answers, fraction of the tokens.
What It Does
Every tool call, DB query, file read, and RAG retrieval your agent makes is 70-95% boilerplate. Headroom compresses it away before it hits the model. The LLM sees less noise, responds faster, and costs less.
Your Agent / App
│
│ tool outputs, logs, DB reads, RAG results, file reads, API responses
▼
Headroom ← proxy, Python library, or framework integration
│
▼
LLM Provider (OpenAI, Anthropic, Google, Bedrock, 100+ via LiteLLM)
Headroom works as a transparent proxy (zero code changes), a Python function (compress()), or a framework integration (LangChain, Agno, Strands, LiteLLM, MCP).
Quick Start
=== "Proxy (Zero Code Changes)"
```bash
pip install "headroom-ai[all]"
headroom proxy
```
```bash
# Point any tool at the proxy
ANTHROPIC_BASE_URL=http://localhost:8787 claude
OPENAI_BASE_URL=http://localhost:8787/v1 your-app
```
That's it. Your existing code works unchanged, with 40-90% fewer tokens.
Want an always-on local runtime instead? See [Persistent Installs →](persistent-installs.md).
=== "Python SDK"
```python
from headroom import compress
result = compress(messages, model="claude-sonnet-4-5-20250929")
response = client.messages.create(
model="claude-sonnet-4-5-20250929",
messages=result.messages,
)
print(f"Saved {result.tokens_saved} tokens ({result.compression_ratio:.0%})")
```
Works with any Python LLM client. [Full SDK guide →](sdk.md)
=== "Coding Agents"
```bash
headroom wrap claude # Claude Code
headroom wrap copilot -- --model claude-sonnet-4-20250514
headroom wrap codex # OpenAI Codex CLI
headroom wrap aider # Aider
headroom wrap cursor # Cursor
headroom wrap openclaw # OpenClaw plugin bootstrap
```
Starts the proxy, points your tool at it, compresses everything automatically.
If you prefer an always-on proxy that `wrap` can reuse or recover, see [Persistent Installs →](persistent-installs.md).
=== "TypeScript SDK"
```typescript
import { compress } from 'headroom-ai';
const result = await compress(messages, { model: 'claude-sonnet-4-5-20250929' });
// Use result.messages with any LLM client
console.log(`Saved ${result.tokensSaved} tokens`);
```
Works with Vercel AI SDK, OpenAI Node SDK, and Anthropic TS SDK. [Full TS guide →](typescript-sdk.md)
=== "LiteLLM Callback"
```python
import litellm
from headroom.integrations.litellm_callback import HeadroomCallback
litellm.callbacks = [HeadroomCallback()]
# All 100+ providers now compressed automatically
```
Framework Integrations
LangChain
Wrap any chat model. Supports memory, retrievers, tools, streaming, async.
from headroom.integrations import HeadroomChatModel
llm = HeadroomChatModel(ChatOpenAI(model="gpt-4o"))
Agno
Full agent framework integration with observability hooks.
from headroom.integrations.agno import HeadroomAgnoModel
model = HeadroomAgnoModel(Claude(id="claude-sonnet-4-20250514"))
agent = Agent(model=model)
Strands
Model wrapping + tool output hook provider for Strands Agents.
from headroom.integrations.strands import HeadroomStrandsModel
model = HeadroomStrandsModel(wrapped_model=bedrock_model)
agent = Agent(model=model)
MCP Tools
Three tools for Claude Code, Cursor, or any MCP client: headroom_compress, headroom_retrieve, headroom_stats.
headroom mcp install && claude
TypeScript SDK
compress(), Vercel AI SDK middleware, OpenAI and Anthropic client wrappers.
npm install headroom-ai
OpenClaw
ContextEngine plugin for OpenClaw agents. Auto-compresses context in assemble().
headroom wrap openclaw
All integration patterns →{ .md-button }
How It Works
Headroom runs a two-stage pipeline on every request:
graph LR
A[Your Prompt] --> B[CacheAligner]
B --> C[ContentRouter]
C --> E[LLM Provider]
C -->|JSON| F[SmartCrusher]
C -->|Code| G[CodeCompressor]
C -->|Text| H[Kompress]
C -->|Logs| I[LogCompressor]
F --> E
G --> E
H --> E
I --> E
Stage 1: CacheAligner — Stabilizes message prefixes so the provider's KV cache actually hits. Claude offers a 90% read discount on cached prefixes; CacheAligner makes that work.
Stage 2: ContentRouter — Auto-detects content type (JSON, code, logs, search results, diffs, HTML, plain text) and routes each to the optimal compressor:
| Content Type | Compressor | How It Works |
|---|---|---|
| JSON arrays | SmartCrusher | Statistical analysis: keeps errors, anomalies, boundaries. No hardcoded rules. |
| Source code | CodeCompressor | AST-aware (tree-sitter). Preserves function signatures, collapses bodies. |
| Plain text | Kompress | ModernBERT token classification. Removes redundant tokens while preserving meaning. |
| Build/test logs | LogCompressor | Keeps failures, errors, warnings. Drops passing noise. |
| Search results | SearchCompressor | Ranks by relevance to user query, keeps top matches. |
| Git diffs | DiffCompressor | Preserves change hunks, drops unchanged context. |
| HTML | HTMLExtractor | Strips markup, extracts readable content. |
Context management is handled automatically inside the pipeline (live-zone-only compression): Headroom compresses only the newest content blocks (the latest user message and tool results) and never drops messages from history. The system prompt, tool definitions, and older turns — the provider cache hot zone — are left untouched so prompt caching keeps working.
Nothing is lost. Compressed content goes into the CCR store (Compress-Cache-Retrieve). The LLM gets a headroom_retrieve tool and can fetch full originals when it needs more detail.
Results
100 production log entries. One critical error buried at position 67.
| Metric | Baseline | Headroom |
|---|---|---|
| Input tokens | 10,144 | 1,260 |
| Correct answers | 4/4 | 4/4 |
87.6% fewer tokens. Same answer. The FATAL error was automatically preserved — not by keyword matching, but by statistical analysis of field variance.
Real Workloads
| Scenario | Before | After | Savings |
|---|---|---|---|
| Code search (100 results) | 17,765 | 1,408 | 92% |
| SRE incident debugging | 65,694 | 5,118 | 92% |
| Codebase exploration | 78,502 | 41,254 | 47% |
| GitHub issue triage | 54,174 | 14,761 | 73% |
Accuracy Benchmarks
| Benchmark | Category | N | Accuracy | Compression |
|---|---|---|---|---|
| GSM8K | Math | 100 | 0.870 | 0.000 delta |
| TruthfulQA | Factual | 100 | 0.560 | +0.030 delta |
| SQuAD v2 | QA | 100 | 97% | 19% reduction |
| BFCL | Tool/Function | 100 | 97% | 32% reduction |
| CCR Needle | Lossless | 50 | 100% | 77% reduction |
Full benchmark methodology → | Known limitations →
Key Features
Cloud Providers
Works with any LLM provider out of the box:
headroom proxy # Direct Anthropic/OpenAI
headroom proxy --backend bedrock --region us-east-1 # AWS Bedrock
headroom proxy --backend vertex_ai --region us-central1 # Google Vertex AI
headroom proxy --backend azure # Azure OpenAI
headroom proxy --backend openrouter # OpenRouter (400+ models)
Or via LiteLLM for 100+ providers (Together, Groq, Fireworks, Ollama, vLLM, etc.).
Installation
pip install headroom-ai # Core library (Python)
pip install "headroom-ai[all]" # Everything (recommended)
npm install headroom-ai # TypeScript / Node.js
pip install "headroom-ai[proxy]" # Proxy server + MCP tools
pip install "headroom-ai[ml]" # ML compression (Kompress, requires torch)
pip install "headroom-ai[langchain]" # LangChain integration
pip install "headroom-ai[agno]" # Agno integration
pip install "headroom-ai[evals]" # Evaluation framework
Requires Python 3.10+.
Next Steps
- Quickstart — Running in 5 minutes
- Integration Guide — Every way to add Headroom to your stack
- Architecture — How the pipeline works under the hood
- Benchmarks — Accuracy and latency data
- Limitations — When compression helps and when it doesn't
- Filesystem Contract — Canonical config/workspace env vars and paths
Apache 2.0 — Free for commercial use. GitHub | PyPI | Discord