From 3ebcd89d4663da4d496b2e9a717568ee224f92e9 Mon Sep 17 00:00:00 2001
From: chopratejas
Date: Thu, 19 Feb 2026 10:17:33 -0800
Subject: [PATCH] Rewrite README + add Integration Guide
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
README: 694 → 203 lines. Crisp, scannable, links to docs.
- compress() as the hero quickstart (not proxy)
- Integration table: compress(), LiteLLM, ASGI, proxy, Agno, LangChain
- LangChain marked as experimental
- "Already have a proxy?" callout linking to Integration Guide
- Architecture: ContentRouter (not SmartCrusher) as the primary compressor
New: docs/integration-guide.md
- Detailed setup for every integration path
- compress() with Anthropic, OpenAI, LiteLLM, raw HTTP
- LiteLLM callback + LiteLLM proxy ASGI middleware
- ASGI middleware for any FastAPI/Starlette app
- Compression hooks for advanced customization
- FAQ section
Fix: compress() uses default pipeline (CacheAligner + ContentRouter +
IntelligentContext) instead of manually specifying SmartCrusher.
---
README.md | 615 +++++++-------------------------------
docs/integration-guide.md | 292 ++++++++++++++++++
headroom/compress.py | 18 +-
3 files changed, 403 insertions(+), 522 deletions(-)
create mode 100644 docs/integration-guide.md
diff --git a/README.md b/README.md
index 51d9a0028..c8130f857 100644
--- a/README.md
+++ b/README.md
@@ -29,7 +29,6 @@
-
---
## Demo
@@ -40,570 +39,164 @@
---
-## Does It Actually Work? A Real Test
-
-**The setup:** 100 production log entries. One critical error buried at position 67.
-
-
-BEFORE: 100 log entries (18,952 chars) - click to expand
-
-```json
-[
- {"timestamp": "2024-12-15T00:00:00Z", "level": "INFO", "service": "api-gateway", "message": "Request processed successfully - latency=50ms", "request_id": "req-000000", "status_code": 200},
- {"timestamp": "2024-12-15T01:01:00Z", "level": "INFO", "service": "user-service", "message": "Request processed successfully - latency=51ms", "request_id": "req-000001", "status_code": 200},
- {"timestamp": "2024-12-15T02:02:00Z", "level": "INFO", "service": "inventory", "message": "Request processed successfully - latency=52ms", "request_id": "req-000002", "status_code": 200},
- // ... 64 more INFO entries ...
- {"timestamp": "2024-12-15T03:47:23Z", "level": "FATAL", "service": "payment-gateway", "message": "Connection pool exhausted", "error_code": "PG-5523", "resolution": "Increase max_connections to 500 in config/database.yml", "affected_transactions": 1847},
- // ... 32 more INFO entries ...
-]
-```
-
-
-**AFTER:** Headroom compresses to 6 entries (1,155 chars):
-
-```json
-[
- {"timestamp": "2024-12-15T00:00:00Z", "level": "INFO", "service": "api-gateway", ...},
- {"timestamp": "2024-12-15T01:01:00Z", "level": "INFO", "service": "user-service", ...},
- {"timestamp": "2024-12-15T02:02:00Z", "level": "INFO", "service": "inventory", ...},
- {"timestamp": "2024-12-15T03:47:23Z", "level": "FATAL", "service": "payment-gateway", "error_code": "PG-5523", "resolution": "Increase max_connections to 500 in config/database.yml", "affected_transactions": 1847},
- {"timestamp": "2024-12-15T02:38:00Z", "level": "INFO", "service": "inventory", ...},
- {"timestamp": "2024-12-15T03:39:00Z", "level": "INFO", "service": "auth", ...}
-]
-```
-
-**What happened:** First 3 items + the FATAL error + last 2 items. The critical error at position 67 was automatically preserved.
-
----
-
-**The question we asked Claude:** "What caused the outage? What's the error code? What's the fix?"
-
-| | Baseline | Headroom |
-|--|----------|----------|
-| Input tokens | 10,144 | 1,260 |
-| Correct answers | **4/4** | **4/4** |
-
-Both responses: *"payment-gateway service, error PG-5523, fix: Increase max_connections to 500, 1,847 transactions affected"*
-
-**87.6% fewer tokens. Same answer.**
-
-Run it yourself: `python examples/needle_in_haystack_test.py`
-
----
-
-## Accuracy Benchmarks
-
-> **Headroom's guarantee: compress without losing accuracy.**
-
-We validate against established open-source benchmarks. Full methodology and reproducible tests: [Benchmarks Documentation](https://chopratejas.github.io/headroom/benchmarks/)
-
-| Benchmark | Metric | Result | Status |
-|-----------|--------|--------|--------|
-| [Scrapinghub Article Extraction](https://huggingface.co/datasets/allenai/scrapinghub-article-extraction-benchmark) | F1 Score | **0.919** (baseline: 0.958) | :white_check_mark: |
-| [Scrapinghub Article Extraction](https://huggingface.co/datasets/allenai/scrapinghub-article-extraction-benchmark) | Recall | **98.2%** | :white_check_mark: |
-| [Scrapinghub Article Extraction](https://huggingface.co/datasets/allenai/scrapinghub-article-extraction-benchmark) | Compression | **94.9%** | :white_check_mark: |
-| SmartCrusher (JSON) | Accuracy | **100%** (4/4 correct) | :white_check_mark: |
-| SmartCrusher (JSON) | Compression | **87.6%** | :white_check_mark: |
-| Multi-Tool Agent | Accuracy | **100%** (all findings) | :white_check_mark: |
-| Multi-Tool Agent | Compression | **76.3%** | :white_check_mark: |
-
-**Why recall matters most**: For LLM applications, capturing all relevant information is critical. 98.2% recall means nearly all content is preserved — LLMs can answer questions accurately from compressed context.
-
-
-Run benchmarks yourself
+## Quick Start
```bash
-# Install with benchmark dependencies
-pip install "headroom-ai[evals,html]" datasets
-
-# Run HTML extraction benchmark (no API key needed)
-pytest tests/test_evals/test_html_oss_benchmarks.py::TestExtractionBenchmark -v -s
-
-# Run QA accuracy tests (requires OPENAI_API_KEY)
-pytest tests/test_evals/test_html_oss_benchmarks.py::TestQAAccuracyPreservation -v -s
+pip install "headroom-ai[all]"
```
-
+```python
+from headroom import compress
+
+messages = [
+ {"role": "user", "content": "What caused the outage?"},
+ {"role": "tool", "content": huge_log_output, "tool_call_id": "call_1"},
+]
+
+result = compress(messages, model="claude-sonnet-4-5-20250929")
+# result.messages → same format, 50-90% fewer tokens
+# result.tokens_saved → 8,000
+# result.compression_ratio → 0.87
+
+response = client.messages.create(model="claude-sonnet-4-5-20250929", messages=result.messages)
+```
+
+**Same answer. 87% fewer tokens.**
---
-## Multi-Tool Agent Test: Real Function Calling
+## How to Use Headroom
-**The setup:** An Agno agent with 4 tools (GitHub Issues, ArXiv Papers, Code Search, Database Logs) investigating a memory leak. Total tool output: 62,323 chars (~15,580 tokens).
+Headroom is a compression library, not just a proxy. Use whichever integration fits your stack:
-```python
-from agno.agent import Agent
-from agno.models.anthropic import Claude
-from headroom.integrations.agno import HeadroomAgnoModel
+| You have... | Use this | Code |
+|-------------|----------|------|
+| Any Python app | `compress()` | `result = compress(messages, model="gpt-4o")` |
+| LiteLLM | Callback | `litellm.callbacks = [HeadroomCallback()]` |
+| Python proxy (FastAPI) | ASGI Middleware | `app.add_middleware(CompressionMiddleware)` |
+| Claude Code / Cursor | Proxy | `ANTHROPIC_BASE_URL=http://localhost:8787 claude` |
+| Agno agents | Wrap model | `HeadroomAgnoModel(your_model)` |
+| LangChain | Wrap model | `HeadroomChatModel(your_llm)` *(experimental)* |
-# Wrap your model - that's it!
-base_model = Claude(id="claude-sonnet-4-20250514")
-model = HeadroomAgnoModel(wrapped_model=base_model)
-
-agent = Agent(model=model, tools=[search_github, search_arxiv, search_code, query_db])
-response = agent.run("Investigate the memory leak and recommend a fix")
-```
-
-**Results with Claude Sonnet:**
-
-| | Baseline | Headroom |
-|--|----------|----------|
-| Tokens sent to API | 15,662 | 6,100 |
-| API requests | 2 | 2 |
-| Tool calls | 4 | 4 |
-| Duration | 26.5s | 27.0s |
-
-**76.3% fewer tokens. Same comprehensive answer.**
-
-Both found: Issue #42 (memory leak), the `cleanup_worker()` fix, OutOfMemoryError logs (7.8GB/8GB, 847 threads), and relevant research papers.
-
-Run it yourself: `python examples/multi_tool_agent_test.py`
+**Already have a proxy?** You don't need another one. See the **[Integration Guide](docs/integration-guide.md)** for detailed setup with LiteLLM, ASGI middleware, and direct `compress()` usage.
---
## How It Works
-> Headroom optimizes LLM context *before* it hits the provider —
-> without changing your agent logic or tools.
-
-```mermaid
-flowchart LR
- User["Your App"]
- Entry["Headroom"]
- Transform["Context
Optimization"]
- LLM["LLM Provider"]
- Response["Response"]
-
- User --> Entry --> Transform --> LLM --> Response
+```
+Your App → Headroom → LLM Provider
+ ↓
+ CacheAligner: stabilizes prefix for KV cache hits
+ ContentRouter: routes to optimal compressor per content type
+ → SmartCrusher (JSON) | CodeCompressor (code) | LLMLingua (text)
+ IntelligentContext: score-based token fitting
+ CCR: stores originals for retrieval if LLM needs more
```
-### Inside Headroom
-
-```mermaid
-flowchart TB
-
-subgraph Pipeline["Transform Pipeline"]
- CA["Cache Aligner
Stabilizes dynamic tokens"]
- SC["Smart Crusher
Removes redundant tool output"]
- CM["Intelligent Context
Score-based token fitting"]
- CA --> SC --> CM
-end
-
-subgraph CCR["CCR: Compress-Cache-Retrieve"]
- Store[("Compressed
Store")]
- Tool["Retrieve Tool"]
- Tool <--> Store
-end
-
-LLM["LLM Provider"]
-
-CM --> LLM
-SC -. "Stores originals" .-> Store
-LLM -. "Requests full context
if needed" .-> Tool
-```
-
-> Headroom never throws data away.
-> It compresses aggressively and retrieves precisely.
-
-### What actually happens
-
-1. **Headroom intercepts context** — Tool outputs, logs, search results, and intermediate agent steps.
-
-2. **Dynamic content is stabilized** — Timestamps, UUIDs, request IDs are normalized so prompts cache cleanly.
-
-3. **Low-signal content is removed** — Repetitive or redundant data is crushed, not truncated.
-
-4. **Original data is preserved** — Full content is stored separately and retrieved *only if the LLM asks*.
-
-5. **Provider caches finally work** — Headroom aligns prompts so OpenAI, Anthropic, and Google caches actually hit.
-
-For deep technical details, see [Architecture Documentation](docs/ARCHITECTURE.md).
-
----
-
-## Why Headroom?
-
-- **Zero code changes** - works as a transparent proxy
-- **47-92% savings** - depends on your workload (tool-heavy = more savings)
-- **Image compression** - 40-90% reduction via trained ML router (OpenAI, Anthropic, Google)
-- **Reversible compression** - LLM retrieves original data via CCR
-- **Content-aware** - code, logs, JSON, images each handled optimally
-- **Provider caching** - automatic prefix optimization for cache hits
-- **Framework native** - LangChain, Agno, MCP, agents supported
-
----
-
-## 30-Second Quickstart
-
-### Option 1: Proxy (Zero Code Changes)
-
-```bash
-pip install "headroom-ai[all]" # Recommended for best performance
-headroom proxy --port 8787
-```
-
-> **Note:** First startup downloads ML models (~500MB) for optimal compression. This is a one-time download.
-
-**Dashboard:** Open http://localhost:8787/dashboard to see real-time stats, token savings, and request history.
-
-Point your tools at the proxy:
-
-```bash
-# Claude Code
-ANTHROPIC_BASE_URL=http://localhost:8787 claude
-
-# Any OpenAI-compatible client
-OPENAI_BASE_URL=http://localhost:8787/v1 cursor
-```
-
-**Enable Persistent Memory** - Claude remembers across conversations:
-
-```bash
-headroom proxy --memory
-```
-
-Memory auto-detects your provider (Anthropic, OpenAI, Gemini) and uses the appropriate format:
-- **Anthropic**: Uses native memory tool (`memory_20250818`) - works with Claude Code subscriptions
-- **OpenAI/Gemini/Others**: Uses function calling format
-- All providers share the same semantic vector store for search
-
-Set `x-headroom-user-id` header for per-user memory isolation (defaults to 'default').
-
-**Claude Code Subscription Users** - Use MCP for CCR (Compress-Cache-Retrieve):
-
-If you use Claude Code with a subscription (not API key), you need MCP to enable the `headroom_retrieve` tool:
-
-```bash
-# One-time setup
-pip install "headroom-ai[mcp]"
-headroom mcp install
-
-# Every time you code
-headroom proxy # Terminal 1
-claude # Terminal 2 - now has headroom_retrieve!
-```
-
-What this does:
-- Configures Claude Code to use Headroom's MCP server (`~/.claude/mcp.json`)
-- When the proxy compresses large tool outputs, Claude sees markers like `[47 items compressed... hash=abc123]`
-- Claude can call `headroom_retrieve` to get the full original content when needed
-
-Check your setup:
-```bash
-headroom mcp status
-```
-
-
-Why MCP for subscriptions?
-
-- **API users** can inject custom tools directly via the Messages API
-- **Subscription users** use Claude Code's built-in tool set and can't inject tools programmatically
-- **MCP** (Model Context Protocol) is Claude's official way to extend tools - it works with subscriptions
-
-The MCP server exposes `headroom_retrieve` so Claude can request uncompressed content when the compressed summary isn't enough.
-
-
-**Using AWS Bedrock, Google Vertex, or Azure?** Route through Headroom:
-
-```bash
-# AWS Bedrock - Terminal 1: Start proxy
-export AWS_ACCESS_KEY_ID="AKIA..."
-export AWS_SECRET_ACCESS_KEY="..."
-export AWS_REGION="us-east-1"
-headroom proxy --backend bedrock --region us-east-1
-
-# AWS Bedrock - Terminal 2: Run Claude Code
-export ANTHROPIC_API_KEY="sk-ant-dummy" # Any value works! Headroom ignores it.
-export ANTHROPIC_BASE_URL="http://localhost:8787"
-# IMPORTANT: Do NOT set CLAUDE_CODE_USE_BEDROCK=1 (Headroom handles Bedrock routing)
-claude
-```
-
-
-VS Code settings.json for Bedrock (click to expand)
-
-```json
-{
- "claudeCode.environmentVariables": [
- { "name": "ANTHROPIC_API_KEY", "value": "sk-ant-dummy" },
- { "name": "ANTHROPIC_BASE_URL", "value": "http://localhost:8787" },
- { "name": "AWS_ACCESS_KEY_ID", "value": "AKIA..." },
- { "name": "AWS_SECRET_ACCESS_KEY", "value": "..." },
- { "name": "AWS_REGION", "value": "us-east-1" }
- ]
-}
-```
-
-**Do NOT include** `CLAUDE_CODE_USE_BEDROCK` - Headroom handles the Bedrock routing.
-
-
-**Using OpenRouter?** Access 400+ models through a single API:
-
-```bash
-# OpenRouter - Terminal 1: Start proxy
-export OPENROUTER_API_KEY="sk-or-v1-..."
-headroom proxy --backend openrouter
-
-# OpenRouter - Terminal 2: Run your client
-export ANTHROPIC_API_KEY="sk-ant-dummy" # Any value works! Headroom ignores it.
-export ANTHROPIC_BASE_URL="http://localhost:8787"
-# Use OpenRouter model names in your requests:
-# - anthropic/claude-3.5-sonnet
-# - openai/gpt-4o
-# - google/gemini-pro
-# - meta-llama/llama-3-70b-instruct
-# See all models: https://openrouter.ai/models
-```
-
-```bash
-# Google Vertex AI
-headroom proxy --backend vertex_ai --region us-central1
-
-# Azure OpenAI
-headroom proxy --backend azure --region eastus
-```
-
-### Option 2: LangChain Integration
-
-```bash
-pip install "headroom-ai[langchain]"
-```
-
-```python
-from langchain_openai import ChatOpenAI
-from headroom.integrations import HeadroomChatModel
-
-# Wrap your model - that's it!
-llm = HeadroomChatModel(ChatOpenAI(model="gpt-4o"))
-
-# Use exactly like before
-response = llm.invoke("Hello!")
-```
-
-See the full [LangChain Integration Guide](docs/langchain.md) for memory, retrievers, agents, and more.
-
-### Option 3: Agno Integration
-
-```bash
-pip install "headroom-ai[agno]"
-```
-
-```python
-from agno.agent import Agent
-from agno.models.openai import OpenAIChat
-from headroom.integrations.agno import HeadroomAgnoModel
-
-# Wrap your model - that's it!
-model = HeadroomAgnoModel(OpenAIChat(id="gpt-4o"))
-agent = Agent(model=model)
-
-# Use exactly like before
-response = agent.run("Hello!")
-
-# Check savings
-print(f"Tokens saved: {model.total_tokens_saved}")
-```
-
-See the full [Agno Integration Guide](docs/agno.md) for hooks, multi-provider support, and more.
-
----
-
-## Framework Integrations
-
-| Framework | Integration | Docs |
-|-----------|-------------|------|
-| **LangChain** | `HeadroomChatModel`, memory, retrievers, agents | [Guide](docs/langchain.md) |
-| **Agno** | `HeadroomAgnoModel`, hooks, multi-provider | [Guide](docs/agno.md) |
-| **MCP** | Claude Code subscription support via `headroom mcp install` | [Guide](docs/mcp.md) |
-| **Any OpenAI Client** | Proxy server | [Guide](docs/proxy.md) |
-
----
-
-## Features
-
-| Feature | Description | Docs |
-|---------|-------------|------|
-| **Image Compression** | 40-90% token reduction for images via trained ML router | [Image Compression](docs/image-compression.md) |
-| **Memory** | Persistent memory across conversations (zero-latency inline extraction) | [Memory](docs/memory.md) |
-| **Universal Compression** | ML-based content detection + structure-preserving compression | [Compression](docs/compression.md) |
-| **SmartCrusher** | Compresses JSON tool outputs statistically | [Transforms](docs/transforms.md) |
-| **CacheAligner** | Stabilizes prefixes for provider caching | [Transforms](docs/transforms.md) |
-| **IntelligentContext** | Score-based context dropping with TOIN-learned importance | [Transforms](docs/transforms.md) |
-| **CCR** | Reversible compression with automatic retrieval | [CCR Guide](docs/ccr.md) |
-| **MCP Server** | Claude Code subscription support via `headroom mcp install` | [MCP Guide](docs/mcp.md) |
-| **LangChain** | Memory, retrievers, agents, streaming | [LangChain](docs/langchain.md) |
-| **Agno** | Agent framework integration with hooks | [Agno](docs/agno.md) |
-| **Text Utilities** | Opt-in compression for search/logs | [Text Compression](docs/text-compression.md) |
-| **LLMLingua-2** | ML-based 20x compression (opt-in) | [LLMLingua](docs/llmlingua.md) |
-| **Code-Aware** | AST-based code compression (tree-sitter) | [Transforms](docs/transforms.md) |
-| **Evals Framework** | Prove compression preserves accuracy (12+ datasets) | [Evals](headroom/evals/README.md) |
-
----
-
-## Evaluation Framework: Prove It Works
-
-Skeptical? Good. We built a comprehensive evaluation framework to **prove** compression preserves accuracy.
-
-```bash
-# Install evals
-pip install "headroom-ai[evals]"
-
-# Quick sanity check (5 samples)
-python -m headroom.evals quick
-
-# Run on real datasets
-python -m headroom.evals benchmark --dataset hotpotqa -n 100
-```
-
-### How Evals Work
-
-```
-Original Context ───► LLM ───► Response A
- │
-Compressed Context ─► LLM ───► Response B
- │
- Compare A vs B │
- ─────────────────
- F1 Score: 0.95
- Semantic Similarity: 0.97
- Ground Truth Match: ✓
- ─────────────────
- PASS: Accuracy preserved
-```
-
-### Available Datasets (12+)
-
-| Category | Datasets |
-|----------|----------|
-| **RAG** | HotpotQA, Natural Questions, TriviaQA, MS MARCO, SQuAD |
-| **Long Context** | LongBench (4K-128K tokens), NarrativeQA |
-| **Tool Use** | BFCL (function calling), ToolBench, Built-in samples |
-| **Code** | CodeSearchNet, HumanEval |
-
-### CI Integration
-
-```yaml
-# GitHub Actions
-- name: Run Compression Evals
- run: python -m headroom.evals quick -n 20
- env:
- ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
-```
-
-Exit code 0 if accuracy ≥ 90%, 1 otherwise.
-
-See the full [Evals Documentation](headroom/evals/README.md) for datasets, metrics, and programmatic API.
+Headroom never throws data away. It compresses aggressively and retrieves precisely.
---
## Verified Performance
-These numbers are from actual API calls, not estimates:
+| Scenario | Tokens Before | Tokens After | Savings |
+|----------|--------------|-------------|---------|
+| Code search (100 results) | 17,765 | 1,408 | **92%** |
+| SRE incident debugging | 65,694 | 5,118 | **92%** |
+| Codebase exploration | 78,502 | 41,254 | **47%** |
+| GitHub issue triage | 54,174 | 14,761 | **73%** |
-| Scenario | Before | After | Savings | Verified |
-|----------|--------|-------|---------|----------|
-| Code search (100 results) | 17,765 tokens | 1,408 tokens | 92% | Claude Sonnet |
-| SRE incident debugging | 65,694 tokens | 5,118 tokens | 92% | GPT-4o |
-| Codebase exploration | 78,502 tokens | 41,254 tokens | 47% | GPT-4o |
-| GitHub issue triage | 54,174 tokens | 14,761 tokens | 73% | GPT-4o |
-
-**Overhead**: ~1-5ms compression latency
-
-**When savings are highest**: Tool-heavy workloads (search, logs, database queries)
-**When savings are lowest**: Conversation-heavy workloads with minimal tool use
+**Overhead**: 1-5ms. **Accuracy**: [benchmarked](docs/benchmarks.md) across 12+ datasets.
---
-## Providers
+## Integrations
-| Provider | Token Counting | Cache Optimization |
-|----------|----------------|-------------------|
-| OpenAI | tiktoken (exact) | Automatic prefix caching |
-| Anthropic | Official API | cache_control blocks |
-| Google | Official API | Context caching |
-| Cohere | Official API | - |
-| Mistral | Official tokenizer | - |
-
-New models auto-supported via naming pattern detection.
+| Integration | Status | Docs |
+|-------------|--------|------|
+| `compress()` — one function | **Stable** | [Integration Guide](docs/integration-guide.md) |
+| LiteLLM callback | **Stable** | [Integration Guide](docs/integration-guide.md#litellm) |
+| ASGI middleware | **Stable** | [Integration Guide](docs/integration-guide.md#asgi-middleware) |
+| Proxy server | **Stable** | [Proxy Docs](docs/proxy.md) |
+| Agno | **Stable** | [Agno Guide](docs/agno.md) |
+| MCP (Claude Code) | **Stable** | [MCP Guide](docs/mcp.md) |
+| Strands | **Stable** | [Strands Guide](docs/strands.md) |
+| LangChain | **Experimental** | [LangChain Guide](docs/langchain.md) |
---
-## Safety Guarantees
+## Features
-- **Never removes human content** - user/assistant messages preserved
-- **Never breaks tool ordering** - tool calls and responses stay paired
-- **Parse failures are no-ops** - malformed content passes through unchanged
-- **Compression is reversible** - LLM retrieves original data via CCR
+| Feature | What it does |
+|---------|-------------|
+| **Content Router** | Auto-detects content type, routes to optimal compressor |
+| **SmartCrusher** | Statistically compresses JSON arrays (tool outputs, API responses) |
+| **CodeCompressor** | AST-aware code compression (Python, JS, Go, Rust, Java) |
+| **LLMLingua-2** | ML-based 20x text compression |
+| **CCR** | Reversible compression — LLM retrieves originals when needed |
+| **CacheAligner** | Stabilizes prefixes for provider KV cache hits |
+| **IntelligentContext** | Score-based context management with learned importance |
+| **Image Compression** | 40-90% token reduction via trained ML router |
+| **Memory** | Persistent memory across conversations |
+| **Compression Hooks** | Customize compression with pre/post hooks |
+| **Query Echo** | Re-injects user question after compressed data for better attention |
+
+---
+
+## Cloud Providers
+
+```bash
+headroom proxy --backend bedrock --region us-east-1 # AWS Bedrock
+headroom proxy --backend vertex_ai --region us-central1 # Google Vertex
+headroom proxy --backend azure # Azure OpenAI
+headroom proxy --backend openrouter # OpenRouter (400+ models)
+```
---
## Installation
```bash
-# Recommended: Install everything for best compression performance
-pip install "headroom-ai[all]"
-
-# Or install specific components
-pip install headroom-ai # SDK only
-pip install "headroom-ai[proxy]" # Proxy server
-pip install "headroom-ai[mcp]" # MCP server for Claude Code subscriptions
-pip install "headroom-ai[langchain]" # LangChain integration
-pip install "headroom-ai[agno]" # Agno agent framework
-pip install "headroom-ai[evals]" # Evaluation framework
-pip install "headroom-ai[code]" # AST-based code compression
-pip install "headroom-ai[llmlingua]" # ML-based compression
+pip install headroom-ai # Core library
+pip install "headroom-ai[all]" # Everything (recommended)
+pip install "headroom-ai[proxy]" # Proxy server
+pip install "headroom-ai[mcp]" # MCP for Claude Code
+pip install "headroom-ai[agno]" # Agno integration
+pip install "headroom-ai[langchain]" # LangChain (experimental)
+pip install "headroom-ai[evals]" # Evaluation framework
```
-**Requirements**: Python 3.10+
-
-> **First-time startup:** Headroom downloads ML models (~500MB) on first run for optimal compression. This is cached locally and only happens once.
+Python 3.10+
---
## Documentation
-| Guide | Description |
-|-------|-------------|
-| [Memory Guide](docs/memory.md) | Persistent memory for LLMs |
-| [Compression Guide](docs/compression.md) | Universal compression with ML detection |
-| [Evals Framework](headroom/evals/README.md) | Prove compression preserves accuracy |
-| [LangChain Integration](docs/langchain.md) | Full LangChain support |
-| [Agno Integration](docs/agno.md) | Full Agno agent framework support |
-| [SDK Guide](docs/sdk.md) | Fine-grained control |
-| [Proxy Guide](docs/proxy.md) | Production deployment |
-| [Configuration](docs/configuration.md) | All options |
+| | |
+|---|---|
+| [Integration Guide](docs/integration-guide.md) | LiteLLM, ASGI, compress(), proxy |
+| [Proxy Docs](docs/proxy.md) | Proxy server configuration |
+| [Architecture](docs/ARCHITECTURE.md) | How the pipeline works |
| [CCR Guide](docs/ccr.md) | Reversible compression |
-| [MCP Guide](docs/mcp.md) | Claude Code subscription support |
-| [Metrics](docs/metrics.md) | Monitoring |
-| [Troubleshooting](docs/troubleshooting.md) | Common issues |
-
----
-
-## Who's Using Headroom?
-
-> Add your project here! [Open a PR](https://github.com/chopratejas/headroom/pulls) or [start a discussion](https://github.com/chopratejas/headroom/discussions).
+| [Benchmarks](docs/benchmarks.md) | Accuracy validation |
+| [Evals Framework](headroom/evals/README.md) | Prove compression preserves accuracy |
+| [Memory](docs/memory.md) | Persistent memory |
+| [Agno](docs/agno.md) | Agno agent framework |
+| [MCP](docs/mcp.md) | Claude Code subscriptions |
+| [Configuration](docs/configuration.md) | All options |
---
## Contributing
```bash
-git clone https://github.com/chopratejas/headroom.git
-cd headroom
-pip install -e ".[dev]"
-pytest
+git clone https://github.com/chopratejas/headroom.git && cd headroom
+pip install -e ".[dev]" && pytest
```
-See [CONTRIBUTING.md](CONTRIBUTING.md) for details.
-
---
## License
-Apache License 2.0 - see [LICENSE](LICENSE).
-
----
-
-
- Built for the AI developer community
-
+Apache License 2.0 — see [LICENSE](LICENSE).
diff --git a/docs/integration-guide.md b/docs/integration-guide.md
new file mode 100644
index 000000000..659649c30
--- /dev/null
+++ b/docs/integration-guide.md
@@ -0,0 +1,292 @@
+# Integration Guide
+
+You don't need to run the Headroom proxy. Headroom is a compression library that works with **any** LLM client, proxy, or framework.
+
+## Pick Your Path
+
+| You have... | Use this | Setup |
+|-------------|----------|-------|
+| Any Python app | [`compress()`](#compress-function) | 2 lines |
+| LiteLLM | [LiteLLM callback](#litellm) | 1 line |
+| A Python proxy (FastAPI, custom) | [ASGI middleware](#asgi-middleware) | 1 line |
+| Claude Code / Cursor | [Headroom proxy](#proxy) | 1 env var |
+| Agno agents | [Agno integration](#agno) | Wrap model |
+| LangChain | [LangChain integration](#langchain) | Wrap model |
+| Non-Python app | [Headroom proxy](#proxy) | HTTP |
+
+---
+
+## compress() Function
+
+The simplest integration. Works with any LLM client.
+
+```python
+from headroom import compress
+
+# Before sending to your LLM:
+result = compress(messages, model="claude-sonnet-4-5-20250929")
+response = your_client.create(messages=result.messages) # Fewer tokens, same answer
+
+print(f"Saved {result.tokens_saved} tokens ({result.compression_ratio:.0%})")
+```
+
+### With Anthropic SDK
+
+```python
+from anthropic import Anthropic
+from headroom import compress
+
+client = Anthropic()
+messages = [
+ {"role": "user", "content": "What went wrong?"},
+ {"role": "assistant", "content": "Let me check.", "tool_use": [...]},
+ {"role": "user", "content": [{"type": "tool_result", "content": huge_json}]},
+]
+
+compressed = compress(messages, model="claude-sonnet-4-5-20250929")
+response = client.messages.create(
+ model="claude-sonnet-4-5-20250929",
+ messages=compressed.messages,
+ max_tokens=1000,
+)
+```
+
+### With OpenAI SDK
+
+```python
+from openai import OpenAI
+from headroom import compress
+
+client = OpenAI()
+messages = [
+ {"role": "user", "content": "Analyze these results"},
+ {"role": "tool", "content": big_json_output, "tool_call_id": "call_1"},
+]
+
+compressed = compress(messages, model="gpt-4o")
+response = client.chat.completions.create(
+ model="gpt-4o",
+ messages=compressed.messages,
+)
+```
+
+### With LiteLLM (direct)
+
+```python
+import litellm
+from headroom import compress
+
+messages = [...]
+compressed = compress(messages, model="bedrock/claude-sonnet")
+response = litellm.completion(model="bedrock/claude-sonnet", messages=compressed.messages)
+```
+
+### With any HTTP client
+
+```python
+import httpx
+from headroom import compress
+
+compressed = compress(messages, model="claude-sonnet-4-5-20250929")
+httpx.post("https://api.anthropic.com/v1/messages", json={
+ "model": "claude-sonnet-4-5-20250929",
+ "messages": compressed.messages,
+}, headers={"X-Api-Key": api_key, "anthropic-version": "2023-06-01"})
+```
+
+### What compress() returns
+
+```python
+result = compress(messages, model="gpt-4o")
+result.messages # list[dict] — compressed messages, same format as input
+result.tokens_before # int — original token count
+result.tokens_after # int — compressed token count
+result.tokens_saved # int — tokens removed
+result.compression_ratio # float — 0.0 (no savings) to 1.0 (100% removed)
+result.transforms_applied # list[str] — what ran (e.g., ["router:smart_crusher:0.35"])
+```
+
+---
+
+## LiteLLM
+
+If you're already using LiteLLM as your LLM gateway, add Headroom as a callback:
+
+```python
+import litellm
+from headroom.integrations.litellm_callback import HeadroomCallback
+
+litellm.callbacks = [HeadroomCallback()]
+
+# All calls now compressed automatically
+response = litellm.completion(model="gpt-4o", messages=[...])
+response = litellm.completion(model="bedrock/claude-sonnet", messages=[...])
+response = litellm.completion(model="azure/gpt-4o", messages=[...])
+```
+
+The callback compresses messages in LiteLLM's `pre_call_hook` before they're sent to the provider. Works with all 100+ LiteLLM-supported providers.
+
+### With LiteLLM Proxy
+
+If you run LiteLLM as a proxy server, use the ASGI middleware instead:
+
+```python
+# In your LiteLLM proxy startup
+from litellm.proxy.proxy_server import app
+from headroom.integrations.asgi import CompressionMiddleware
+
+app.add_middleware(CompressionMiddleware)
+```
+
+Or use the callback in your LiteLLM config:
+
+```yaml
+# litellm_config.yaml
+litellm_settings:
+ callbacks: ["headroom.integrations.litellm_callback.HeadroomCallback"]
+```
+
+---
+
+## ASGI Middleware
+
+Drop-in middleware for any ASGI application (FastAPI, Starlette, LiteLLM proxy, custom proxies).
+
+```python
+from headroom.integrations.asgi import CompressionMiddleware
+
+# FastAPI
+app = FastAPI()
+app.add_middleware(CompressionMiddleware)
+
+# Starlette
+app = Starlette(routes=[...])
+app.add_middleware(CompressionMiddleware)
+
+# LiteLLM proxy
+from litellm.proxy.proxy_server import app
+app.add_middleware(CompressionMiddleware)
+```
+
+The middleware intercepts POST requests to `/v1/messages`, `/v1/chat/completions`, `/v1/responses`, and `/chat/completions`. All other requests pass through untouched.
+
+Response headers include:
+- `x-headroom-compressed: true` — compression was applied
+- `x-headroom-tokens-saved: 1234` — tokens removed
+
+---
+
+## Proxy
+
+The Headroom proxy is a standalone HTTP server. Best for non-Python apps or tools that only support base URL configuration (Claude Code, Cursor).
+
+```bash
+pip install "headroom-ai[all]"
+headroom proxy --port 8787
+```
+
+```bash
+# Claude Code
+ANTHROPIC_BASE_URL=http://localhost:8787 claude
+
+# Cursor / Any OpenAI client
+OPENAI_BASE_URL=http://localhost:8787/v1 cursor
+```
+
+### With Cloud Providers
+
+```bash
+# AWS Bedrock
+headroom proxy --backend bedrock --region us-east-1
+
+# Google Vertex AI
+headroom proxy --backend vertex_ai --region us-central1
+
+# Azure OpenAI
+headroom proxy --backend azure
+
+# OpenRouter (400+ models)
+OPENROUTER_API_KEY=sk-or-... headroom proxy --backend openrouter
+```
+
+See [Proxy Documentation](proxy.md) for all options.
+
+---
+
+## Agno
+
+Full integration with the Agno agent framework.
+
+```python
+from agno.agent import Agent
+from agno.models.anthropic import Claude
+from headroom.integrations.agno import HeadroomAgnoModel
+
+model = HeadroomAgnoModel(Claude(id="claude-sonnet-4-20250514"))
+agent = Agent(model=model, tools=[your_tools])
+response = agent.run("Investigate the issue")
+
+print(f"Tokens saved: {model.total_tokens_saved}")
+```
+
+See [Agno Guide](agno.md) for hooks, multi-provider, and streaming.
+
+---
+
+## LangChain
+
+> **Experimental.** Core compression works. Streaming callbacks and async chains are still being tested.
+
+```python
+from langchain_openai import ChatOpenAI
+from headroom.integrations import HeadroomChatModel
+
+llm = HeadroomChatModel(ChatOpenAI(model="gpt-4o"))
+response = llm.invoke("Hello!")
+```
+
+See [LangChain Guide](langchain.md) for details and known limitations.
+
+---
+
+## Compression Hooks (Advanced)
+
+Customize compression behavior without modifying Headroom's code:
+
+```python
+from headroom import compress, CompressionHooks, CompressContext
+
+class MyHooks(CompressionHooks):
+ def pre_compress(self, messages, ctx):
+ # Modify messages before compression (dedup, filter, inject)
+ return messages
+
+ def compute_biases(self, messages, ctx):
+ # Per-message compression aggressiveness
+ # >1.0 = keep more, <1.0 = compress more
+ return {5: 1.5, 6: 0.5} # Keep message 5, compress message 6
+
+ def post_compress(self, event):
+ # Observe results (logging, analytics, learning)
+ print(f"Saved {event.tokens_saved} tokens")
+
+result = compress(messages, model="gpt-4o", hooks=MyHooks())
+```
+
+See [Architecture](ARCHITECTURE.md) for how hooks integrate with the pipeline.
+
+---
+
+## FAQ
+
+**Q: Does Headroom change the response format?**
+No. Your LLM returns the same response format. Headroom only modifies the input messages.
+
+**Q: What if compression removes something the LLM needs?**
+Headroom stores originals in CCR (Compress-Cache-Retrieve). The LLM can call `headroom_retrieve` to get full uncompressed content. Compression summaries tell the LLM what's available.
+
+**Q: Does it work with streaming?**
+Yes. Compression happens before the request is sent. Streaming responses are unaffected.
+
+**Q: How much latency does it add?**
+1-5ms for compression. The token savings typically save more time on the LLM side than compression adds.
diff --git a/headroom/compress.py b/headroom/compress.py
index 187c9ac31..bec533f73 100644
--- a/headroom/compress.py
+++ b/headroom/compress.py
@@ -183,17 +183,13 @@ def _get_pipeline() -> Any:
if _pipeline is not None:
return _pipeline
- from headroom.transforms import ContentRouter, SmartCrusher, TransformPipeline
+ from headroom.transforms import TransformPipeline
- _pipeline = TransformPipeline(
- transforms=[
- ContentRouter(),
- SmartCrusher(),
- ],
- # No provider needed — pipeline uses tokenizer registry which
- # auto-detects the right tokenizer per model:
- # OpenAI → tiktoken (exact), Anthropic → calibrated estimation,
- # Open models → HuggingFace (if installed)
- )
+ # Default pipeline: CacheAligner → ContentRouter → IntelligentContext
+ # CacheAligner: stabilizes prefix for provider KV cache hits
+ # ContentRouter: routes to the right compressor per content type
+ # (SmartCrusher for JSON, CodeCompressor for code, LLMLingua for text)
+ # IntelligentContext: enforces token limits with score-based dropping
+ _pipeline = TransformPipeline()
logger.debug("Headroom compression pipeline initialized")
return _pipeline