From 3ebcd89d4663da4d496b2e9a717568ee224f92e9 Mon Sep 17 00:00:00 2001 From: chopratejas Date: Thu, 19 Feb 2026 10:17:33 -0800 Subject: [PATCH] Rewrite README + add Integration Guide MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit README: 694 → 203 lines. Crisp, scannable, links to docs. - compress() as the hero quickstart (not proxy) - Integration table: compress(), LiteLLM, ASGI, proxy, Agno, LangChain - LangChain marked as experimental - "Already have a proxy?" callout linking to Integration Guide - Architecture: ContentRouter (not SmartCrusher) as the primary compressor New: docs/integration-guide.md - Detailed setup for every integration path - compress() with Anthropic, OpenAI, LiteLLM, raw HTTP - LiteLLM callback + LiteLLM proxy ASGI middleware - ASGI middleware for any FastAPI/Starlette app - Compression hooks for advanced customization - FAQ section Fix: compress() uses default pipeline (CacheAligner + ContentRouter + IntelligentContext) instead of manually specifying SmartCrusher. --- README.md | 615 +++++++------------------------------- docs/integration-guide.md | 292 ++++++++++++++++++ headroom/compress.py | 18 +- 3 files changed, 403 insertions(+), 522 deletions(-) create mode 100644 docs/integration-guide.md diff --git a/README.md b/README.md index 51d9a0028..c8130f857 100644 --- a/README.md +++ b/README.md @@ -29,7 +29,6 @@

- --- ## Demo @@ -40,570 +39,164 @@ --- -## Does It Actually Work? A Real Test - -**The setup:** 100 production log entries. One critical error buried at position 67. - -
-BEFORE: 100 log entries (18,952 chars) - click to expand - -```json -[ - {"timestamp": "2024-12-15T00:00:00Z", "level": "INFO", "service": "api-gateway", "message": "Request processed successfully - latency=50ms", "request_id": "req-000000", "status_code": 200}, - {"timestamp": "2024-12-15T01:01:00Z", "level": "INFO", "service": "user-service", "message": "Request processed successfully - latency=51ms", "request_id": "req-000001", "status_code": 200}, - {"timestamp": "2024-12-15T02:02:00Z", "level": "INFO", "service": "inventory", "message": "Request processed successfully - latency=52ms", "request_id": "req-000002", "status_code": 200}, - // ... 64 more INFO entries ... - {"timestamp": "2024-12-15T03:47:23Z", "level": "FATAL", "service": "payment-gateway", "message": "Connection pool exhausted", "error_code": "PG-5523", "resolution": "Increase max_connections to 500 in config/database.yml", "affected_transactions": 1847}, - // ... 32 more INFO entries ... -] -``` -
- -**AFTER:** Headroom compresses to 6 entries (1,155 chars): - -```json -[ - {"timestamp": "2024-12-15T00:00:00Z", "level": "INFO", "service": "api-gateway", ...}, - {"timestamp": "2024-12-15T01:01:00Z", "level": "INFO", "service": "user-service", ...}, - {"timestamp": "2024-12-15T02:02:00Z", "level": "INFO", "service": "inventory", ...}, - {"timestamp": "2024-12-15T03:47:23Z", "level": "FATAL", "service": "payment-gateway", "error_code": "PG-5523", "resolution": "Increase max_connections to 500 in config/database.yml", "affected_transactions": 1847}, - {"timestamp": "2024-12-15T02:38:00Z", "level": "INFO", "service": "inventory", ...}, - {"timestamp": "2024-12-15T03:39:00Z", "level": "INFO", "service": "auth", ...} -] -``` - -**What happened:** First 3 items + the FATAL error + last 2 items. The critical error at position 67 was automatically preserved. - ---- - -**The question we asked Claude:** "What caused the outage? What's the error code? What's the fix?" - -| | Baseline | Headroom | -|--|----------|----------| -| Input tokens | 10,144 | 1,260 | -| Correct answers | **4/4** | **4/4** | - -Both responses: *"payment-gateway service, error PG-5523, fix: Increase max_connections to 500, 1,847 transactions affected"* - -**87.6% fewer tokens. Same answer.** - -Run it yourself: `python examples/needle_in_haystack_test.py` - ---- - -## Accuracy Benchmarks - -> **Headroom's guarantee: compress without losing accuracy.** - -We validate against established open-source benchmarks. Full methodology and reproducible tests: [Benchmarks Documentation](https://chopratejas.github.io/headroom/benchmarks/) - -| Benchmark | Metric | Result | Status | -|-----------|--------|--------|--------| -| [Scrapinghub Article Extraction](https://huggingface.co/datasets/allenai/scrapinghub-article-extraction-benchmark) | F1 Score | **0.919** (baseline: 0.958) | :white_check_mark: | -| [Scrapinghub Article Extraction](https://huggingface.co/datasets/allenai/scrapinghub-article-extraction-benchmark) | Recall | **98.2%** | :white_check_mark: | -| [Scrapinghub Article Extraction](https://huggingface.co/datasets/allenai/scrapinghub-article-extraction-benchmark) | Compression | **94.9%** | :white_check_mark: | -| SmartCrusher (JSON) | Accuracy | **100%** (4/4 correct) | :white_check_mark: | -| SmartCrusher (JSON) | Compression | **87.6%** | :white_check_mark: | -| Multi-Tool Agent | Accuracy | **100%** (all findings) | :white_check_mark: | -| Multi-Tool Agent | Compression | **76.3%** | :white_check_mark: | - -**Why recall matters most**: For LLM applications, capturing all relevant information is critical. 98.2% recall means nearly all content is preserved — LLMs can answer questions accurately from compressed context. - -
-Run benchmarks yourself +## Quick Start ```bash -# Install with benchmark dependencies -pip install "headroom-ai[evals,html]" datasets - -# Run HTML extraction benchmark (no API key needed) -pytest tests/test_evals/test_html_oss_benchmarks.py::TestExtractionBenchmark -v -s - -# Run QA accuracy tests (requires OPENAI_API_KEY) -pytest tests/test_evals/test_html_oss_benchmarks.py::TestQAAccuracyPreservation -v -s +pip install "headroom-ai[all]" ``` -
+```python +from headroom import compress + +messages = [ + {"role": "user", "content": "What caused the outage?"}, + {"role": "tool", "content": huge_log_output, "tool_call_id": "call_1"}, +] + +result = compress(messages, model="claude-sonnet-4-5-20250929") +# result.messages → same format, 50-90% fewer tokens +# result.tokens_saved → 8,000 +# result.compression_ratio → 0.87 + +response = client.messages.create(model="claude-sonnet-4-5-20250929", messages=result.messages) +``` + +**Same answer. 87% fewer tokens.** --- -## Multi-Tool Agent Test: Real Function Calling +## How to Use Headroom -**The setup:** An Agno agent with 4 tools (GitHub Issues, ArXiv Papers, Code Search, Database Logs) investigating a memory leak. Total tool output: 62,323 chars (~15,580 tokens). +Headroom is a compression library, not just a proxy. Use whichever integration fits your stack: -```python -from agno.agent import Agent -from agno.models.anthropic import Claude -from headroom.integrations.agno import HeadroomAgnoModel +| You have... | Use this | Code | +|-------------|----------|------| +| Any Python app | `compress()` | `result = compress(messages, model="gpt-4o")` | +| LiteLLM | Callback | `litellm.callbacks = [HeadroomCallback()]` | +| Python proxy (FastAPI) | ASGI Middleware | `app.add_middleware(CompressionMiddleware)` | +| Claude Code / Cursor | Proxy | `ANTHROPIC_BASE_URL=http://localhost:8787 claude` | +| Agno agents | Wrap model | `HeadroomAgnoModel(your_model)` | +| LangChain | Wrap model | `HeadroomChatModel(your_llm)` *(experimental)* | -# Wrap your model - that's it! -base_model = Claude(id="claude-sonnet-4-20250514") -model = HeadroomAgnoModel(wrapped_model=base_model) - -agent = Agent(model=model, tools=[search_github, search_arxiv, search_code, query_db]) -response = agent.run("Investigate the memory leak and recommend a fix") -``` - -**Results with Claude Sonnet:** - -| | Baseline | Headroom | -|--|----------|----------| -| Tokens sent to API | 15,662 | 6,100 | -| API requests | 2 | 2 | -| Tool calls | 4 | 4 | -| Duration | 26.5s | 27.0s | - -**76.3% fewer tokens. Same comprehensive answer.** - -Both found: Issue #42 (memory leak), the `cleanup_worker()` fix, OutOfMemoryError logs (7.8GB/8GB, 847 threads), and relevant research papers. - -Run it yourself: `python examples/multi_tool_agent_test.py` +**Already have a proxy?** You don't need another one. See the **[Integration Guide](docs/integration-guide.md)** for detailed setup with LiteLLM, ASGI middleware, and direct `compress()` usage. --- ## How It Works -> Headroom optimizes LLM context *before* it hits the provider — -> without changing your agent logic or tools. - -```mermaid -flowchart LR - User["Your App"] - Entry["Headroom"] - Transform["Context
Optimization"] - LLM["LLM Provider"] - Response["Response"] - - User --> Entry --> Transform --> LLM --> Response +``` +Your App → Headroom → LLM Provider + ↓ + CacheAligner: stabilizes prefix for KV cache hits + ContentRouter: routes to optimal compressor per content type + → SmartCrusher (JSON) | CodeCompressor (code) | LLMLingua (text) + IntelligentContext: score-based token fitting + CCR: stores originals for retrieval if LLM needs more ``` -### Inside Headroom - -```mermaid -flowchart TB - -subgraph Pipeline["Transform Pipeline"] - CA["Cache Aligner
Stabilizes dynamic tokens"] - SC["Smart Crusher
Removes redundant tool output"] - CM["Intelligent Context
Score-based token fitting"] - CA --> SC --> CM -end - -subgraph CCR["CCR: Compress-Cache-Retrieve"] - Store[("Compressed
Store")] - Tool["Retrieve Tool"] - Tool <--> Store -end - -LLM["LLM Provider"] - -CM --> LLM -SC -. "Stores originals" .-> Store -LLM -. "Requests full context
if needed" .-> Tool -``` - -> Headroom never throws data away. -> It compresses aggressively and retrieves precisely. - -### What actually happens - -1. **Headroom intercepts context** — Tool outputs, logs, search results, and intermediate agent steps. - -2. **Dynamic content is stabilized** — Timestamps, UUIDs, request IDs are normalized so prompts cache cleanly. - -3. **Low-signal content is removed** — Repetitive or redundant data is crushed, not truncated. - -4. **Original data is preserved** — Full content is stored separately and retrieved *only if the LLM asks*. - -5. **Provider caches finally work** — Headroom aligns prompts so OpenAI, Anthropic, and Google caches actually hit. - -For deep technical details, see [Architecture Documentation](docs/ARCHITECTURE.md). - ---- - -## Why Headroom? - -- **Zero code changes** - works as a transparent proxy -- **47-92% savings** - depends on your workload (tool-heavy = more savings) -- **Image compression** - 40-90% reduction via trained ML router (OpenAI, Anthropic, Google) -- **Reversible compression** - LLM retrieves original data via CCR -- **Content-aware** - code, logs, JSON, images each handled optimally -- **Provider caching** - automatic prefix optimization for cache hits -- **Framework native** - LangChain, Agno, MCP, agents supported - ---- - -## 30-Second Quickstart - -### Option 1: Proxy (Zero Code Changes) - -```bash -pip install "headroom-ai[all]" # Recommended for best performance -headroom proxy --port 8787 -``` - -> **Note:** First startup downloads ML models (~500MB) for optimal compression. This is a one-time download. - -**Dashboard:** Open http://localhost:8787/dashboard to see real-time stats, token savings, and request history. - -Point your tools at the proxy: - -```bash -# Claude Code -ANTHROPIC_BASE_URL=http://localhost:8787 claude - -# Any OpenAI-compatible client -OPENAI_BASE_URL=http://localhost:8787/v1 cursor -``` - -**Enable Persistent Memory** - Claude remembers across conversations: - -```bash -headroom proxy --memory -``` - -Memory auto-detects your provider (Anthropic, OpenAI, Gemini) and uses the appropriate format: -- **Anthropic**: Uses native memory tool (`memory_20250818`) - works with Claude Code subscriptions -- **OpenAI/Gemini/Others**: Uses function calling format -- All providers share the same semantic vector store for search - -Set `x-headroom-user-id` header for per-user memory isolation (defaults to 'default'). - -**Claude Code Subscription Users** - Use MCP for CCR (Compress-Cache-Retrieve): - -If you use Claude Code with a subscription (not API key), you need MCP to enable the `headroom_retrieve` tool: - -```bash -# One-time setup -pip install "headroom-ai[mcp]" -headroom mcp install - -# Every time you code -headroom proxy # Terminal 1 -claude # Terminal 2 - now has headroom_retrieve! -``` - -What this does: -- Configures Claude Code to use Headroom's MCP server (`~/.claude/mcp.json`) -- When the proxy compresses large tool outputs, Claude sees markers like `[47 items compressed... hash=abc123]` -- Claude can call `headroom_retrieve` to get the full original content when needed - -Check your setup: -```bash -headroom mcp status -``` - -
-Why MCP for subscriptions? - -- **API users** can inject custom tools directly via the Messages API -- **Subscription users** use Claude Code's built-in tool set and can't inject tools programmatically -- **MCP** (Model Context Protocol) is Claude's official way to extend tools - it works with subscriptions - -The MCP server exposes `headroom_retrieve` so Claude can request uncompressed content when the compressed summary isn't enough. -
- -**Using AWS Bedrock, Google Vertex, or Azure?** Route through Headroom: - -```bash -# AWS Bedrock - Terminal 1: Start proxy -export AWS_ACCESS_KEY_ID="AKIA..." -export AWS_SECRET_ACCESS_KEY="..." -export AWS_REGION="us-east-1" -headroom proxy --backend bedrock --region us-east-1 - -# AWS Bedrock - Terminal 2: Run Claude Code -export ANTHROPIC_API_KEY="sk-ant-dummy" # Any value works! Headroom ignores it. -export ANTHROPIC_BASE_URL="http://localhost:8787" -# IMPORTANT: Do NOT set CLAUDE_CODE_USE_BEDROCK=1 (Headroom handles Bedrock routing) -claude -``` - -
-VS Code settings.json for Bedrock (click to expand) - -```json -{ - "claudeCode.environmentVariables": [ - { "name": "ANTHROPIC_API_KEY", "value": "sk-ant-dummy" }, - { "name": "ANTHROPIC_BASE_URL", "value": "http://localhost:8787" }, - { "name": "AWS_ACCESS_KEY_ID", "value": "AKIA..." }, - { "name": "AWS_SECRET_ACCESS_KEY", "value": "..." }, - { "name": "AWS_REGION", "value": "us-east-1" } - ] -} -``` - -**Do NOT include** `CLAUDE_CODE_USE_BEDROCK` - Headroom handles the Bedrock routing. -
- -**Using OpenRouter?** Access 400+ models through a single API: - -```bash -# OpenRouter - Terminal 1: Start proxy -export OPENROUTER_API_KEY="sk-or-v1-..." -headroom proxy --backend openrouter - -# OpenRouter - Terminal 2: Run your client -export ANTHROPIC_API_KEY="sk-ant-dummy" # Any value works! Headroom ignores it. -export ANTHROPIC_BASE_URL="http://localhost:8787" -# Use OpenRouter model names in your requests: -# - anthropic/claude-3.5-sonnet -# - openai/gpt-4o -# - google/gemini-pro -# - meta-llama/llama-3-70b-instruct -# See all models: https://openrouter.ai/models -``` - -```bash -# Google Vertex AI -headroom proxy --backend vertex_ai --region us-central1 - -# Azure OpenAI -headroom proxy --backend azure --region eastus -``` - -### Option 2: LangChain Integration - -```bash -pip install "headroom-ai[langchain]" -``` - -```python -from langchain_openai import ChatOpenAI -from headroom.integrations import HeadroomChatModel - -# Wrap your model - that's it! -llm = HeadroomChatModel(ChatOpenAI(model="gpt-4o")) - -# Use exactly like before -response = llm.invoke("Hello!") -``` - -See the full [LangChain Integration Guide](docs/langchain.md) for memory, retrievers, agents, and more. - -### Option 3: Agno Integration - -```bash -pip install "headroom-ai[agno]" -``` - -```python -from agno.agent import Agent -from agno.models.openai import OpenAIChat -from headroom.integrations.agno import HeadroomAgnoModel - -# Wrap your model - that's it! -model = HeadroomAgnoModel(OpenAIChat(id="gpt-4o")) -agent = Agent(model=model) - -# Use exactly like before -response = agent.run("Hello!") - -# Check savings -print(f"Tokens saved: {model.total_tokens_saved}") -``` - -See the full [Agno Integration Guide](docs/agno.md) for hooks, multi-provider support, and more. - ---- - -## Framework Integrations - -| Framework | Integration | Docs | -|-----------|-------------|------| -| **LangChain** | `HeadroomChatModel`, memory, retrievers, agents | [Guide](docs/langchain.md) | -| **Agno** | `HeadroomAgnoModel`, hooks, multi-provider | [Guide](docs/agno.md) | -| **MCP** | Claude Code subscription support via `headroom mcp install` | [Guide](docs/mcp.md) | -| **Any OpenAI Client** | Proxy server | [Guide](docs/proxy.md) | - ---- - -## Features - -| Feature | Description | Docs | -|---------|-------------|------| -| **Image Compression** | 40-90% token reduction for images via trained ML router | [Image Compression](docs/image-compression.md) | -| **Memory** | Persistent memory across conversations (zero-latency inline extraction) | [Memory](docs/memory.md) | -| **Universal Compression** | ML-based content detection + structure-preserving compression | [Compression](docs/compression.md) | -| **SmartCrusher** | Compresses JSON tool outputs statistically | [Transforms](docs/transforms.md) | -| **CacheAligner** | Stabilizes prefixes for provider caching | [Transforms](docs/transforms.md) | -| **IntelligentContext** | Score-based context dropping with TOIN-learned importance | [Transforms](docs/transforms.md) | -| **CCR** | Reversible compression with automatic retrieval | [CCR Guide](docs/ccr.md) | -| **MCP Server** | Claude Code subscription support via `headroom mcp install` | [MCP Guide](docs/mcp.md) | -| **LangChain** | Memory, retrievers, agents, streaming | [LangChain](docs/langchain.md) | -| **Agno** | Agent framework integration with hooks | [Agno](docs/agno.md) | -| **Text Utilities** | Opt-in compression for search/logs | [Text Compression](docs/text-compression.md) | -| **LLMLingua-2** | ML-based 20x compression (opt-in) | [LLMLingua](docs/llmlingua.md) | -| **Code-Aware** | AST-based code compression (tree-sitter) | [Transforms](docs/transforms.md) | -| **Evals Framework** | Prove compression preserves accuracy (12+ datasets) | [Evals](headroom/evals/README.md) | - ---- - -## Evaluation Framework: Prove It Works - -Skeptical? Good. We built a comprehensive evaluation framework to **prove** compression preserves accuracy. - -```bash -# Install evals -pip install "headroom-ai[evals]" - -# Quick sanity check (5 samples) -python -m headroom.evals quick - -# Run on real datasets -python -m headroom.evals benchmark --dataset hotpotqa -n 100 -``` - -### How Evals Work - -``` -Original Context ───► LLM ───► Response A - │ -Compressed Context ─► LLM ───► Response B - │ - Compare A vs B │ - ───────────────── - F1 Score: 0.95 - Semantic Similarity: 0.97 - Ground Truth Match: ✓ - ───────────────── - PASS: Accuracy preserved -``` - -### Available Datasets (12+) - -| Category | Datasets | -|----------|----------| -| **RAG** | HotpotQA, Natural Questions, TriviaQA, MS MARCO, SQuAD | -| **Long Context** | LongBench (4K-128K tokens), NarrativeQA | -| **Tool Use** | BFCL (function calling), ToolBench, Built-in samples | -| **Code** | CodeSearchNet, HumanEval | - -### CI Integration - -```yaml -# GitHub Actions -- name: Run Compression Evals - run: python -m headroom.evals quick -n 20 - env: - ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} -``` - -Exit code 0 if accuracy ≥ 90%, 1 otherwise. - -See the full [Evals Documentation](headroom/evals/README.md) for datasets, metrics, and programmatic API. +Headroom never throws data away. It compresses aggressively and retrieves precisely. --- ## Verified Performance -These numbers are from actual API calls, not estimates: +| Scenario | Tokens Before | Tokens After | Savings | +|----------|--------------|-------------|---------| +| Code search (100 results) | 17,765 | 1,408 | **92%** | +| SRE incident debugging | 65,694 | 5,118 | **92%** | +| Codebase exploration | 78,502 | 41,254 | **47%** | +| GitHub issue triage | 54,174 | 14,761 | **73%** | -| Scenario | Before | After | Savings | Verified | -|----------|--------|-------|---------|----------| -| Code search (100 results) | 17,765 tokens | 1,408 tokens | 92% | Claude Sonnet | -| SRE incident debugging | 65,694 tokens | 5,118 tokens | 92% | GPT-4o | -| Codebase exploration | 78,502 tokens | 41,254 tokens | 47% | GPT-4o | -| GitHub issue triage | 54,174 tokens | 14,761 tokens | 73% | GPT-4o | - -**Overhead**: ~1-5ms compression latency - -**When savings are highest**: Tool-heavy workloads (search, logs, database queries) -**When savings are lowest**: Conversation-heavy workloads with minimal tool use +**Overhead**: 1-5ms. **Accuracy**: [benchmarked](docs/benchmarks.md) across 12+ datasets. --- -## Providers +## Integrations -| Provider | Token Counting | Cache Optimization | -|----------|----------------|-------------------| -| OpenAI | tiktoken (exact) | Automatic prefix caching | -| Anthropic | Official API | cache_control blocks | -| Google | Official API | Context caching | -| Cohere | Official API | - | -| Mistral | Official tokenizer | - | - -New models auto-supported via naming pattern detection. +| Integration | Status | Docs | +|-------------|--------|------| +| `compress()` — one function | **Stable** | [Integration Guide](docs/integration-guide.md) | +| LiteLLM callback | **Stable** | [Integration Guide](docs/integration-guide.md#litellm) | +| ASGI middleware | **Stable** | [Integration Guide](docs/integration-guide.md#asgi-middleware) | +| Proxy server | **Stable** | [Proxy Docs](docs/proxy.md) | +| Agno | **Stable** | [Agno Guide](docs/agno.md) | +| MCP (Claude Code) | **Stable** | [MCP Guide](docs/mcp.md) | +| Strands | **Stable** | [Strands Guide](docs/strands.md) | +| LangChain | **Experimental** | [LangChain Guide](docs/langchain.md) | --- -## Safety Guarantees +## Features -- **Never removes human content** - user/assistant messages preserved -- **Never breaks tool ordering** - tool calls and responses stay paired -- **Parse failures are no-ops** - malformed content passes through unchanged -- **Compression is reversible** - LLM retrieves original data via CCR +| Feature | What it does | +|---------|-------------| +| **Content Router** | Auto-detects content type, routes to optimal compressor | +| **SmartCrusher** | Statistically compresses JSON arrays (tool outputs, API responses) | +| **CodeCompressor** | AST-aware code compression (Python, JS, Go, Rust, Java) | +| **LLMLingua-2** | ML-based 20x text compression | +| **CCR** | Reversible compression — LLM retrieves originals when needed | +| **CacheAligner** | Stabilizes prefixes for provider KV cache hits | +| **IntelligentContext** | Score-based context management with learned importance | +| **Image Compression** | 40-90% token reduction via trained ML router | +| **Memory** | Persistent memory across conversations | +| **Compression Hooks** | Customize compression with pre/post hooks | +| **Query Echo** | Re-injects user question after compressed data for better attention | + +--- + +## Cloud Providers + +```bash +headroom proxy --backend bedrock --region us-east-1 # AWS Bedrock +headroom proxy --backend vertex_ai --region us-central1 # Google Vertex +headroom proxy --backend azure # Azure OpenAI +headroom proxy --backend openrouter # OpenRouter (400+ models) +``` --- ## Installation ```bash -# Recommended: Install everything for best compression performance -pip install "headroom-ai[all]" - -# Or install specific components -pip install headroom-ai # SDK only -pip install "headroom-ai[proxy]" # Proxy server -pip install "headroom-ai[mcp]" # MCP server for Claude Code subscriptions -pip install "headroom-ai[langchain]" # LangChain integration -pip install "headroom-ai[agno]" # Agno agent framework -pip install "headroom-ai[evals]" # Evaluation framework -pip install "headroom-ai[code]" # AST-based code compression -pip install "headroom-ai[llmlingua]" # ML-based compression +pip install headroom-ai # Core library +pip install "headroom-ai[all]" # Everything (recommended) +pip install "headroom-ai[proxy]" # Proxy server +pip install "headroom-ai[mcp]" # MCP for Claude Code +pip install "headroom-ai[agno]" # Agno integration +pip install "headroom-ai[langchain]" # LangChain (experimental) +pip install "headroom-ai[evals]" # Evaluation framework ``` -**Requirements**: Python 3.10+ - -> **First-time startup:** Headroom downloads ML models (~500MB) on first run for optimal compression. This is cached locally and only happens once. +Python 3.10+ --- ## Documentation -| Guide | Description | -|-------|-------------| -| [Memory Guide](docs/memory.md) | Persistent memory for LLMs | -| [Compression Guide](docs/compression.md) | Universal compression with ML detection | -| [Evals Framework](headroom/evals/README.md) | Prove compression preserves accuracy | -| [LangChain Integration](docs/langchain.md) | Full LangChain support | -| [Agno Integration](docs/agno.md) | Full Agno agent framework support | -| [SDK Guide](docs/sdk.md) | Fine-grained control | -| [Proxy Guide](docs/proxy.md) | Production deployment | -| [Configuration](docs/configuration.md) | All options | +| | | +|---|---| +| [Integration Guide](docs/integration-guide.md) | LiteLLM, ASGI, compress(), proxy | +| [Proxy Docs](docs/proxy.md) | Proxy server configuration | +| [Architecture](docs/ARCHITECTURE.md) | How the pipeline works | | [CCR Guide](docs/ccr.md) | Reversible compression | -| [MCP Guide](docs/mcp.md) | Claude Code subscription support | -| [Metrics](docs/metrics.md) | Monitoring | -| [Troubleshooting](docs/troubleshooting.md) | Common issues | - ---- - -## Who's Using Headroom? - -> Add your project here! [Open a PR](https://github.com/chopratejas/headroom/pulls) or [start a discussion](https://github.com/chopratejas/headroom/discussions). +| [Benchmarks](docs/benchmarks.md) | Accuracy validation | +| [Evals Framework](headroom/evals/README.md) | Prove compression preserves accuracy | +| [Memory](docs/memory.md) | Persistent memory | +| [Agno](docs/agno.md) | Agno agent framework | +| [MCP](docs/mcp.md) | Claude Code subscriptions | +| [Configuration](docs/configuration.md) | All options | --- ## Contributing ```bash -git clone https://github.com/chopratejas/headroom.git -cd headroom -pip install -e ".[dev]" -pytest +git clone https://github.com/chopratejas/headroom.git && cd headroom +pip install -e ".[dev]" && pytest ``` -See [CONTRIBUTING.md](CONTRIBUTING.md) for details. - --- ## License -Apache License 2.0 - see [LICENSE](LICENSE). - ---- - -

- Built for the AI developer community -

+Apache License 2.0 — see [LICENSE](LICENSE). diff --git a/docs/integration-guide.md b/docs/integration-guide.md new file mode 100644 index 000000000..659649c30 --- /dev/null +++ b/docs/integration-guide.md @@ -0,0 +1,292 @@ +# Integration Guide + +You don't need to run the Headroom proxy. Headroom is a compression library that works with **any** LLM client, proxy, or framework. + +## Pick Your Path + +| You have... | Use this | Setup | +|-------------|----------|-------| +| Any Python app | [`compress()`](#compress-function) | 2 lines | +| LiteLLM | [LiteLLM callback](#litellm) | 1 line | +| A Python proxy (FastAPI, custom) | [ASGI middleware](#asgi-middleware) | 1 line | +| Claude Code / Cursor | [Headroom proxy](#proxy) | 1 env var | +| Agno agents | [Agno integration](#agno) | Wrap model | +| LangChain | [LangChain integration](#langchain) | Wrap model | +| Non-Python app | [Headroom proxy](#proxy) | HTTP | + +--- + +## compress() Function + +The simplest integration. Works with any LLM client. + +```python +from headroom import compress + +# Before sending to your LLM: +result = compress(messages, model="claude-sonnet-4-5-20250929") +response = your_client.create(messages=result.messages) # Fewer tokens, same answer + +print(f"Saved {result.tokens_saved} tokens ({result.compression_ratio:.0%})") +``` + +### With Anthropic SDK + +```python +from anthropic import Anthropic +from headroom import compress + +client = Anthropic() +messages = [ + {"role": "user", "content": "What went wrong?"}, + {"role": "assistant", "content": "Let me check.", "tool_use": [...]}, + {"role": "user", "content": [{"type": "tool_result", "content": huge_json}]}, +] + +compressed = compress(messages, model="claude-sonnet-4-5-20250929") +response = client.messages.create( + model="claude-sonnet-4-5-20250929", + messages=compressed.messages, + max_tokens=1000, +) +``` + +### With OpenAI SDK + +```python +from openai import OpenAI +from headroom import compress + +client = OpenAI() +messages = [ + {"role": "user", "content": "Analyze these results"}, + {"role": "tool", "content": big_json_output, "tool_call_id": "call_1"}, +] + +compressed = compress(messages, model="gpt-4o") +response = client.chat.completions.create( + model="gpt-4o", + messages=compressed.messages, +) +``` + +### With LiteLLM (direct) + +```python +import litellm +from headroom import compress + +messages = [...] +compressed = compress(messages, model="bedrock/claude-sonnet") +response = litellm.completion(model="bedrock/claude-sonnet", messages=compressed.messages) +``` + +### With any HTTP client + +```python +import httpx +from headroom import compress + +compressed = compress(messages, model="claude-sonnet-4-5-20250929") +httpx.post("https://api.anthropic.com/v1/messages", json={ + "model": "claude-sonnet-4-5-20250929", + "messages": compressed.messages, +}, headers={"X-Api-Key": api_key, "anthropic-version": "2023-06-01"}) +``` + +### What compress() returns + +```python +result = compress(messages, model="gpt-4o") +result.messages # list[dict] — compressed messages, same format as input +result.tokens_before # int — original token count +result.tokens_after # int — compressed token count +result.tokens_saved # int — tokens removed +result.compression_ratio # float — 0.0 (no savings) to 1.0 (100% removed) +result.transforms_applied # list[str] — what ran (e.g., ["router:smart_crusher:0.35"]) +``` + +--- + +## LiteLLM + +If you're already using LiteLLM as your LLM gateway, add Headroom as a callback: + +```python +import litellm +from headroom.integrations.litellm_callback import HeadroomCallback + +litellm.callbacks = [HeadroomCallback()] + +# All calls now compressed automatically +response = litellm.completion(model="gpt-4o", messages=[...]) +response = litellm.completion(model="bedrock/claude-sonnet", messages=[...]) +response = litellm.completion(model="azure/gpt-4o", messages=[...]) +``` + +The callback compresses messages in LiteLLM's `pre_call_hook` before they're sent to the provider. Works with all 100+ LiteLLM-supported providers. + +### With LiteLLM Proxy + +If you run LiteLLM as a proxy server, use the ASGI middleware instead: + +```python +# In your LiteLLM proxy startup +from litellm.proxy.proxy_server import app +from headroom.integrations.asgi import CompressionMiddleware + +app.add_middleware(CompressionMiddleware) +``` + +Or use the callback in your LiteLLM config: + +```yaml +# litellm_config.yaml +litellm_settings: + callbacks: ["headroom.integrations.litellm_callback.HeadroomCallback"] +``` + +--- + +## ASGI Middleware + +Drop-in middleware for any ASGI application (FastAPI, Starlette, LiteLLM proxy, custom proxies). + +```python +from headroom.integrations.asgi import CompressionMiddleware + +# FastAPI +app = FastAPI() +app.add_middleware(CompressionMiddleware) + +# Starlette +app = Starlette(routes=[...]) +app.add_middleware(CompressionMiddleware) + +# LiteLLM proxy +from litellm.proxy.proxy_server import app +app.add_middleware(CompressionMiddleware) +``` + +The middleware intercepts POST requests to `/v1/messages`, `/v1/chat/completions`, `/v1/responses`, and `/chat/completions`. All other requests pass through untouched. + +Response headers include: +- `x-headroom-compressed: true` — compression was applied +- `x-headroom-tokens-saved: 1234` — tokens removed + +--- + +## Proxy + +The Headroom proxy is a standalone HTTP server. Best for non-Python apps or tools that only support base URL configuration (Claude Code, Cursor). + +```bash +pip install "headroom-ai[all]" +headroom proxy --port 8787 +``` + +```bash +# Claude Code +ANTHROPIC_BASE_URL=http://localhost:8787 claude + +# Cursor / Any OpenAI client +OPENAI_BASE_URL=http://localhost:8787/v1 cursor +``` + +### With Cloud Providers + +```bash +# AWS Bedrock +headroom proxy --backend bedrock --region us-east-1 + +# Google Vertex AI +headroom proxy --backend vertex_ai --region us-central1 + +# Azure OpenAI +headroom proxy --backend azure + +# OpenRouter (400+ models) +OPENROUTER_API_KEY=sk-or-... headroom proxy --backend openrouter +``` + +See [Proxy Documentation](proxy.md) for all options. + +--- + +## Agno + +Full integration with the Agno agent framework. + +```python +from agno.agent import Agent +from agno.models.anthropic import Claude +from headroom.integrations.agno import HeadroomAgnoModel + +model = HeadroomAgnoModel(Claude(id="claude-sonnet-4-20250514")) +agent = Agent(model=model, tools=[your_tools]) +response = agent.run("Investigate the issue") + +print(f"Tokens saved: {model.total_tokens_saved}") +``` + +See [Agno Guide](agno.md) for hooks, multi-provider, and streaming. + +--- + +## LangChain + +> **Experimental.** Core compression works. Streaming callbacks and async chains are still being tested. + +```python +from langchain_openai import ChatOpenAI +from headroom.integrations import HeadroomChatModel + +llm = HeadroomChatModel(ChatOpenAI(model="gpt-4o")) +response = llm.invoke("Hello!") +``` + +See [LangChain Guide](langchain.md) for details and known limitations. + +--- + +## Compression Hooks (Advanced) + +Customize compression behavior without modifying Headroom's code: + +```python +from headroom import compress, CompressionHooks, CompressContext + +class MyHooks(CompressionHooks): + def pre_compress(self, messages, ctx): + # Modify messages before compression (dedup, filter, inject) + return messages + + def compute_biases(self, messages, ctx): + # Per-message compression aggressiveness + # >1.0 = keep more, <1.0 = compress more + return {5: 1.5, 6: 0.5} # Keep message 5, compress message 6 + + def post_compress(self, event): + # Observe results (logging, analytics, learning) + print(f"Saved {event.tokens_saved} tokens") + +result = compress(messages, model="gpt-4o", hooks=MyHooks()) +``` + +See [Architecture](ARCHITECTURE.md) for how hooks integrate with the pipeline. + +--- + +## FAQ + +**Q: Does Headroom change the response format?** +No. Your LLM returns the same response format. Headroom only modifies the input messages. + +**Q: What if compression removes something the LLM needs?** +Headroom stores originals in CCR (Compress-Cache-Retrieve). The LLM can call `headroom_retrieve` to get full uncompressed content. Compression summaries tell the LLM what's available. + +**Q: Does it work with streaming?** +Yes. Compression happens before the request is sent. Streaming responses are unaffected. + +**Q: How much latency does it add?** +1-5ms for compression. The token savings typically save more time on the LLM side than compression adds. diff --git a/headroom/compress.py b/headroom/compress.py index 187c9ac31..bec533f73 100644 --- a/headroom/compress.py +++ b/headroom/compress.py @@ -183,17 +183,13 @@ def _get_pipeline() -> Any: if _pipeline is not None: return _pipeline - from headroom.transforms import ContentRouter, SmartCrusher, TransformPipeline + from headroom.transforms import TransformPipeline - _pipeline = TransformPipeline( - transforms=[ - ContentRouter(), - SmartCrusher(), - ], - # No provider needed — pipeline uses tokenizer registry which - # auto-detects the right tokenizer per model: - # OpenAI → tiktoken (exact), Anthropic → calibrated estimation, - # Open models → HuggingFace (if installed) - ) + # Default pipeline: CacheAligner → ContentRouter → IntelligentContext + # CacheAligner: stabilizes prefix for provider KV cache hits + # ContentRouter: routes to the right compressor per content type + # (SmartCrusher for JSON, CodeCompressor for code, LLMLingua for text) + # IntelligentContext: enforces token limits with score-based dropping + _pipeline = TransformPipeline() logger.debug("Headroom compression pipeline initialized") return _pipeline