headroom/docs
Tejas Chopra 2fd9552102 Add image token compression with trained ML router
Introduces automatic image compression for LLM requests, reducing token
usage by 40-90% while maintaining answer accuracy.

Key features:
- Trained MiniLM classifier (93.7% accuracy) hosted on HuggingFace
- SigLIP-based image analysis for content-aware routing
- Provider-specific compression:
  - OpenAI: detail="low" parameter
  - Anthropic: PIL resize to 512px
  - Google: PIL resize to 768px (tile-optimized)
- Four compression techniques: full_low, preserve, crop, transcode
- Integration in both Headroom proxy and SDK (ContentRouter)

New files:
- headroom/image/ module with ImageCompressor API
- docs/image-compression.md user documentation
- tests/test_image_compressor.py (51 tests)

Model: chopratejas/technique-router on HuggingFace (~128MB)
2026-01-25 22:40:43 -08:00
..
plans Add dynamic anchor selection with content deduplication 2026-01-21 00:19:05 -08:00
agno.md Add feature coverage and known limitations to Agno docs 2026-01-16 16:25:31 -08:00
api.md Add IntelligentContextManager for semantic-aware context management 2026-01-18 22:22:48 -08:00
ARCHITECTURE.md Add image token compression with trained ML router 2026-01-25 22:40:43 -08:00
ccr.md v0.2.2: Add CCR Response Handler, Context Tracker, and restructure docs 2026-01-14 13:03:41 -08:00
compression.md Add universal compression module with ML-based content detection 2026-01-15 15:26:14 -08:00
configuration.md Add IntelligentContextManager for semantic-aware context management 2026-01-18 22:22:48 -08:00
errors.md v0.2.2: Add CCR Response Handler, Context Tracker, and restructure docs 2026-01-14 13:03:41 -08:00
getting-started.md Prepare for OSS release v0.2.0 2026-01-07 11:36:44 -08:00
image-compression.md Add image token compression with trained ML router 2026-01-25 22:40:43 -08:00
langchain.md Add seamless LangChain integration 2026-01-14 16:03:34 -08:00
llmlingua.md v0.2.2: Add CCR Response Handler, Context Tracker, and restructure docs 2026-01-14 13:03:41 -08:00
macos-deployment.md docs(deployment): add macOS LaunchAgent deployment guide and templates 2026-01-19 13:34:52 -08:00
memory.md Update memory documentation for HierarchicalMemory system 2026-01-22 23:37:14 -08:00
metrics.md v0.2.2: Add CCR Response Handler, Context Tracker, and restructure docs 2026-01-14 13:03:41 -08:00
proxy.md docs update 2026-01-19 23:35:16 +01:00
quickstart.md Publish headroom-ai v0.2.0 to PyPI with DevEx fixes 2026-01-10 14:51:08 -08:00
README.md Add image token compression with trained ML router 2026-01-25 22:40:43 -08:00
sdk.md v0.2.2: Add CCR Response Handler, Context Tracker, and restructure docs 2026-01-14 13:03:41 -08:00
text-compression.md v0.2.2: Add CCR Response Handler, Context Tracker, and restructure docs 2026-01-14 13:03:41 -08:00
transforms.md Remove hardcoded source hint system from ContentRouter 2026-01-23 00:28:22 -08:00
troubleshooting.md Publish headroom-ai v0.2.0 to PyPI with DevEx fixes 2026-01-10 14:51:08 -08:00

Headroom Documentation

Welcome to the Headroom documentation.

Getting Started

Guide Description
Quickstart 5-minute setup
SDK Guide Python SDK usage
Proxy Guide Proxy server deployment

Framework Integrations

Framework Description
LangChain Chat models, memory, retrievers, agents, streaming
Agno Model wrapper, hooks, multi-provider support
MCP See CCR Guide for tool compression

Core Concepts

Topic Description
Universal Compression ML-based content detection + structure preservation
Image Compression 40-90% token reduction for images via trained ML router
Transforms How compression works
CCR Reversible compression architecture
Configuration All configuration options

Advanced

Topic Description
Text Compression Opt-in utilities for search/logs
LLMLingua ML-based compression
Metrics Monitoring and observability
Errors Error handling

Deployment & Operations

Guide Description
macOS Deployment Run proxy as background service on macOS

Reference

Topic Description
API Reference Complete API docs
Architecture Internal design
Troubleshooting Common issues

Overview

Headroom is the Context Optimization Layer for LLM applications. It reduces your LLM costs by 50-90% through intelligent context compression.

How It Works

  1. Universal Compression — ML-based content detection with structure-preserving compression
  2. SmartCrusher — Compresses JSON tool outputs, keeping errors, anomalies, and relevant items
  3. CacheAligner — Stabilizes message prefixes so provider caching works
  4. RollingWindow — Manages context limits without breaking tool call pairs
  5. CCR — Caches original data so compression is reversible

Safety Guarantees

  • Never removes human content
  • Never breaks tool call ordering
  • Parse failures pass through unchanged
  • LLM can always retrieve original data

Getting Help