mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
Bumps the pip-minor-patch group with 1 update in the / directory: [ruff](https://github.com/astral-sh/ruff). Updates `ruff` from 0.15.22 to 0.16.2 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/astral-sh/ruff/releases">ruff's releases</a>.</em></p> <blockquote> <h2>0.16.2</h2> <h2>Release Notes</h2> <p>Released on 2026-08-06.</p> <h3>Bug fixes</h3> <ul> <li>[<code>flake8-pyi</code>] Avoid false positives on <code>singledispatch</code> functions (<code>PYI041</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27335">#27335</a>)</li> </ul> <h3>Server</h3> <ul> <li>Register formatting capabilities dynamically to exclude TOML files (<a href="https://redirect.github.com/astral-sh/ruff/pull/27332">#27332</a>)</li> </ul> <h3>Contributors</h3> <ul> <li><a href="https://github.com/MeGaGiGaGon"><code>@MeGaGiGaGon</code></a></li> <li><a href="https://github.com/charliermarsh"><code>@charliermarsh</code></a></li> <li><a href="https://github.com/epage"><code>@epage</code></a></li> <li><a href="https://github.com/sharkdp"><code>@sharkdp</code></a></li> <li><a href="https://github.com/ntBre"><code>@ntBre</code></a></li> </ul> <h2>Install ruff 0.16.2</h2> <h3>Install prebuilt binaries via shell script</h3> <pre lang="sh"><code>curl --proto '=https' --tlsv1.2 -LsSf https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-installer.sh | sh </code></pre> <h3>Install prebuilt binaries via powershell script</h3> <pre lang="sh"><code>powershell -ExecutionPolicy Bypass -c "irm https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-installer.ps1 | iex" </code></pre> <h2>Download ruff 0.16.2</h2> <table> <thead> <tr> <th>File</th> <th>Platform</th> <th>Checksum</th> </tr> </thead> <tbody> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-apple-darwin.tar.gz">ruff-aarch64-apple-darwin.tar.gz</a></td> <td>Apple Silicon macOS</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-apple-darwin.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-x86_64-apple-darwin.tar.gz">ruff-x86_64-apple-darwin.tar.gz</a></td> <td>Intel macOS</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-x86_64-apple-darwin.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-pc-windows-msvc.zip">ruff-aarch64-pc-windows-msvc.zip</a></td> <td>ARM64 Windows</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-pc-windows-msvc.zip.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-i686-pc-windows-msvc.zip">ruff-i686-pc-windows-msvc.zip</a></td> <td>x86 Windows</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-i686-pc-windows-msvc.zip.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-x86_64-pc-windows-msvc.zip">ruff-x86_64-pc-windows-msvc.zip</a></td> <td>x64 Windows</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-x86_64-pc-windows-msvc.zip.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-unknown-linux-gnu.tar.gz">ruff-aarch64-unknown-linux-gnu.tar.gz</a></td> <td>ARM64 Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-i686-unknown-linux-gnu.tar.gz">ruff-i686-unknown-linux-gnu.tar.gz</a></td> <td>x86 Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-i686-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-powerpc64-unknown-linux-gnu.tar.gz">ruff-powerpc64-unknown-linux-gnu.tar.gz</a></td> <td>PPC64 Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-powerpc64-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-powerpc64le-unknown-linux-gnu.tar.gz">ruff-powerpc64le-unknown-linux-gnu.tar.gz</a></td> <td>PPC64LE Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-powerpc64le-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-riscv64gc-unknown-linux-gnu.tar.gz">ruff-riscv64gc-unknown-linux-gnu.tar.gz</a></td> <td>RISCV Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-riscv64gc-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-s390x-unknown-linux-gnu.tar.gz">ruff-s390x-unknown-linux-gnu.tar.gz</a></td> <td>S390x Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-s390x-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> </tbody> </table> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md">ruff's changelog</a>.</em></p> <blockquote> <h2>0.16.2</h2> <p>Released on 2026-08-06.</p> <h3>Bug fixes</h3> <ul> <li>[<code>flake8-pyi</code>] Avoid false positives on <code>singledispatch</code> functions (<code>PYI041</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27335">#27335</a>)</li> </ul> <h3>Server</h3> <ul> <li>Register formatting capabilities dynamically to exclude TOML files (<a href="https://redirect.github.com/astral-sh/ruff/pull/27332">#27332</a>)</li> </ul> <h3>Contributors</h3> <ul> <li><a href="https://github.com/MeGaGiGaGon"><code>@MeGaGiGaGon</code></a></li> <li><a href="https://github.com/charliermarsh"><code>@charliermarsh</code></a></li> <li><a href="https://github.com/epage"><code>@epage</code></a></li> <li><a href="https://github.com/sharkdp"><code>@sharkdp</code></a></li> <li><a href="https://github.com/ntBre"><code>@ntBre</code></a></li> </ul> <h2>0.16.1</h2> <p>Released on 2026-07-30.</p> <h3>Preview features</h3> <ul> <li>Add an option to opt out of human-readable names (<a href="https://redirect.github.com/astral-sh/ruff/pull/27160">#27160</a>)</li> <li>[<code>flake8-pytest-style</code>] Make fixes safe by default and unsafe only when comments are present (<code>PT018</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27201">#27201</a>)</li> <li>[<code>pyupgrade</code>] Skip fix when a defaulted <code>TypeVar</code> precedes a non-defaulted one (<code>UP040</code>, <code>UP046</code>, <code>UP047</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27133">#27133</a>)</li> <li>[<code>ruff</code>] Fix false positive with unpacked arguments (<code>RUF065</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/26959">#26959</a>)</li> </ul> <h3>Bug fixes</h3> <ul> <li>Bump <code>gen-lsp-types</code> to gracefully handle unknown enumeration values in LSP messages (<a href="https://redirect.github.com/astral-sh/ruff/pull/27230">#27230</a>)</li> <li>[<code>flake8-bugbear</code>] Mark <code>range</code> as immutable (<code>B008</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27247">#27247</a>)</li> <li>[<code>flake8-comprehensions</code>] NFKC-normalize keyword names in <code>C408</code> fix (<a href="https://redirect.github.com/astral-sh/ruff/pull/26813">#26813</a>)</li> <li>[<code>flake8-return</code>] Fix false positive when variable is read in <code>finally</code> clause (<code>RET504</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/25441">#25441</a>)</li> <li>[<code>pydocstyle</code>] Skip section detection inside RST directive bodies (<code>D214</code>, <code>D405</code>, <code>D413</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/23635">#23635</a>)</li> <li>[<code>refurb</code>] Parenthesize <code>yield</code> arguments in the <code>FURB192</code> fix (<a href="https://redirect.github.com/astral-sh/ruff/pull/27192">#27192</a>)</li> </ul> <h3>Rule changes</h3> <ul> <li>[<code>flake8-pytest-style</code>] Mark <code>PT022</code> fixes as unsafe (<a href="https://redirect.github.com/astral-sh/ruff/pull/26440">#26440</a>)</li> <li>[<code>refurb</code>] Mark fixes that remove unknown separators as unsafe (<code>FURB105</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27200">#27200</a>)</li> </ul> <h3>Server</h3> <ul> <li>Fix indexing of excluded nested Ruff workspaces (<a href="https://redirect.github.com/astral-sh/ruff/pull/27303">#27303</a>)</li> <li>Lint TOML files in the LSP (<a href="https://redirect.github.com/astral-sh/ruff/pull/26862">#26862</a>)</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="5b48a04097"><code>5b48a04</code></a> Bump 0.16.2 (<a href="https://redirect.github.com/astral-sh/ruff/issues/27555">#27555</a>)</li> <li><a href="1b9e5fc483"><code>1b9e5fc</code></a> Update Swatinem/rust-cache action to v2.9.2 (<a href="https://redirect.github.com/astral-sh/ruff/issues/27568">#27568</a>)</li> <li><a href="c4e86fc039"><code>c4e86fc</code></a> [ty] Add helper extension methods for half-range and equality constraints (<a href="https://redirect.github.com/astral-sh/ruff/issues/2">#2</a>...</li> <li><a href="17a00de2e2"><code>17a00de</code></a> [ty] Reuse primer commands in memory reports (<a href="https://redirect.github.com/astral-sh/ruff/issues/27553">#27553</a>)</li> <li><a href="6ea296b969"><code>6ea296b</code></a> [ty] Normalize type labels in structured docstrings (<a href="https://redirect.github.com/astral-sh/ruff/issues/26923">#26923</a>)</li> <li><a href="2fc445f005"><code>2fc445f</code></a> [ty] Diagnose invalid <strong>getattr</strong> calls (<a href="https://redirect.github.com/astral-sh/ruff/issues/27502">#27502</a>)</li> <li><a href="22c7823c4e"><code>22c7823</code></a> [ty] Enable (but downrank) auto-import completion suggestions from stub-only ...</li> <li><a href="05160d507f"><code>05160d5</code></a> [ty] Diagnose invalid descriptor <code>__get__</code> calls (<a href="https://redirect.github.com/astral-sh/ruff/issues/27400">#27400</a>)</li> <li><a href="baea3d0dce"><code>baea3d0</code></a> [ty] Expose strict analysis options in the playground (<a href="https://redirect.github.com/astral-sh/ruff/issues/27543">#27543</a>)</li> <li><a href="c88946ebeb"><code>c88946e</code></a> [ty] Bump ecosystem-analyzer for strict project settings (<a href="https://redirect.github.com/astral-sh/ruff/issues/27542">#27542</a>)</li> <li>Additional commits viewable in <a href="https://github.com/astral-sh/ruff/compare/0.15.22...0.16.2">compare view</a></li> </ul> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: JerrettDavis <mxjerrett@gmail.com>
372 lines
11 KiB
Markdown
372 lines
11 KiB
Markdown
# Transform Reference
|
|
|
|
Headroom provides several transforms that work together to optimize LLM context.
|
|
|
|
## SmartCrusher
|
|
|
|
Statistical compression for JSON tool outputs.
|
|
|
|
### How It Works
|
|
|
|
SmartCrusher analyzes JSON arrays and selectively keeps important items:
|
|
|
|
1. **First/Last items** - Context for pagination and recency
|
|
2. **Error items** - 100% preservation of error states
|
|
3. **Anomalies** - Statistical outliers (> 2 std dev from mean)
|
|
4. **Relevant items** - Matches to user's query via BM25/embeddings
|
|
5. **Change points** - Significant transitions in data
|
|
|
|
### Configuration
|
|
|
|
```python
|
|
from headroom import SmartCrusherConfig
|
|
|
|
config = SmartCrusherConfig(
|
|
min_tokens_to_crush=200, # Only compress if > 200 tokens
|
|
max_items_after_crush=50, # Keep at most 50 items
|
|
keep_first=3, # Always keep first 3 items
|
|
keep_last=2, # Always keep last 2 items
|
|
relevance_threshold=0.3, # Keep items with relevance > 0.3
|
|
anomaly_std_threshold=2.0, # Keep items > 2 std dev from mean
|
|
preserve_errors=True, # Always keep error items
|
|
)
|
|
```
|
|
|
|
### Example
|
|
|
|
```python
|
|
from headroom import SmartCrusher
|
|
|
|
crusher = SmartCrusher(config)
|
|
|
|
# Before: 1000 search results (45,000 tokens)
|
|
tool_output = {"results": [...1000 items...]}
|
|
|
|
# After: ~50 important items (4,500 tokens) - 90% reduction
|
|
compressed = crusher.crush(tool_output, query="user's question")
|
|
```
|
|
|
|
### What Gets Preserved
|
|
|
|
| Category | Preserved | Why |
|
|
|----------|-----------|-----|
|
|
| Errors | 100% | Critical for debugging |
|
|
| First N | 100% | Context/pagination |
|
|
| Last N | 100% | Recency |
|
|
| Anomalies | All | Unusual values matter |
|
|
| Relevant | Top K | Match user's query |
|
|
| Others | Sampled | Statistical representation |
|
|
|
|
---
|
|
|
|
## CacheAligner
|
|
|
|
Prefix stabilization for improved cache hit rates.
|
|
|
|
### The Problem
|
|
|
|
LLM providers cache request prefixes. But dynamic content breaks caching:
|
|
|
|
```
|
|
"You are helpful. Today is January 7, 2025." # Changes daily = no cache
|
|
```
|
|
|
|
### The Solution
|
|
|
|
CacheAligner extracts dynamic content to stabilize the prefix:
|
|
|
|
```python
|
|
from headroom import CacheAligner
|
|
|
|
aligner = CacheAligner()
|
|
result = aligner.align(messages)
|
|
|
|
# Static prefix (cacheable):
|
|
# "You are helpful."
|
|
|
|
# Dynamic content moved to end:
|
|
# [Current date context]
|
|
```
|
|
|
|
### Configuration
|
|
|
|
```python
|
|
from headroom import CacheAlignerConfig
|
|
|
|
config = CacheAlignerConfig(
|
|
extract_dates=True, # Move dates to dynamic section
|
|
normalize_whitespace=True, # Consistent spacing
|
|
stable_prefix_min_tokens=100, # Min prefix size for alignment
|
|
)
|
|
```
|
|
|
|
### Cache Hit Improvement
|
|
|
|
| Scenario | Before | After |
|
|
|----------|--------|-------|
|
|
| Daily date in prompt | 0% hits | ~95% hits |
|
|
| Dynamic user context | ~10% hits | ~80% hits |
|
|
| Consistent prompts | ~90% hits | ~95% hits |
|
|
|
|
---
|
|
|
|
## Context management
|
|
|
|
Context management is handled automatically inside the pipeline
|
|
(live-zone-only compression). Headroom **never** drops messages from the
|
|
conversation history and does not do position-based or score-based context
|
|
management. It compresses only the newest content blocks (the latest user
|
|
message and the latest tool result / tool output), type-aware and reversible
|
|
via CCR. The cache hot zone — system prompt, tools, and older turns — is never
|
|
mutated, which preserves provider prompt caching.
|
|
|
|
> The earlier position-based `RollingWindow` and score-based
|
|
> `IntelligentContextManager` transforms have been removed and are no longer
|
|
> part of Headroom.
|
|
|
|
---
|
|
|
|
## LLMLinguaCompressor — RETIRED
|
|
|
|
The earlier LLMLingua-2 integration (`LLMLinguaCompressor`,
|
|
`LLMLinguaConfig`, `is_llmlingua_model_loaded`, `unload_llmlingua_model`,
|
|
the `headroom-ai[llmlingua]` extra, and the `--llmlingua` proxy flag)
|
|
was retired in 0.9.x and replaced by **Kompress** (ModernBERT).
|
|
`pip install 'headroom-ai[llmlingua]'` no longer resolves; use the
|
|
`[ml]` extra instead. The Kompress transform shipped with the proxy
|
|
runs as Transform 4 in the live-zone pipeline (see
|
|
[ARCHITECTURE.md](ARCHITECTURE.md)).
|
|
|
|
---
|
|
|
|
## CodeAwareCompressor (Optional)
|
|
|
|
AST-based compression for source code using tree-sitter.
|
|
|
|
### When to Use
|
|
|
|
| Transform | Best For | Speed | Compression |
|
|
|-----------|----------|-------|-------------|
|
|
| SmartCrusher | JSON arrays | ~1ms | 70-90% |
|
|
| **CodeAwareCompressor** | Source code | ~10-50ms | 40-70% |
|
|
| Kompress (ML) | Any text | 50-200ms | 80-95% |
|
|
|
|
### Key Benefits
|
|
|
|
- **Syntax validity guaranteed** — Output always parses correctly
|
|
- **Preserves critical structure** — Imports, signatures, types, error handlers
|
|
- **Multi-language support** — Python, JavaScript, TypeScript, Go, Rust, Java, C, C++
|
|
- **Lightweight** — ~50MB vs ~1GB for the ML compressor
|
|
|
|
### Installation
|
|
|
|
```bash
|
|
pip install "headroom-ai[code]" # Adds tree-sitter-language-pack
|
|
```
|
|
|
|
### Configuration
|
|
|
|
```python
|
|
from headroom.transforms import CodeAwareCompressor, CodeCompressorConfig, DocstringMode
|
|
|
|
config = CodeCompressorConfig(
|
|
preserve_imports=True, # Always keep imports
|
|
preserve_signatures=True, # Always keep function signatures
|
|
preserve_type_annotations=True, # Keep type hints
|
|
preserve_error_handlers=True, # Keep try/except blocks
|
|
preserve_decorators=True, # Keep decorators
|
|
docstring_mode=DocstringMode.FIRST_LINE, # FULL, FIRST_LINE, REMOVE
|
|
target_compression_rate=0.2, # Keep 20% of tokens
|
|
max_body_lines=5, # Lines to keep per function body
|
|
min_tokens_for_compression=100, # Skip small content
|
|
language_hint=None, # Auto-detect if None
|
|
)
|
|
|
|
compressor = CodeAwareCompressor(config)
|
|
```
|
|
|
|
### Example
|
|
|
|
```python
|
|
from headroom.transforms import CodeAwareCompressor
|
|
|
|
compressor = CodeAwareCompressor()
|
|
|
|
code = '''
|
|
import os
|
|
from typing import List
|
|
|
|
def process_items(items: List[str]) -> List[str]:
|
|
"""Process a list of items."""
|
|
results = []
|
|
for item in items:
|
|
if not item:
|
|
continue
|
|
processed = item.strip().lower()
|
|
results.append(processed)
|
|
return results
|
|
'''
|
|
|
|
result = compressor.compress(code, language="python")
|
|
print(result.compressed)
|
|
# import os
|
|
# from typing import List
|
|
#
|
|
# def process_items(items: List[str]) -> List[str]:
|
|
# """Process a list of items."""
|
|
# results = []
|
|
# for item in items:
|
|
# # ... (5 lines compressed)
|
|
# pass
|
|
|
|
print(f"Compression: {result.compression_ratio:.0%}") # ~55%
|
|
print(f"Syntax valid: {result.syntax_valid}") # True
|
|
```
|
|
|
|
### Supported Languages
|
|
|
|
| Tier | Languages | Support Level |
|
|
|------|-----------|---------------|
|
|
| 1 | Python, JavaScript, TypeScript | Full AST analysis |
|
|
| 2 | Go, Rust, Java, C, C++ | Function body compression |
|
|
|
|
### Memory Management
|
|
|
|
```python
|
|
from headroom.transforms import is_tree_sitter_available, unload_tree_sitter
|
|
|
|
# Check if tree-sitter is installed
|
|
print(is_tree_sitter_available()) # True/False
|
|
|
|
# Free memory when done (parsers are lazy-loaded)
|
|
unload_tree_sitter()
|
|
```
|
|
|
|
---
|
|
|
|
## ContentRouter
|
|
|
|
Intelligent compression orchestrator that routes content to the optimal compressor.
|
|
|
|
### How It Works
|
|
|
|
ContentRouter analyzes content and selects the best compression strategy:
|
|
|
|
1. **Detect content type** — JSON, code, logs, search results, plain text
|
|
2. **Consider source hints** — File paths, tool names for high-confidence routing
|
|
3. **Route to compressor** — SmartCrusher, CodeAwareCompressor, SearchCompressor, etc.
|
|
4. **Log decisions** — Transparent routing for debugging
|
|
|
|
### Configuration
|
|
|
|
```python
|
|
from headroom.transforms import ContentRouter, ContentRouterConfig, CompressionStrategy
|
|
|
|
config = ContentRouterConfig(
|
|
min_section_tokens=100, # Minimum tokens to compress
|
|
enable_code_aware=True, # Use CodeAwareCompressor for code
|
|
enable_search_compression=True, # Use SearchCompressor for grep output
|
|
enable_log_compression=True, # Use LogCompressor for logs
|
|
default_strategy=CompressionStrategy.TEXT, # Fallback strategy
|
|
)
|
|
|
|
router = ContentRouter(config)
|
|
```
|
|
|
|
### Example
|
|
|
|
```python
|
|
from headroom.transforms import ContentRouter
|
|
|
|
router = ContentRouter()
|
|
|
|
# Router auto-detects content type and routes to optimal compressor
|
|
result = router.compress(content)
|
|
|
|
print(result.strategy_used) # CompressionStrategy.CODE_AWARE, SMART_CRUSHER, etc.
|
|
print(result.routing_log) # List of routing decisions
|
|
```
|
|
|
|
### Compression Strategies
|
|
|
|
| Strategy | Used For | Compressor |
|
|
|----------|----------|------------|
|
|
| CODE_AWARE | Source code | CodeAwareCompressor |
|
|
| SMART_CRUSHER | JSON arrays | SmartCrusher |
|
|
| SEARCH | Grep/find output | SearchCompressor |
|
|
| LOG | Log files | LogCompressor |
|
|
| TEXT | Plain text | TextCompressor |
|
|
| PASSTHROUGH | Small content | None |
|
|
|
|
(The earlier `LLMLINGUA` strategy was retired with the LLMLingua integration; ML compression is now provided by Kompress.)
|
|
|
|
### Content Detection
|
|
|
|
The router automatically detects content types by analyzing the content itself:
|
|
|
|
- **Source code**: Detected by syntax patterns, indentation, keywords
|
|
- **JSON arrays**: Detected by JSON structure with array elements
|
|
- **Search results**: Detected by `file:line:` patterns
|
|
- **Log output**: Detected by timestamp and log level patterns
|
|
- **Plain text**: Fallback for prose content
|
|
|
|
No manual hints required - the router inspects content directly.
|
|
|
|
### TOIN Integration
|
|
|
|
ContentRouter records all compressions to TOIN (Tool Output Intelligence Network) for cross-user learning:
|
|
|
|
- **All strategies tracked**: Code, search, logs, text, and ML compressions are recorded
|
|
- **Retrieval feedback**: When users retrieve original content via CCR, TOIN learns which compressions need expansion
|
|
- **Pattern learning**: TOIN builds signatures for each content type to improve future compressions
|
|
|
|
This enables the feedback loop where compression decisions improve based on actual user behavior across all content types, not just JSON arrays.
|
|
|
|
---
|
|
|
|
## TransformPipeline
|
|
|
|
Combine transforms for optimal results.
|
|
|
|
```python
|
|
from headroom import TransformPipeline, SmartCrusher, CacheAligner
|
|
|
|
pipeline = TransformPipeline(
|
|
[
|
|
SmartCrusher(), # First: compress tool outputs
|
|
CacheAligner(), # Then: stabilize prefix
|
|
]
|
|
)
|
|
|
|
result = pipeline.transform(messages)
|
|
print(f"Saved {result.tokens_saved} tokens")
|
|
```
|
|
|
|
### With ML compression (Optional, Kompress)
|
|
|
|
The earlier hand-assembled `TransformPipeline([..., LLMLinguaCompressor(), ...])` recipe is no longer supported. ML compression now ships as part of the live-zone pipeline when the `[ml]` extra is installed; see [ARCHITECTURE.md](ARCHITECTURE.md) for the current placement.
|
|
|
|
### Recommended Order
|
|
|
|
| Order | Transform | Purpose |
|
|
|-------|-----------|---------|
|
|
| 1 | CacheAligner | Stabilize prefix for caching |
|
|
| 2 | SmartCrusher | Compress JSON tool outputs |
|
|
| 3 | Kompress (ML) | ML compression on remaining text (optional, `[ml]` extra) |
|
|
|
|
**Why this order?**
|
|
- CacheAligner first to maximize prefix stability
|
|
- SmartCrusher handles JSON arrays efficiently
|
|
- Kompress compresses remaining long text
|
|
|
|
---
|
|
|
|
## Safety Guarantees
|
|
|
|
All transforms follow strict safety rules:
|
|
|
|
1. **Never remove human content** - User/assistant text is sacred
|
|
2. **Never break tool ordering** - Calls and results stay paired
|
|
3. **Parse failures are no-ops** - Malformed content passes through
|
|
4. **Preserves recency** - Last N turns always kept
|
|
5. **100% error preservation** - Error items never dropped
|