mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
Bumps the pip-minor-patch group with 1 update in the / directory: [ruff](https://github.com/astral-sh/ruff). Updates `ruff` from 0.15.22 to 0.16.2 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/astral-sh/ruff/releases">ruff's releases</a>.</em></p> <blockquote> <h2>0.16.2</h2> <h2>Release Notes</h2> <p>Released on 2026-08-06.</p> <h3>Bug fixes</h3> <ul> <li>[<code>flake8-pyi</code>] Avoid false positives on <code>singledispatch</code> functions (<code>PYI041</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27335">#27335</a>)</li> </ul> <h3>Server</h3> <ul> <li>Register formatting capabilities dynamically to exclude TOML files (<a href="https://redirect.github.com/astral-sh/ruff/pull/27332">#27332</a>)</li> </ul> <h3>Contributors</h3> <ul> <li><a href="https://github.com/MeGaGiGaGon"><code>@MeGaGiGaGon</code></a></li> <li><a href="https://github.com/charliermarsh"><code>@charliermarsh</code></a></li> <li><a href="https://github.com/epage"><code>@epage</code></a></li> <li><a href="https://github.com/sharkdp"><code>@sharkdp</code></a></li> <li><a href="https://github.com/ntBre"><code>@ntBre</code></a></li> </ul> <h2>Install ruff 0.16.2</h2> <h3>Install prebuilt binaries via shell script</h3> <pre lang="sh"><code>curl --proto '=https' --tlsv1.2 -LsSf https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-installer.sh | sh </code></pre> <h3>Install prebuilt binaries via powershell script</h3> <pre lang="sh"><code>powershell -ExecutionPolicy Bypass -c "irm https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-installer.ps1 | iex" </code></pre> <h2>Download ruff 0.16.2</h2> <table> <thead> <tr> <th>File</th> <th>Platform</th> <th>Checksum</th> </tr> </thead> <tbody> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-apple-darwin.tar.gz">ruff-aarch64-apple-darwin.tar.gz</a></td> <td>Apple Silicon macOS</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-apple-darwin.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-x86_64-apple-darwin.tar.gz">ruff-x86_64-apple-darwin.tar.gz</a></td> <td>Intel macOS</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-x86_64-apple-darwin.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-pc-windows-msvc.zip">ruff-aarch64-pc-windows-msvc.zip</a></td> <td>ARM64 Windows</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-pc-windows-msvc.zip.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-i686-pc-windows-msvc.zip">ruff-i686-pc-windows-msvc.zip</a></td> <td>x86 Windows</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-i686-pc-windows-msvc.zip.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-x86_64-pc-windows-msvc.zip">ruff-x86_64-pc-windows-msvc.zip</a></td> <td>x64 Windows</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-x86_64-pc-windows-msvc.zip.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-unknown-linux-gnu.tar.gz">ruff-aarch64-unknown-linux-gnu.tar.gz</a></td> <td>ARM64 Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-aarch64-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-i686-unknown-linux-gnu.tar.gz">ruff-i686-unknown-linux-gnu.tar.gz</a></td> <td>x86 Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-i686-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-powerpc64-unknown-linux-gnu.tar.gz">ruff-powerpc64-unknown-linux-gnu.tar.gz</a></td> <td>PPC64 Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-powerpc64-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-powerpc64le-unknown-linux-gnu.tar.gz">ruff-powerpc64le-unknown-linux-gnu.tar.gz</a></td> <td>PPC64LE Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-powerpc64le-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-riscv64gc-unknown-linux-gnu.tar.gz">ruff-riscv64gc-unknown-linux-gnu.tar.gz</a></td> <td>RISCV Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-riscv64gc-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> <tr> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-s390x-unknown-linux-gnu.tar.gz">ruff-s390x-unknown-linux-gnu.tar.gz</a></td> <td>S390x Linux</td> <td><a href="https://releases.astral.sh/github/ruff/releases/download/0.16.2/ruff-s390x-unknown-linux-gnu.tar.gz.sha256">checksum</a></td> </tr> </tbody> </table> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/astral-sh/ruff/blob/main/CHANGELOG.md">ruff's changelog</a>.</em></p> <blockquote> <h2>0.16.2</h2> <p>Released on 2026-08-06.</p> <h3>Bug fixes</h3> <ul> <li>[<code>flake8-pyi</code>] Avoid false positives on <code>singledispatch</code> functions (<code>PYI041</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27335">#27335</a>)</li> </ul> <h3>Server</h3> <ul> <li>Register formatting capabilities dynamically to exclude TOML files (<a href="https://redirect.github.com/astral-sh/ruff/pull/27332">#27332</a>)</li> </ul> <h3>Contributors</h3> <ul> <li><a href="https://github.com/MeGaGiGaGon"><code>@MeGaGiGaGon</code></a></li> <li><a href="https://github.com/charliermarsh"><code>@charliermarsh</code></a></li> <li><a href="https://github.com/epage"><code>@epage</code></a></li> <li><a href="https://github.com/sharkdp"><code>@sharkdp</code></a></li> <li><a href="https://github.com/ntBre"><code>@ntBre</code></a></li> </ul> <h2>0.16.1</h2> <p>Released on 2026-07-30.</p> <h3>Preview features</h3> <ul> <li>Add an option to opt out of human-readable names (<a href="https://redirect.github.com/astral-sh/ruff/pull/27160">#27160</a>)</li> <li>[<code>flake8-pytest-style</code>] Make fixes safe by default and unsafe only when comments are present (<code>PT018</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27201">#27201</a>)</li> <li>[<code>pyupgrade</code>] Skip fix when a defaulted <code>TypeVar</code> precedes a non-defaulted one (<code>UP040</code>, <code>UP046</code>, <code>UP047</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27133">#27133</a>)</li> <li>[<code>ruff</code>] Fix false positive with unpacked arguments (<code>RUF065</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/26959">#26959</a>)</li> </ul> <h3>Bug fixes</h3> <ul> <li>Bump <code>gen-lsp-types</code> to gracefully handle unknown enumeration values in LSP messages (<a href="https://redirect.github.com/astral-sh/ruff/pull/27230">#27230</a>)</li> <li>[<code>flake8-bugbear</code>] Mark <code>range</code> as immutable (<code>B008</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27247">#27247</a>)</li> <li>[<code>flake8-comprehensions</code>] NFKC-normalize keyword names in <code>C408</code> fix (<a href="https://redirect.github.com/astral-sh/ruff/pull/26813">#26813</a>)</li> <li>[<code>flake8-return</code>] Fix false positive when variable is read in <code>finally</code> clause (<code>RET504</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/25441">#25441</a>)</li> <li>[<code>pydocstyle</code>] Skip section detection inside RST directive bodies (<code>D214</code>, <code>D405</code>, <code>D413</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/23635">#23635</a>)</li> <li>[<code>refurb</code>] Parenthesize <code>yield</code> arguments in the <code>FURB192</code> fix (<a href="https://redirect.github.com/astral-sh/ruff/pull/27192">#27192</a>)</li> </ul> <h3>Rule changes</h3> <ul> <li>[<code>flake8-pytest-style</code>] Mark <code>PT022</code> fixes as unsafe (<a href="https://redirect.github.com/astral-sh/ruff/pull/26440">#26440</a>)</li> <li>[<code>refurb</code>] Mark fixes that remove unknown separators as unsafe (<code>FURB105</code>) (<a href="https://redirect.github.com/astral-sh/ruff/pull/27200">#27200</a>)</li> </ul> <h3>Server</h3> <ul> <li>Fix indexing of excluded nested Ruff workspaces (<a href="https://redirect.github.com/astral-sh/ruff/pull/27303">#27303</a>)</li> <li>Lint TOML files in the LSP (<a href="https://redirect.github.com/astral-sh/ruff/pull/26862">#26862</a>)</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="5b48a04097"><code>5b48a04</code></a> Bump 0.16.2 (<a href="https://redirect.github.com/astral-sh/ruff/issues/27555">#27555</a>)</li> <li><a href="1b9e5fc483"><code>1b9e5fc</code></a> Update Swatinem/rust-cache action to v2.9.2 (<a href="https://redirect.github.com/astral-sh/ruff/issues/27568">#27568</a>)</li> <li><a href="c4e86fc039"><code>c4e86fc</code></a> [ty] Add helper extension methods for half-range and equality constraints (<a href="https://redirect.github.com/astral-sh/ruff/issues/2">#2</a>...</li> <li><a href="17a00de2e2"><code>17a00de</code></a> [ty] Reuse primer commands in memory reports (<a href="https://redirect.github.com/astral-sh/ruff/issues/27553">#27553</a>)</li> <li><a href="6ea296b969"><code>6ea296b</code></a> [ty] Normalize type labels in structured docstrings (<a href="https://redirect.github.com/astral-sh/ruff/issues/26923">#26923</a>)</li> <li><a href="2fc445f005"><code>2fc445f</code></a> [ty] Diagnose invalid <strong>getattr</strong> calls (<a href="https://redirect.github.com/astral-sh/ruff/issues/27502">#27502</a>)</li> <li><a href="22c7823c4e"><code>22c7823</code></a> [ty] Enable (but downrank) auto-import completion suggestions from stub-only ...</li> <li><a href="05160d507f"><code>05160d5</code></a> [ty] Diagnose invalid descriptor <code>__get__</code> calls (<a href="https://redirect.github.com/astral-sh/ruff/issues/27400">#27400</a>)</li> <li><a href="baea3d0dce"><code>baea3d0</code></a> [ty] Expose strict analysis options in the playground (<a href="https://redirect.github.com/astral-sh/ruff/issues/27543">#27543</a>)</li> <li><a href="c88946ebeb"><code>c88946e</code></a> [ty] Bump ecosystem-analyzer for strict project settings (<a href="https://redirect.github.com/astral-sh/ruff/issues/27542">#27542</a>)</li> <li>Additional commits viewable in <a href="https://github.com/astral-sh/ruff/compare/0.15.22...0.16.2">compare view</a></li> </ul> </details> <br /> --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: JerrettDavis <mxjerrett@gmail.com>
398 lines
12 KiB
Markdown
398 lines
12 KiB
Markdown
# Universal Compression
|
|
|
|
Headroom's Universal Compression module provides intelligent, automatic compression with ML-based content detection and structure preservation.
|
|
|
|
## Overview
|
|
|
|
Universal Compression combines several techniques:
|
|
|
|
1. **ML-based Detection** - Automatically detects content type (JSON, code, logs, text) using Magika
|
|
2. **Structure Preservation** - Keeps keys, signatures, and templates intact via structure masks
|
|
3. **Intelligent Compression** - Compresses content while preserving meaning with the optional ML compressor (Kompress)
|
|
4. **Reversible via CCR** - Stores originals for retrieval when LLM needs full context
|
|
|
|
## Quick Start
|
|
|
|
### One-Liner
|
|
|
|
```python
|
|
from headroom.compression import compress
|
|
|
|
result = compress(content)
|
|
print(result.compressed)
|
|
print(f"Saved {result.savings_percentage:.0f}% tokens")
|
|
```
|
|
|
|
### With Configuration
|
|
|
|
```python
|
|
from headroom.compression import UniversalCompressor, UniversalCompressorConfig
|
|
|
|
config = UniversalCompressorConfig(
|
|
compression_ratio_target=0.5, # Keep 50% of content
|
|
use_entropy_preservation=True, # Preserve UUIDs, hashes
|
|
)
|
|
|
|
compressor = UniversalCompressor(config=config)
|
|
result = compressor.compress(content)
|
|
```
|
|
|
|
---
|
|
|
|
## How It Works
|
|
|
|
### Detection Flow
|
|
|
|
```
|
|
┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐
|
|
│ Content │───>│ Detect │───>│ Extract │───>│ Compress │
|
|
│ Input │ │ Type │ │ Structure │ │ Content │
|
|
└─────────────┘ └─────────────┘ └─────────────┘ └─────────────┘
|
|
│ │ │
|
|
▼ ▼ ▼
|
|
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
|
|
│ Magika │ │ Handler │ │ Kompress │
|
|
│ (ML) │ │ (JSON, │ │ (ML, opt- │
|
|
│ │ │ Code...) │ │ in [ml]) │
|
|
└─────────────┘ └─────────────┘ └─────────────┘
|
|
```
|
|
|
|
### Structure Masks
|
|
|
|
Structure masks identify what to preserve:
|
|
|
|
| Content Type | What's Preserved | What's Compressed |
|
|
|--------------|------------------|-------------------|
|
|
| **JSON** | Keys, brackets, booleans, nulls, short values, UUIDs | Long string values, whitespace |
|
|
| **Code** | Imports, function signatures, class definitions, types | Function bodies, comments |
|
|
| **Logs** | Timestamps, log levels, error messages | Repeated patterns, verbose details |
|
|
| **Text** | High-entropy tokens (IDs, hashes) | Low-information content |
|
|
|
|
---
|
|
|
|
## Configuration
|
|
|
|
### UniversalCompressorConfig
|
|
|
|
```python
|
|
from headroom.compression import UniversalCompressorConfig
|
|
|
|
config = UniversalCompressorConfig(
|
|
# Detection
|
|
use_magika=True, # Use ML-based detection (requires magika)
|
|
# Compression
|
|
# (Note: the legacy `use_llmlingua` flag was retired with the
|
|
# LLMLingua-2 integration. The optional ML compressor is now Kompress,
|
|
# installed via `headroom-ai[ml]` and configured separately.)
|
|
compression_ratio_target=0.3, # Keep 30% of content (70% reduction)
|
|
min_content_length=100, # Skip content shorter than this
|
|
# Structure preservation
|
|
use_entropy_preservation=True, # Preserve high-entropy tokens
|
|
entropy_threshold=0.85, # Entropy threshold for preservation
|
|
# CCR
|
|
ccr_enabled=True, # Store originals for retrieval
|
|
)
|
|
```
|
|
|
|
### Configuration Options
|
|
|
|
| Option | Default | Description |
|
|
|--------|---------|-------------|
|
|
| `use_magika` | `True` | Use ML-based content detection |
|
|
| `use_llmlingua` | `True` | Use LLMLingua for compression |
|
|
| `compression_ratio_target` | `0.3` | Target ratio (0.3 = keep 30%) |
|
|
| `min_content_length` | `100` | Minimum chars to compress |
|
|
| `use_entropy_preservation` | `True` | Preserve high-entropy tokens |
|
|
| `entropy_threshold` | `0.85` | Entropy threshold (0.0-1.0) |
|
|
| `ccr_enabled` | `True` | Enable CCR storage |
|
|
|
|
---
|
|
|
|
## Content Handlers
|
|
|
|
### JSON Handler
|
|
|
|
Preserves JSON structure while compressing values:
|
|
|
|
```python
|
|
from headroom.compression.handlers.json_handler import JSONStructureHandler
|
|
|
|
handler = JSONStructureHandler(
|
|
preserve_short_values=True, # Keep values < 20 chars
|
|
short_value_threshold=20, # Threshold for "short"
|
|
preserve_high_entropy=True, # Keep UUIDs, hashes
|
|
entropy_threshold=0.85, # Entropy threshold
|
|
max_array_items_full=3, # Keep first N array items full
|
|
max_number_digits=10, # Preserve numbers up to N digits
|
|
)
|
|
```
|
|
|
|
**What's Preserved:**
|
|
- All keys (navigational - LLM sees schema)
|
|
- Structural syntax (`{`, `}`, `[`, `]`, `:`, `,`)
|
|
- Booleans and nulls (semantically important)
|
|
- High-entropy strings (UUIDs, hashes - identifiers)
|
|
- Short numbers (often IDs)
|
|
|
|
**Example:**
|
|
|
|
```python
|
|
# Before
|
|
{"id": "usr_abc123", "name": "Alice Johnson", "bio": "A long description that goes on and on..."}
|
|
|
|
# After (structure preserved, long values compressed)
|
|
{"id": "usr_abc123", "name": "Alice Johnson", "bio": "A long...[compressed]..."}
|
|
```
|
|
|
|
### Code Handler
|
|
|
|
Preserves code structure using AST parsing (tree-sitter) or regex fallback:
|
|
|
|
```python
|
|
from headroom.compression.handlers.code_handler import CodeStructureHandler
|
|
|
|
handler = CodeStructureHandler(
|
|
preserve_comments=False, # Preserve comments as structural
|
|
use_tree_sitter=True, # Use tree-sitter for parsing
|
|
default_language="python", # Default when detection fails
|
|
)
|
|
```
|
|
|
|
**What's Preserved:**
|
|
- Import statements
|
|
- Function/method signatures
|
|
- Class definitions
|
|
- Type annotations
|
|
- Decorators
|
|
|
|
**What's Compressed:**
|
|
- Function bodies (implementations)
|
|
- Comments (unless `preserve_comments=True`)
|
|
|
|
**Example:**
|
|
|
|
```python
|
|
# Before
|
|
def process_data(items: List[str]) -> Dict[str, int]:
|
|
"""Process items and count occurrences."""
|
|
result = {}
|
|
for item in items:
|
|
item = item.strip().lower()
|
|
if item in result:
|
|
result[item] += 1
|
|
else:
|
|
result[item] = 1
|
|
return result
|
|
|
|
# After (signature preserved, body compressed)
|
|
def process_data(items: List[str]) -> Dict[str, int]:
|
|
"""Process items and count occurrences."""
|
|
result = {}
|
|
for item in items:
|
|
...[compressed]...
|
|
```
|
|
|
|
### Supported Languages
|
|
|
|
| Language | Parser | Support Level |
|
|
|----------|--------|---------------|
|
|
| Python | tree-sitter | Full AST |
|
|
| JavaScript | tree-sitter | Full AST |
|
|
| TypeScript | tree-sitter | Full AST |
|
|
| Go | tree-sitter | Full AST |
|
|
| Rust | tree-sitter | Full AST |
|
|
| Java | tree-sitter | Full AST |
|
|
| C | tree-sitter | Full AST |
|
|
| C++ | tree-sitter | Full AST |
|
|
|
|
---
|
|
|
|
## Compression Result
|
|
|
|
```python
|
|
from headroom.compression import compress
|
|
|
|
result = compress(content)
|
|
|
|
# Access result fields
|
|
print(result.compressed) # Compressed content
|
|
print(result.original) # Original content
|
|
print(result.compression_ratio) # e.g., 0.35 (35% of original size)
|
|
print(result.tokens_before) # Estimated tokens before
|
|
print(result.tokens_after) # Estimated tokens after
|
|
print(result.tokens_saved) # tokens_before - tokens_after
|
|
print(result.savings_percentage) # e.g., 65.0 (65% savings)
|
|
|
|
# Detection info
|
|
print(result.content_type) # ContentType.JSON, CODE, etc.
|
|
print(result.detection_confidence) # 0.0-1.0
|
|
|
|
# Structure info
|
|
print(result.handler_used) # "json", "code", etc.
|
|
print(result.preservation_ratio) # Fraction preserved as structure
|
|
|
|
# CCR info
|
|
print(result.ccr_key) # Key for retrieval (if CCR enabled)
|
|
```
|
|
|
|
---
|
|
|
|
## Batch Compression
|
|
|
|
For multiple contents, batch compression is more efficient:
|
|
|
|
```python
|
|
from headroom.compression import UniversalCompressor
|
|
|
|
compressor = UniversalCompressor()
|
|
|
|
contents = [
|
|
'{"users": [...]}',
|
|
"def hello(): pass",
|
|
"Plain text content",
|
|
]
|
|
|
|
results = compressor.compress_batch(contents)
|
|
|
|
for result in results:
|
|
print(f"{result.content_type}: {result.savings_percentage:.0f}% saved")
|
|
```
|
|
|
|
---
|
|
|
|
## Custom Handlers
|
|
|
|
Register custom handlers for specific content types:
|
|
|
|
```python
|
|
from headroom.compression import UniversalCompressor
|
|
from headroom.compression.detector import ContentType
|
|
from headroom.compression.handlers.base import BaseStructureHandler, HandlerResult
|
|
from headroom.compression.masks import StructureMask
|
|
|
|
|
|
class LogStructureHandler(BaseStructureHandler):
|
|
"""Custom handler for log content."""
|
|
|
|
def __init__(self):
|
|
super().__init__(name="log")
|
|
|
|
def can_handle(self, content: str) -> bool:
|
|
return "[INFO]" in content or "[ERROR]" in content
|
|
|
|
def _extract_mask(self, content, tokens, **kwargs):
|
|
# Mark timestamps and log levels as structural
|
|
mask = [False] * len(content)
|
|
# ... (custom logic)
|
|
return HandlerResult(
|
|
mask=StructureMask(tokens=tokens, mask=mask),
|
|
handler_name=self.name,
|
|
confidence=0.9,
|
|
)
|
|
|
|
|
|
# Register the custom handler
|
|
compressor = UniversalCompressor()
|
|
compressor.register_handler(ContentType.TEXT, LogStructureHandler())
|
|
```
|
|
|
|
---
|
|
|
|
## CCR Integration
|
|
|
|
Universal Compression integrates with CCR (Compress-Cache-Retrieve) for reversible compression:
|
|
|
|
```python
|
|
from headroom.compression import UniversalCompressor, UniversalCompressorConfig
|
|
|
|
config = UniversalCompressorConfig(ccr_enabled=True)
|
|
compressor = UniversalCompressor(config=config)
|
|
|
|
result = compressor.compress(large_content)
|
|
|
|
# CCR key for retrieval
|
|
if result.ccr_key:
|
|
print(f"Original stored with key: {result.ccr_key}")
|
|
# LLM can request original via CCR when needed
|
|
```
|
|
|
|
See [CCR Guide](ccr.md) for full CCR documentation.
|
|
|
|
---
|
|
|
|
## Performance
|
|
|
|
| Content Type | Compression | Speed | Accuracy |
|
|
|--------------|-------------|-------|----------|
|
|
| JSON (large arrays) | 70-90% | ~1ms | Keys preserved |
|
|
| Code (Python) | 50-70% | ~10ms | Signatures preserved |
|
|
| Plain text | 60-80% | ~5ms | High-entropy preserved |
|
|
|
|
**Overhead:** ~1-10ms per compression depending on content size and type.
|
|
|
|
---
|
|
|
|
## Installation
|
|
|
|
```bash
|
|
# Basic compression (fallback to simple compression)
|
|
pip install headroom-ai
|
|
|
|
# With ML detection (recommended)
|
|
pip install "headroom-ai[magika]"
|
|
|
|
# With LLMLingua compression
|
|
pip install "headroom-ai[llmlingua]"
|
|
|
|
# With AST-based code handling
|
|
pip install "headroom-ai[code]"
|
|
|
|
# Everything
|
|
pip install "headroom-ai[all]"
|
|
```
|
|
|
|
---
|
|
|
|
## Example: Full Pipeline
|
|
|
|
```python
|
|
from headroom.compression import UniversalCompressor, UniversalCompressorConfig
|
|
|
|
# Configure for aggressive compression
|
|
config = UniversalCompressorConfig(
|
|
compression_ratio_target=0.25, # Keep 25%
|
|
use_magika=True,
|
|
use_llmlingua=True,
|
|
ccr_enabled=True,
|
|
)
|
|
|
|
compressor = UniversalCompressor(config=config)
|
|
|
|
# Compress JSON API response
|
|
json_content = """
|
|
{
|
|
"users": [
|
|
{"id": "usr_123", "name": "Alice", "bio": "Software engineer..."},
|
|
{"id": "usr_456", "name": "Bob", "bio": "Product manager..."}
|
|
],
|
|
"total": 2,
|
|
"page": 1
|
|
}
|
|
"""
|
|
|
|
result = compressor.compress(json_content)
|
|
|
|
print(f"Type: {result.content_type}") # ContentType.JSON
|
|
print(f"Handler: {result.handler_used}") # json
|
|
print(f"Saved: {result.savings_percentage:.0f}%") # ~60%
|
|
print(f"Structure: {result.preservation_ratio:.0%} preserved") # ~40%
|
|
print(f"CCR Key: {result.ccr_key}") # For retrieval
|
|
```
|
|
|
|
---
|
|
|
|
## See Also
|
|
|
|
- [Transforms Reference](transforms.md) - Other compression transforms
|
|
- [CCR Guide](ccr.md) - Reversible compression architecture
|
|
- [Text Compression](text-compression.md) - Opt-in utilities for search/logs
|