## Description
Reported by [@JoaoMarcos44](https://github.com/JoaoMarcos44) via an
independent security audit — thanks for the careful, well-documented
report.
Fixes two confirmed findings from a June 2026 security audit of
`headroom/compression/` (the `UniversalCompressor` utility). Both are
real defects in shipped, public, tested code; note that this module is
**not** on the proxy hot path (the proxy uses `headroom/transforms/`),
so real-world blast radius is module-local rather than proxy-wide.
- **SEC-01 (entropy bypass):** `use_entropy_preservation` was a silent
no-op. `compress()` tokenized content at character level
(`list(content)`) and fed single-char tokens to `compute_entropy_mask`,
whose `min_token_length` guard skipped every one — so high-entropy
secrets (API keys, OAuth tokens, UUIDs, hashes) were never preserved
despite the feature being enabled.
- **SEC-02 (JSON corruption):** the `_simple_compress` truncation
fallback (used when Kompress is unavailable or raises) inserted a
separator containing raw newlines. When that fallback ran on a span
inside a JSON string value it produced invalid JSON (RFC 8259 §7),
crashing downstream `json.loads()`.
The other three audited items need no code change and were verified, not
assumed: SEC-03 (surrogate DoS) is already caught by the `try/except` in
`code_handler._extract_mask` and falls back to regex — non-reproducible
even with `tree_sitter_language_pack` installed; SEC-04 (prompt
injection) is out of a compressor's scope; SEC-05 (SQLite race) is a
misread (`CompressionStore` defaults to `InMemoryBackend`; the SQLite
backend uses WAL + busy_timeout + a lock).
Closes #
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
## Changes Made
- Add `compute_entropy_mask_for_content()` (`masks.py`): scores
whitespace-delimited words and maps high-entropy ones back to character
positions, returning a char-aligned mask. The existing token-level
`compute_entropy_mask` is left intact.
- Introduce `SECRET_ENTROPY_MIN_LENGTH = 20` as the default word-length
floor. Normalized Shannon entropy rates short-but-diverse words (e.g.
"detailed") nearly as high as a real secret, so a length floor is the
discriminator; 20 matches the entropy-detection floor used by secret
scanners (trufflehog, detect-secrets) and prevents over-preserving prose
(which would otherwise block legitimate compression).
- Wire the content-level entropy pass into
`UniversalCompressor.compress()` (scores `content`, not the char-level
`tokens`).
- Replace the `_simple_compress` separator `"\n...[compressed]...\n"`
with the control-char-free `" ...[compressed]... "`.
- Add regression tests at the mask level and end-to-end.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ ruff check headroom/compression/
All checks passed!
$ mypy headroom/compression/masks.py headroom/compression/universal.py
Success: no issues found in 2 source files
$ pytest tests/test_compression/test_masks.py tests/test_compression/test_universal.py \
tests/test_compression/test_json_handler.py tests/test_compression/test_code_handler.py -q
======================= 111 passed, 2 warnings in 10.76s =======================
```
## Real Behavior Proof
- Environment: macOS, Python 3.12 in repo `.venv`;
`tree_sitter_language_pack` and Kompress present.
- Exact command / steps: reproduced each finding by calling
`UniversalCompressor.compress()` directly before/after the fix — SEC-01:
`compute_entropy_mask(list("k="+secret))` preserved 0 of N tokens
(inert); after fix `compute_entropy_mask_for_content` preserves the
secret's char range and the end-to-end test shows a 43-char secret
dropped with preservation off / kept with it on. SEC-02:
`compress(json.dumps({...long value...}), content_type=JSON)` with
`use_kompress=False` raised `JSONDecodeError` before the fix and
round-trips through `json.loads()` after.
- Observed result: SEC-01 entropy preservation now functions; SEC-02
output is valid JSON on both the Kompress and fallback paths; the
previously-failing `test_compression_reduces_tokens` passes again (no
over-preservation).
- Not tested: `tests/test_compression/test_evals.py` and
`test_llm_eval.py` (require external API/model access); the
proxy/transforms live path is unaffected since it does not import
`UniversalCompressor`.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable
## Additional Notes
CHANGELOG not updated (handled by the release tooling). The audit also
flagged SEC-03/04/05 — left unchanged by design, with verification
rationale in the Description.
LLMLingua was the original ML text compressor (BERT-based). Kompress
(ModernBERT, trained on 330K structured tool outputs) replaced it with
better compression quality and simpler architecture.
Removed across 35 files:
- Deleted headroom/transforms/llmlingua_compressor.py
- Deleted tests/test_transforms/test_llmlingua_compressor.py
- Deleted tests/test_proxy_llmlingua.py
- Removed all enable_llmlingua config, _get_llmlingua methods,
LLMLingua fallback paths, LLMLINGUA strategy enum values
- Removed CLI flags, model configs, compression handler references
- Simplified ContentRouter: Kompress is primary and only text compressor