feat(config): fold TOML array-of-tables to csv-schema via SmartCrusher (#1799)

## Description

Adds a schema-fold tier to the structured-config compressor (introduced
in #1784). TOML files containing an `[[array-of-tables]]` are parsed
with the stdlib `tomllib` reference parser and bridged to SmartCrusher's
lossless `csv-schema` renderer, which folds the repeated per-record keys
into a single schema over the rows. On lockfiles and override-lists —
where the repeated keys dominate the byte count — this is a large win.

**Stacked on #1784 — review only the top commit** (`feat(config): fold
TOML array-of-tables to csv-schema via SmartCrusher`). The base commit
is #1784's config-compressor PR; this PR will collapse to the single new
commit once #1784 merges.

Faithfulness is guaranteed by construction, not by a heuristic:
- `tomllib` is the reference TOML parser, so the extracted records are
ground-truth.
- `csv-schema` is a lossless JSON renderer (`smart_crusher.py` documents
it as such), so the model reads a faithful, reformatted view of the
exact parsed data.
- Byte-exact recovery rides the existing CCR path — the original is
persisted to the `CompressionStore` and a `Retrieve original: hash=…`
marker is emitted. The fold is only emitted when that store write
succeeds, so nothing is ever unrecoverable.

Scope is deliberately **TOML-only**: `tomllib` is stdlib, whereas PyYAML
is only a *transitive* dependency (not declared in `pyproject.toml`),
and INI record-sections would need a bespoke dict-of-dicts→records
transform. Those flavors can follow in a separate PR with an explicit
dependency decision.

## Type of Change

- [x] New feature (non-breaking change which adds functionality)

## Changes Made

- `headroom/transforms/config_compressor.py`: added Tier 3
(`_schema_fold`) — TOML `[[array-of-tables]]` → `tomllib` → JSON →
`SmartCrusher(csv-schema)`. New `enable_schema_fold` config flag
(default on; auto-off in lossless mode since it rides `enable_ccr`). The
fold competes with the reversible text tiers and is adopted only when
strictly smaller. Added `_load_toml` (stdlib parser with tomli backport)
and `_json_default` (TOML date/time → ISO; bail on any other
non-serializable value).
- Recovery reuses the existing `CompressionStore` + `Retrieve original:
hash=` marker; no new CCR plumbing.
- `tests/test_transforms_config_compressor.py`: 14 new tests covering
the fold, big-win assertion, byte-exact CCR round-trip, lossless-mode
disable, flag-off, non-TOML skip, no-array skip, small-array
`passthrough` decline, store-failure fallback, savings-floor rejection,
unparseable/non-serializable bails, `_load_toml`/`_json_default` units,
and a datetime-valued fold.

## Testing

- [x] New and existing unit tests pass locally
- [x] New tests added for the new behavior

### Test Output

```
tests/test_transforms_config_compressor.py ............................. [ 59%]
headroom/transforms/config_compressor.py     127      0     36      0   100%
============================== 49 passed in 0.61s ==============================
```

Must-stay-green suites (`test_lossless_mode`,
`test_lossless_excluded_compaction`,
`test_transforms_content_detection`,
`test_compression_fidelity_regression`) — 48 passed.
Router/tabular/smart_crusher regression — 73 + 82 passed. `mypy
--strict` clean on the changed module.

## Real Behavior Proof

- Environment: local, Python 3.11.0, macOS (darwin), `HF_HUB_OFFLINE=1`
- Exact command / steps: parsed a 25-record `[[tool.mypy.overrides]]`
TOML through
`ConfigCompressor(ConfigCompressorConfig(enable_ccr=True)).compress()`,
then retrieved the CCR hash from the `CompressionStore`.
- Observed result: `strategy=config_schema_fold`, 2765 → 840 chars (30%
of original); the marker hash resolved to the byte-exact original
(`recovered == original` True); with `enable_ccr=False` (lossless mode)
the fold did not run and no marker was emitted; a 3-record long-valued
`[[package]]` array correctly declined (SmartCrusher `passthrough`).
- Not tested: the live proxy end-to-end path and non-TOML flavors
(YAML/INI schema folding is intentionally out of scope for this PR).

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Ashish 2026-07-15 12:58:27 -07:00 committed by GitHub
parent 12a9710665
commit 737b332129
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
5 changed files with 1166 additions and 4 deletions

View file

@ -0,0 +1,283 @@
"""Structured-config compressor for YAML/TOML/INI tool output.
Config files (k8s manifests, CI pipelines, pyproject/Cargo manifests, INI
files) are high-frequency agent payloads with heavy structural repetition,
but they have no native compressor magika tags them SOURCE_CODE and they
fall through to the lossy prose path. This module compresses them in two
lossless-first tiers:
- Tier 1 (reversible, allowed in no-CCR lossless mode): format-agnostic
text-level compaction via `lossless_compaction` identical-line runs and
repeated multi-line stanzas collapse behind exact-inverse markers, with the
round-trip self-verified before the result is adopted.
- Tier 2 (CCR-recoverable, default mode only): whole-line comments and blank
lines are elided behind a summary line carrying a ``Retrieve original:
hash=`` marker; the full original is persisted to the CompressionStore
first, so nothing is lost.
- Tier 3 (CCR-recoverable, default mode only): TOML files with an
``[[array-of-tables]]`` are parsed with the stdlib ``tomllib`` reference
parser and bridged to SmartCrusher's lossless ``csv-schema`` renderer,
which folds the repeated per-record keys into a single schema. Only TOML
is bridged its stdlib parser makes the records ground-truth; YAML/INI
need non-stdlib or bespoke parsing and stay out of scope. Recovery again
rides a stored-original + ``Retrieve original: hash=`` marker.
No new compression algorithm and no new CCR plumbing live here Tier 1 rides
`compact_lossless`, Tier 2/Tier 3 ride the production `CompressionStore`, and
Tier 3's fold rides `SmartCrusher`.
"""
from __future__ import annotations
import datetime as dt
import json
import logging
import re
from dataclasses import dataclass
from typing import Any, cast
from .content_detector import ContentType, detect_content_type
from .lossless_compaction import compact_lossless
logger = logging.getLogger(__name__)
# Whole-line comment prefixes per flavor. INI values can span indented
# continuation lines, so INI only elides column-0 comment lines and keeps
# blanks (configparser keeps blank lines inside multi-line values).
_COMMENT_RES = {
"yaml": re.compile(r"^\s*#"),
"toml": re.compile(r"^\s*#"),
"ini": re.compile(r"^[#;]"),
}
# Content where whole-line elision is unsafe: a '#' line inside a YAML block
# scalar or a TOML multi-line string is data, not a comment. Detection is
# deliberately over-broad — when in doubt, Tier 2 stays off.
_YAML_BLOCK_SCALAR_RE = re.compile(r":\s*[|>][+-]?\d*\s*$", re.MULTILINE)
_TOML_MULTILINE_RE = re.compile(r'"""|\'\'\'')
def _load_toml(content: str) -> dict[str, Any] | None:
"""Parse TOML with the stdlib parser (or the tomli backport); None on error."""
try:
import tomllib
except ModuleNotFoundError: # pragma: no cover - Python < 3.11 only
try:
import tomli as tomllib # type: ignore[no-redef]
except ModuleNotFoundError:
return None
try:
return cast("dict[str, Any]", tomllib.loads(content))
except (tomllib.TOMLDecodeError, ValueError):
return None
def _json_default(value: Any) -> str:
"""Render TOML date/time values as ISO strings; bail on anything else."""
if isinstance(value, dt.datetime | dt.date | dt.time):
return value.isoformat()
raise TypeError(f"unserializable config value: {type(value).__name__}")
@dataclass
class ConfigCompressorConfig:
"""Configuration for structured-config compression."""
# Emit the CCR-marked comment/blank elision tier. The router wires this
# to its ccr_inject_marker setting; lossless mode turns it off.
enable_ccr: bool = True
# Bridge TOML array-of-tables to SmartCrusher csv-schema (Tier 3). Rides
# CCR for recovery, so it only runs when enable_ccr is also on.
enable_schema_fold: bool = True
# Only adopt a result that is strictly smaller than the original.
min_savings_chars: int = 1
@dataclass
class ConfigCompressionResult:
"""Result of structured-config compression."""
compressed: str
original: str
was_modified: bool
flavor: str # "yaml" | "toml" | "ini" | "unknown"
lines_elided: int = 0
ccr_hash: str | None = None
strategy: str = "config"
@property
def compression_ratio(self) -> float:
if not self.original:
return 0.0
return len(self.compressed) / len(self.original)
class ConfigCompressor:
"""Compresses YAML/TOML/INI text via reversible + CCR-recoverable tiers.
Public surface mirrors the other content-type compressors so the router
and tests treat it uniformly.
"""
def __init__(self, config: ConfigCompressorConfig | None = None) -> None:
self.config = config or ConfigCompressorConfig()
def compress(
self,
content: str,
context: str = "",
bias: float = 1.0,
) -> ConfigCompressionResult:
detection = detect_content_type(content)
if detection.content_type is not ContentType.STRUCTURED_CONFIG:
return ConfigCompressionResult(
compressed=content,
original=content,
was_modified=False,
flavor="unknown",
)
flavor = detection.metadata.get("flavor", "yaml")
# Tier 3: schema fold (TOML array-of-tables → csv-schema). Computed
# first so it can compete with the text tiers; it wins on lockfiles
# and override-lists where repeated keys dominate.
schema_fold: tuple[str, str] | None = None
if self.config.enable_ccr and self.config.enable_schema_fold:
schema_fold = self._schema_fold(content, flavor, context, bias)
working = content
lines_elided = 0
ccr_hash: str | None = None
# Tier 2: comment/blank elision behind a CCR marker. Persist first —
# the elided lines are only droppable because the original is stored.
if self.config.enable_ccr and self._elision_safe(content, flavor):
stripped, elided = self._strip_comment_lines(content, flavor)
if elided > 0:
ccr_hash = self._store_original(content, stripped)
if ccr_hash is not None:
marker = (
f"[{elided} comment/blank lines elided. Retrieve original: hash={ccr_hash}]"
)
working = stripped + ("" if stripped.endswith("\n") else "\n") + marker
lines_elided = elided
# Tier 1: reversible run/stanza folding; self-verified round-trip.
compressed = compact_lossless(working, "config")
# Prefer the schema fold when it beats the text tiers.
if schema_fold is not None and len(schema_fold[0]) < len(compressed):
return ConfigCompressionResult(
compressed=schema_fold[0],
original=content,
was_modified=True,
flavor=flavor,
ccr_hash=schema_fold[1],
strategy="config_schema_fold",
)
savings = len(content) - len(compressed)
if savings < self.config.min_savings_chars:
return ConfigCompressionResult(
compressed=content,
original=content,
was_modified=False,
flavor=flavor,
)
return ConfigCompressionResult(
compressed=compressed,
original=content,
was_modified=True,
flavor=flavor,
lines_elided=lines_elided,
ccr_hash=ccr_hash,
)
def _schema_fold(
self, content: str, flavor: str, context: str, bias: float
) -> tuple[str, str] | None:
"""Fold a TOML array-of-tables into SmartCrusher csv-schema.
Returns ``(folded_text_with_marker, ccr_hash)`` when the fold is a
strictly-smaller, faithful representation whose original is safely
stored for CCR retrieval; otherwise ``None`` so the caller keeps the
text tiers. Only TOML is bridged: ``tomllib`` is the reference parser,
so the extracted records are ground-truth and the csv-schema renderer
is itself lossless the model reads a faithful, reformatted view.
"""
if flavor != "toml" or "[[" not in content:
return None
data = _load_toml(content)
if data is None:
return None
try:
json_str = json.dumps(data, ensure_ascii=False, default=_json_default)
except TypeError:
return None # a value we can't represent faithfully → don't fold
from .smart_crusher import SmartCrusher
crusher = SmartCrusher(with_compaction=True, compaction_format="csv-schema")
result = crusher.crush(json_str, context, bias)
# `passthrough` means SmartCrusher only re-canonicalized the JSON and
# applied no schema fold, so there is nothing worth adopting.
if not result.was_modified or result.strategy == "passthrough":
return None
ccr_hash = self._store_original(content, result.compressed)
if ccr_hash is None:
return None # can't recover the original → never emit a lossy form
marker = f"[config folded to schema. Retrieve original: hash={ccr_hash}]"
folded = result.compressed + "\n" + marker
if len(content) - len(folded) < self.config.min_savings_chars:
return None
return folded, ccr_hash
@staticmethod
def _elision_safe(content: str, flavor: str) -> bool:
"""False when a '#' line could be data (block scalars, multi-line strings)."""
if flavor == "yaml":
return not _YAML_BLOCK_SCALAR_RE.search(content)
if flavor == "toml":
return not _TOML_MULTILINE_RE.search(content)
return True
@staticmethod
def _strip_comment_lines(content: str, flavor: str) -> tuple[str, int]:
"""Drop whole-line comments (and, outside INI, blank lines)."""
comment_re = _COMMENT_RES.get(flavor, _COMMENT_RES["yaml"])
keep_blanks = flavor == "ini"
had_trailing = content.endswith("\n")
lines = (content[:-1] if had_trailing else content).split("\n")
kept: list[str] = []
elided = 0
for line in lines:
if comment_re.match(line) or (not keep_blanks and not line.strip()):
elided += 1
else:
kept.append(line)
return "\n".join(kept) + ("\n" if had_trailing else ""), elided
@staticmethod
def _store_original(original: str, compressed: str) -> str | None:
"""Persist the original to the CompressionStore; returns its hash."""
try:
from ..cache.compression_store import get_compression_store
except ImportError as e: # pragma: no cover - store ships with headroom
logger.warning("CCR store import failed; config elision skipped: %s", e)
return None
try:
store: Any = get_compression_store()
stored = store.store(original, compressed, compression_strategy="config")
return str(stored) if stored else None
except Exception as e:
logger.warning("CCR store write failed; config elision skipped: %s", e)
return None
__all__ = [
"ConfigCompressor",
"ConfigCompressorConfig",
"ConfigCompressionResult",
]

View file

@ -10,6 +10,7 @@ Supported content types:
- SEARCH_RESULTS: grep/ripgrep output (file:line:content)
- BUILD_OUTPUT: Compiler, test, lint logs
- GIT_DIFF: Unified diff format
- STRUCTURED_CONFIG: YAML/TOML/INI config files
- PLAIN_TEXT: Generic text (fallback)
"""
@ -31,6 +32,7 @@ class ContentType(Enum):
GIT_DIFF = "diff" # Unified diff format
HTML = "html" # Web pages (needs content extraction, not compression)
TABULAR = "tabular" # CSV/TSV, markdown tables, fixed-width tables
STRUCTURED_CONFIG = "structured_config" # YAML/TOML/INI config files
PLAIN_TEXT = "text" # Fallback
@ -111,6 +113,18 @@ _CODE_PATTERNS = {
],
}
# Structured-config (YAML/TOML/INI) patterns. TOML and INI share the
# `[section]` header shape; the stdlib parsers disambiguate. YAML is
# heuristic-only (PyYAML is not a dependency): key/list/document-marker
# line share plus structure signals, with prose and front-matter guards.
_CONFIG_SECTION_RE = re.compile(r"^\s*\[\[?[\w.\-\"' ]+\]\]?\s*$")
_TOML_ASSIGN_RE = re.compile(r"""^\s*(?:[\w.\-]+|"[^"]+"|'[^']+')\s*=\s*\S""")
_INI_ASSIGN_RE = re.compile(r"^\s*[\w.\-@ ]+?\s*[=:]\s*")
_YAML_KEY_RE = re.compile(r"""^\s*(?:-\s+)?(?:[\w.\-/]+|"[^"]+"|'[^']+')\s*:(?:\s|$)""")
_YAML_LIST_RE = re.compile(r"^\s*-\s+\S")
_YAML_DOC_RE = re.compile(r"^---\s*$|^\.\.\.\s*$")
_CONFIG_COMMENT_RE = re.compile(r"^\s*[#;]")
# Log/build output patterns
_LOG_PATTERNS = [
re.compile(r"\b(ERROR|FAIL|FAILED|FATAL|CRITICAL)\b", re.IGNORECASE),
@ -181,12 +195,19 @@ def detect_content_type(content: str) -> DetectionResult:
if tabular_result and tabular_result.confidence >= 0.6:
return tabular_result
# 7. Check for source code
# 7. Check for structured config (YAML/TOML/INI). Runs after tabular so
# delimited data keeps its claim, and before code so config files with
# code-ish lines route to the structure-aware config compressor.
config_result = _try_detect_structured_config(content)
if config_result and config_result.confidence >= 0.6:
return config_result
# 8. Check for source code
code_result = _try_detect_code(content)
if code_result and code_result.confidence >= 0.5:
return code_result
# 8. Fallback to plain text
# 9. Fallback to plain text
return DetectionResult(ContentType.PLAIN_TEXT, 0.5, {})
@ -581,6 +602,128 @@ def _try_detect_tabular(content: str) -> DetectionResult | None:
return _try_detect_delimited(lines)
def _try_parse_toml(content: str) -> bool:
"""True if `content` parses as TOML (stdlib tomllib, or the tomli backport)."""
try:
import tomllib
except ModuleNotFoundError: # pragma: no cover - Python < 3.11
try:
import tomli as tomllib # type: ignore[no-redef]
except ModuleNotFoundError:
return False
try:
tomllib.loads(content)
return True
except Exception:
return False
def _parse_config_flavor(content: str) -> str | None:
"""Disambiguate `[section]`-shaped config: TOML first, then INI.
Both flavors share the section-header line shape; only the stdlib parsers
can tell them apart reliably. Returns "toml", "ini", or None when neither
parser accepts the content (then it is not claimed as config at all).
"""
if len(content) > 1_000_000:
return None
if _try_parse_toml(content):
return "toml"
import configparser
parser = configparser.ConfigParser(interpolation=None, strict=False)
try:
parser.read_string(content)
except Exception:
return None
return "ini" if parser.sections() else None
def _try_detect_structured_config(content: str) -> DetectionResult | None:
"""Try to detect structured config content (YAML, TOML, INI).
TOML/INI claims are parser-confirmed (stdlib), so they carry high
confidence. YAML has no stdlib parser, so its claim is heuristic:
key/list/document-marker line share plus a structure signal, guarded
against prose and markdown front-matter.
"""
head = content.lstrip()[:1]
if not head or head in "{<":
# JSON objects and markup are never config; JSON arrays and real
# TOML/INI `[section]` headers disambiguate below.
return None
lines = content.split("\n")[:200]
non_empty = [ln for ln in lines if ln.strip()]
if len(non_empty) < 3:
return None
# Comment lines are neutral: excluded from the line-share ratio so
# comment-heavy configs and #-heading markdown don't skew it either way.
body = [ln for ln in non_empty if not _CONFIG_COMMENT_RE.match(ln)]
if len(body) < 3:
return None
# TOML / INI: require a section header plus assignment-dominant body,
# then let the stdlib parsers confirm and disambiguate.
sections = sum(1 for ln in body if _CONFIG_SECTION_RE.match(ln))
if sections >= 1:
assigns = sum(1 for ln in body if _TOML_ASSIGN_RE.match(ln) or _INI_ASSIGN_RE.match(ln))
if assigns >= 2 and (sections + assigns) / len(body) >= 0.6:
flavor = _parse_config_flavor(content)
if flavor is not None:
share = (sections + assigns) / len(body)
return DetectionResult(
ContentType.STRUCTURED_CONFIG,
min(0.95, 0.7 + share * 0.25),
{"flavor": flavor, "sections": sections, "assignments": assigns},
)
# Markdown front-matter guard: a `---` fence closed within 60 lines and
# followed by non-YAML content is a markdown document, not standalone YAML.
if lines and lines[0].strip() == "---":
for idx in range(1, min(len(lines), 60)):
if lines[idx].strip() in ("---", "..."):
tail = [ln for ln in lines[idx + 1 :] if ln.strip()]
tail_yaml = sum(
1 for ln in tail if _YAML_KEY_RE.match(ln) or _YAML_LIST_RE.match(ln)
)
if tail and tail_yaml / len(tail) < 0.3:
return None
break
# YAML heuristic.
yaml_keys = sum(1 for ln in body if _YAML_KEY_RE.match(ln))
yaml_lists = sum(1 for ln in body if _YAML_LIST_RE.match(ln) and not _YAML_KEY_RE.match(ln))
doc_marks = sum(1 for ln in body if _YAML_DOC_RE.match(ln.strip()))
if yaml_keys < 3:
return None
share = (yaml_keys + yaml_lists + doc_marks) / len(body)
if share < 0.6:
return None
# Prose guards: config lines are short field-ish tuples, prose reads like
# sentences (mirrors _looks_like_prose for delimited data).
enders = sum(1 for ln in body if ln.rstrip().endswith((".", "!", "?")))
if enders / len(body) >= 0.5:
return None
avg_words = sum(len(ln.split()) for ln in body) / len(body)
if avg_words > 8:
return None
# Structure signal: nested indentation, a document marker, or a real list.
indents = {
len(ln) - len(ln.lstrip(" "))
for ln in body
if _YAML_KEY_RE.match(ln) or _YAML_LIST_RE.match(ln)
}
if len(indents) < 2 and doc_marks == 0 and yaml_lists < 3:
return None
return DetectionResult(
ContentType.STRUCTURED_CONFIG,
min(0.9, 0.55 + share * 0.35),
{"flavor": "yaml", "keys": yaml_keys, "list_items": yaml_lists},
)
def _try_detect_code(content: str) -> DetectionResult | None:
"""Try to detect source code and identify language."""
lines = content.split("\n")[:100] # Check first 100 lines

View file

@ -62,7 +62,13 @@ from ..tokenizer import Tokenizer
from ..tokenizers.estimator import EstimatingTokenCounter
from . import mixed_content as _mixed_content
from .base import Transform
from .content_detector import ContentType, DetectionResult, _try_detect_log, _try_detect_search
from .content_detector import (
ContentType,
DetectionResult,
_try_detect_log,
_try_detect_search,
_try_detect_structured_config,
)
from .content_detector import detect_content_type as _regex_detect_content_type
from .error_detection import content_has_strong_error_indicators
from .mixed_content import ContentSection, mixed_content_indicators
@ -565,6 +571,15 @@ def _detect_content(content: str) -> DetectionResult:
if override is not None:
return override
# Config misroute guard (native/magika path): magika classifies YAML/TOML/
# INI as SourceCode, which routes to the (default-disabled) code path and
# degrades to prose compression. When the structural config detector
# positively claims the payload, trust it over the SourceCode verdict.
if content_type is ContentType.SOURCE_CODE:
config_override = _try_detect_structured_config(content)
if config_override is not None and config_override.confidence >= 0.7:
return config_override
if content_type is ContentType.PLAIN_TEXT:
regex_result = _regex_detect_content_type(content)
if regex_result.content_type is not ContentType.PLAIN_TEXT:
@ -923,6 +938,7 @@ class CompressionStrategy(Enum):
DIFF = "diff"
HTML = "html"
TABULAR = "tabular"
CONFIG = "config"
MIXED = "mixed"
PASSTHROUGH = "passthrough"
@ -1032,6 +1048,7 @@ class ContentRouterConfig:
enable_search_compressor: Enable search result compression.
enable_log_compressor: Enable build/test log compression.
enable_tabular_compressor: Enable CSV/TSV/markdown-table compression.
enable_config_compressor: Enable YAML/TOML/INI config compression.
enable_image_optimizer: Enable image token optimization.
prefer_code_aware_for_code: Use CodeAware over Kompress for code.
min_section_tokens: Minimum tokens for a section to compress.
@ -1048,6 +1065,7 @@ class ContentRouterConfig:
enable_search_compressor: bool = True
enable_log_compressor: bool = True
enable_tabular_compressor: bool = True # CSV/TSV/markdown tables via SmartCrusher
enable_config_compressor: bool = True # YAML/TOML/INI structural compression
enable_html_extractor: bool = True # HTML content extraction
enable_image_optimizer: bool = True # Image token optimization
@ -1320,6 +1338,7 @@ class ContentRouter(Transform):
self._diff_compressor: Any = None
self._html_extractor: Any = None
self._tabular_compressor: Any = None
self._config_compressor: Any = None
self._kompress: Any = None
# Stage B relevance split (lazy; None until first use, sentinel-checked
# via _relevance_scorer_tried so a failed load isn't retried per call).
@ -1833,6 +1852,7 @@ class ContentRouter(Transform):
ContentType.GIT_DIFF: CompressionStrategy.DIFF,
ContentType.HTML: CompressionStrategy.HTML,
ContentType.TABULAR: CompressionStrategy.TABULAR,
ContentType.STRUCTURED_CONFIG: CompressionStrategy.CONFIG,
ContentType.PLAIN_TEXT: CompressionStrategy.TEXT,
}
@ -1991,6 +2011,7 @@ class ContentRouter(Transform):
CompressionStrategy.SEARCH: "search",
CompressionStrategy.LOG: "log",
CompressionStrategy.DIFF: "diff",
CompressionStrategy.CONFIG: "config",
}.get(strategy)
order = ([primary] if primary else []) + [
k for k in ("search", "paths", "log", "diff", "text") if k != primary
@ -2274,6 +2295,18 @@ class ContentRouter(Transform):
)
decision_reason = "tabular_compressor"
elif strategy == CompressionStrategy.CONFIG:
if self.config.enable_config_compressor:
compressor = self._get_config_compressor()
if compressor:
compressor_name = type(compressor).__name__
result = compressor.compress(content, context=context, bias=bias)
compressed, compressed_tokens = (
result.compressed,
len(result.compressed.split()),
)
decision_reason = "config_compressor"
elif strategy == CompressionStrategy.DIFF:
compressor = self._get_diff_compressor()
if compressor:
@ -2325,6 +2358,7 @@ class ContentRouter(Transform):
CompressionStrategy.SMART_CRUSHER,
CompressionStrategy.CODE_AWARE,
CompressionStrategy.TABULAR,
CompressionStrategy.CONFIG,
}
fallback_no_savings = compressed == content or compressed_tokens >= original_tokens
if fallback_eligible_strategy and fallback_no_savings:
@ -2655,6 +2689,7 @@ class ContentRouter(Transform):
ContentType.GIT_DIFF: CompressionStrategy.DIFF,
ContentType.HTML: CompressionStrategy.HTML,
ContentType.TABULAR: CompressionStrategy.TABULAR,
ContentType.STRUCTURED_CONFIG: CompressionStrategy.CONFIG,
ContentType.PLAIN_TEXT: CompressionStrategy.TEXT,
}
return mapping.get(content_type, self.config.fallback_strategy)
@ -2669,6 +2704,7 @@ class ContentRouter(Transform):
CompressionStrategy.DIFF: ContentType.GIT_DIFF,
CompressionStrategy.HTML: ContentType.HTML,
CompressionStrategy.TABULAR: ContentType.TABULAR,
CompressionStrategy.CONFIG: ContentType.STRUCTURED_CONFIG,
CompressionStrategy.TEXT: ContentType.PLAIN_TEXT,
CompressionStrategy.KOMPRESS: ContentType.PLAIN_TEXT,
CompressionStrategy.PASSTHROUGH: ContentType.PLAIN_TEXT,
@ -2893,6 +2929,19 @@ class ContentRouter(Transform):
logger.debug("TabularCompressor not available")
return self._tabular_compressor
def _get_config_compressor(self) -> Any:
"""Get ConfigCompressor (lazy load)."""
if self._config_compressor is None:
try:
from .config_compressor import ConfigCompressor, ConfigCompressorConfig
self._config_compressor = ConfigCompressor(
ConfigCompressorConfig(enable_ccr=self.config.ccr_inject_marker)
)
except ImportError: # pragma: no cover - defensive; module is pure stdlib
logger.debug("ConfigCompressor not available")
return self._config_compressor
def _get_diff_compressor(self) -> Any:
"""Get DiffCompressor (lazy load). Rust-only — Python implementation
retired in Stage 3b. The wheel (`headroom._core`) is a hard import.

View file

@ -21,6 +21,8 @@ __all__ = [
"collapse_runs",
"expand_runs",
"is_run_collapsed",
"fold_repeated_blocks",
"unfold_repeated_blocks",
"search_heading",
"search_unheading",
"diff_strip_index",
@ -34,6 +36,20 @@ _ANSI_RE = re.compile(r"\x1b\[[0-9;]*m")
# syslog-style run-collapse marker. The count is captured for exact inversion.
_RUN_MARKER_RE = re.compile(r"^\.\.\. \(repeated (\d+) times\)$")
# multi-line block back-reference marker. Length and distance (both in lines,
# in ORIGINAL coordinates) are captured for exact inversion: everything before
# a marker expands to the exact original prefix, so `distance` lines back in
# the expanded output is the block's first occurrence.
_BLOCK_MARKER_RE = re.compile(r"^\.\.\. \(repeats (\d+) lines from (\d+) lines back\)$")
# fold_repeated_blocks search bounds: minimum/maximum block length worth a
# marker, candidate anchors per line, and an input size cap so the scan stays
# negligible on huge payloads.
_FOLD_MIN_BLOCK = 3
_FOLD_MAX_BLOCK = 64
_FOLD_MAX_CANDIDATES = 8
_FOLD_MAX_LINES = 20_000
# grep/ripgrep default row shape: ``path:line:content``. ``line`` is digits;
# ``path`` must not itself look like ``line:content`` (i.e. not start with a
# bare number) so we don't mis-split a heading-form ``line:content`` row.
@ -127,6 +143,79 @@ def is_run_collapsed(text: str) -> bool:
return False
def fold_repeated_blocks(text: str) -> str:
"""Collapse multi-line blocks that repeat earlier content into back-refs.
The block-level generalization of :func:`collapse_runs`: a run of K
consecutive lines (K >= 3) that exactly reproduces K lines seen D lines
earlier becomes ``... (repeats K lines from D lines back)``. The repeats
need not be adjacent, which is what config payloads actually look like
k8s container stanzas repeat with only the ``name:`` line differing, so
their identical tails fold even though no two whole stanzas are
consecutive. Coordinates are in original lines: the fold is only taken
when the block does not overlap its anchor (K <= D), so on expansion the
referenced region is always already reconstructed.
Exact inverse: :func:`unfold_repeated_blocks`.
"""
lines, had_trailing = _split_keep_trailing(text)
n = len(lines)
if n < _FOLD_MIN_BLOCK * 2 or n > _FOLD_MAX_LINES:
return text
positions: dict[str, list[int]] = {}
out: list[str] = []
i = 0
while i < n:
best_len = 0
best_dist = 0
for q in reversed(positions.get(lines[i], ())):
max_len = min(_FOLD_MAX_BLOCK, n - i, i - q)
length = 0
while length < max_len and lines[q + length] == lines[i + length]:
length += 1
if length > best_len:
best_len = length
best_dist = i - q
if best_len >= _FOLD_MIN_BLOCK:
marker = f"... (repeats {best_len} lines from {best_dist} lines back)"
block_chars = sum(len(lines[i + k]) + 1 for k in range(best_len))
if block_chars > len(marker) + 1:
out.append(marker)
for k in range(best_len):
_remember(positions, lines[i + k], i + k)
i += best_len
continue
_remember(positions, lines[i], i)
out.append(lines[i])
i += 1
return _join(out, had_trailing)
def _remember(positions: dict[str, list[int]], line: str, index: int) -> None:
"""Track recent original positions of `line`, bounded per distinct line."""
bucket = positions.setdefault(line, [])
bucket.append(index)
if len(bucket) > _FOLD_MAX_CANDIDATES:
del bucket[0]
def unfold_repeated_blocks(text: str) -> str:
"""Exact inverse of :func:`fold_repeated_blocks`."""
lines, had_trailing = _split_keep_trailing(text)
if not lines:
return text
out: list[str] = []
for line in lines:
m = _BLOCK_MARKER_RE.match(line)
if m:
length, dist = int(m.group(1)), int(m.group(2))
start = len(out) - dist
if start >= 0 and length <= dist:
out.extend(out[start : start + length])
continue
out.append(line)
return _join(out, had_trailing)
def search_heading(text: str) -> str:
"""Convert grep ``path:line:content`` rows into ripgrep --heading form.
@ -279,7 +368,7 @@ def _smaller(candidate: str, original: str) -> bool:
def compact_lossless(content: str, kind: str) -> str:
"""Dispatch format-native lossless compaction by ``kind``.
``kind`` in {'log', 'search', 'diff', 'text'}. For reversible kinds the
``kind`` in {'log', 'search', 'diff', 'text', 'config'}. For reversible kinds the
round-trip is verified internally (modulo the intentionally-dropped
non-semantic bits, e.g. ANSI color for logs); if verification fails or the
result is not smaller, the original content is returned unchanged. Never
@ -322,6 +411,14 @@ def compact_lossless(content: str, kind: str) -> str:
if expand_runs(candidate) != content:
return content
return candidate if _smaller(candidate, content) else content
if kind == "config":
# Structured config (YAML/TOML/INI): single-line runs first, then
# repeated multi-line stanzas. Inverse applies in reverse order.
candidate = fold_repeated_blocks(collapse_runs(content))
if expand_runs(unfold_repeated_blocks(candidate)) != content:
return content
return candidate if _smaller(candidate, content) else content
except Exception:
return content
return content

View file

@ -0,0 +1,590 @@
"""Tests for structured-config (YAML/TOML/INI) detection and compression.
Covers the detection heuristics (including prose / markdown-front-matter
non-claims), the reversible block-fold primitives in lossless_compaction,
the two-tier ConfigCompressor, and the router wiring.
"""
from __future__ import annotations
import pytest
from headroom.parser import CCR_RETRIEVAL_MARKER_RE
from headroom.transforms.config_compressor import (
ConfigCompressor,
ConfigCompressorConfig,
)
from headroom.transforms.content_detector import (
ContentType,
_try_detect_structured_config,
detect_content_type,
)
from headroom.transforms.content_router import (
CompressionStrategy,
ContentRouter,
ContentRouterConfig,
)
from headroom.transforms.lossless_compaction import (
compact_lossless,
expand_runs,
fold_repeated_blocks,
unfold_repeated_blocks,
)
# Reusable fixtures ----------------------------------------------------------
K8S_MANIFEST = """apiVersion: apps/v1
kind: Deployment
metadata:
name: web
labels:
app: web
spec:
replicas: 3
template:
spec:
containers:
- name: web-1
image: nginx:1.25
resources:
limits:
cpu: 500m
memory: 512Mi
requests:
cpu: 250m
memory: 256Mi
- name: web-2
image: nginx:1.25
resources:
limits:
cpu: 500m
memory: 512Mi
requests:
cpu: 250m
memory: 256Mi
- name: web-3
image: nginx:1.25
resources:
limits:
cpu: 500m
memory: 512Mi
requests:
cpu: 250m
memory: 256Mi
"""
PYPROJECT_TOML = """# Build configuration
[package]
name = "demo"
version = "0.1.0"
edition = "2021"
[dependencies]
serde = "1.0"
tokio = "1.38"
[profile.release]
opt-level = 3
lto = true
"""
INI_CONFIG = """[server]
host = 127.0.0.1
port = 8080
; connection tuning
[logging]
level = INFO
file = /var/log/app.log
[auth]
enabled = true
provider = ldap
"""
PROSE = """Note: this document describes the deployment process.
First, you should review the configuration carefully before applying it.
The rollout takes several minutes to complete in most environments.
If anything goes wrong, roll back to the previous release immediately.
Contact the on-call engineer when the dashboard shows sustained errors.
"""
FRONT_MATTER_DOC = """---
title: My Post
tags:
- a
- b
---
# Heading
This is a markdown document body with prose that goes on and on.
More prose here, explaining things in complete English sentences.
And a final line of text to round out the document nicely today.
"""
# Detection -------------------------------------------------------------------
@pytest.mark.parametrize(
"content,flavor",
[
(K8S_MANIFEST, "yaml"),
(PYPROJECT_TOML, "toml"),
(INI_CONFIG, "ini"),
],
)
def test_detects_structured_config(content: str, flavor: str) -> None:
result = detect_content_type(content)
assert result.content_type is ContentType.STRUCTURED_CONFIG
assert result.metadata["flavor"] == flavor
assert result.confidence >= 0.7
@pytest.mark.parametrize(
"content",
[
PROSE,
FRONT_MATTER_DOC,
# grep output: colon shapes must stay SEARCH_RESULTS
"src/main.py:42:def process():\nsrc/util.py:10:import os\nsrc/x.py:5:return 1",
# CSV keeps its tabular claim
"name,age,city\nalice,30,nyc\nbob,25,sf\ncarol,35,la",
# JSON array of dicts stays with SmartCrusher
'[{"id": 1, "name": "a"}, {"id": 2, "name": "b"}]',
# JSON object bodies are never claimed as config
'{\n "key": "value",\n "other": "thing",\n "third": "entry"\n}',
# Python code with colon-ended keywords
"import os\n\n\ndef main():\n if True:\n return os.name\n",
# too short
"key: value",
],
)
def test_does_not_claim_non_config(content: str) -> None:
result = detect_content_type(content)
assert result.content_type is not ContentType.STRUCTURED_CONFIG
def test_multi_document_yaml_stream_is_claimed() -> None:
doc = "---\nname: a\nimage: x:1\nports:\n - 80\n---\nname: b\nimage: y:2\nports:\n - 443\n"
result = _try_detect_structured_config(doc)
assert result is not None
assert result.metadata["flavor"] == "yaml"
def test_toml_beats_ini_when_both_parse() -> None:
# Valid TOML with quoted strings parses under tomllib and wins.
result = _try_detect_structured_config('[a]\nx = "1"\ny = "2"\n[b]\nz = "3"\n')
assert result is not None
assert result.metadata["flavor"] == "toml"
def test_ini_when_toml_rejects_bare_values() -> None:
# Unquoted strings are invalid TOML but fine for configparser.
result = _try_detect_structured_config("[a]\nx = hello\ny = world\n[b]\nz = there\n")
assert result is not None
assert result.metadata["flavor"] == "ini"
def test_flat_yaml_without_structure_not_claimed() -> None:
# Flat key: value lines with one indent level, no docs, no lists: too
# ambiguous with prose-ish "Key: value" notes to claim.
result = _try_detect_structured_config("alpha: 1\nbeta: 2\ngamma: 3\n")
assert result is None
# fold_repeated_blocks / unfold_repeated_blocks -------------------------------
def test_fold_round_trip_k8s() -> None:
folded = fold_repeated_blocks(K8S_MANIFEST)
assert len(folded) < len(K8S_MANIFEST)
assert "lines back)" in folded
assert unfold_repeated_blocks(folded) == K8S_MANIFEST
def test_fold_handles_consecutive_identical_blocks() -> None:
block = "alpha: value-one\nbeta: value-two\ngamma: value-three\n"
text = block * 4
folded = fold_repeated_blocks(text)
assert unfold_repeated_blocks(folded) == text
assert len(folded) < len(text)
def test_fold_skips_short_blocks() -> None:
text = "x: 1\ny: 2\nx: 1\ny: 2\n" # repeats are only 2 lines long
assert fold_repeated_blocks(text) == text
def test_fold_skips_when_marker_not_smaller() -> None:
# Three repeated single-char lines: folding would cost more than it saves.
text = "a\nb\nc\na\nb\nc\n"
assert fold_repeated_blocks(text) == text
def test_unfold_leaves_invalid_marker_untouched() -> None:
text = "line one\n... (repeats 5 lines from 2 lines back)\n"
# length > distance: not a marker fold_repeated_blocks could have emitted.
assert unfold_repeated_blocks(text) == text
def test_compact_lossless_config_verifies_round_trip() -> None:
compacted = compact_lossless(K8S_MANIFEST, "config")
assert len(compacted) < len(K8S_MANIFEST)
assert expand_runs(unfold_repeated_blocks(compacted)) == K8S_MANIFEST
def test_compact_lossless_config_bails_on_marker_collision() -> None:
# Content that already contains a marker-shaped line cannot round-trip,
# so the self-check must return the original unchanged.
lines = [f"item: {i}" for i in range(3)]
block = "\n".join(lines)
text = block + "\n... (repeats 3 lines from 3 lines back)\n" + block + "\n" + block + "\n"
assert compact_lossless(text, "config") == text
# ConfigCompressor ------------------------------------------------------------
def test_tier1_lossless_only_no_marker() -> None:
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=False))
result = comp.compress(K8S_MANIFEST)
assert result.was_modified
assert len(result.compressed) < len(K8S_MANIFEST)
assert not CCR_RETRIEVAL_MARKER_RE.search(result.compressed)
assert result.flavor == "yaml"
assert result.ccr_hash is None
def test_tier2_elides_comments_behind_ccr_marker(monkeypatch) -> None:
from headroom.cache.compression_store import CompressionStore
store = CompressionStore()
monkeypatch.setattr("headroom.cache.compression_store.get_compression_store", lambda: store)
commented = PYPROJECT_TOML.replace(
'serde = "1.0"',
"# pinned for CVE-2024-0001; do not bump without checking the advisory\n"
"# see https://example.com/advisories/CVE-2024-0001 for details\n"
"# owner: platform-team, revisit after the 2.x migration lands\n"
'serde = "1.0"',
)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
result = comp.compress(commented)
assert result.was_modified
assert result.lines_elided > 0
assert result.ccr_hash is not None
assert CCR_RETRIEVAL_MARKER_RE.search(result.compressed)
assert "CVE-2024-0001" not in result.compressed
# The marker hash must resolve to the byte-exact original.
assert store.retrieve(result.ccr_hash).original_content == commented
def test_tier2_skipped_for_yaml_block_scalars(monkeypatch) -> None:
from headroom.cache.compression_store import CompressionStore
monkeypatch.setattr(
"headroom.cache.compression_store.get_compression_store",
lambda: CompressionStore(),
)
doc = (
"config:\n"
" script: |\n"
" # this hash line is DATA inside a block scalar\n"
" echo hi\n"
" replicas: 3\n"
" image: nginx\n"
" ports:\n"
" - 80\n"
)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
result = comp.compress(doc)
# The '#' line inside the block scalar must survive verbatim.
assert "# this hash line is DATA" in result.compressed or not result.was_modified
def test_tier2_skipped_for_toml_multiline_strings(monkeypatch) -> None:
from headroom.cache.compression_store import CompressionStore
monkeypatch.setattr(
"headroom.cache.compression_store.get_compression_store",
lambda: CompressionStore(),
)
doc = '[a]\ntext = """\n# not a comment\n"""\nx = "1"\ny = "2"\n'
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
result = comp.compress(doc)
assert "# not a comment" in result.compressed or not result.was_modified
def test_tier2_ini_keeps_blank_lines(monkeypatch) -> None:
from headroom.cache.compression_store import CompressionStore
monkeypatch.setattr(
"headroom.cache.compression_store.get_compression_store",
lambda: CompressionStore(),
)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
result = comp.compress(INI_CONFIG)
if result.was_modified:
# Column-0 ';' comment goes; blank section separators stay.
assert "; connection tuning" not in result.compressed
assert "\n\n" in result.compressed
def test_store_failure_degrades_to_tier1(monkeypatch) -> None:
class _BrokenStore:
def store(self, *a, **kw): # noqa: ANN002, ANN003
raise RuntimeError("disk full")
monkeypatch.setattr(
"headroom.cache.compression_store.get_compression_store",
lambda: _BrokenStore(),
)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
result = comp.compress(K8S_MANIFEST)
# Tier 1 still folds the manifest; nothing was elided.
assert result.was_modified
assert result.ccr_hash is None
assert result.lines_elided == 0
assert not CCR_RETRIEVAL_MARKER_RE.search(result.compressed)
def test_non_config_content_passes_through() -> None:
comp = ConfigCompressor()
result = comp.compress(PROSE)
assert not result.was_modified
assert result.compressed == PROSE
assert result.flavor == "unknown"
def test_no_savings_returns_original() -> None:
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=False))
result = comp.compress(PYPROJECT_TOML) # compact already; nothing to fold
assert not result.was_modified
assert result.compressed == PYPROJECT_TOML
def test_compression_ratio_zero_for_empty_original() -> None:
from headroom.transforms.config_compressor import ConfigCompressionResult
result = ConfigCompressionResult(
compressed="", original="", was_modified=False, flavor="unknown"
)
assert result.compression_ratio == 0.0
# Tier 3: TOML array-of-tables → SmartCrusher csv-schema -----------------------
def _mypy_overrides_toml(n: int) -> str:
"""A pyproject-style TOML whose ``[[overrides]]`` array dominates the file."""
records = "\n\n".join(
f"[[tool.mypy.overrides]]\n"
f'module = "pkg.sub{i}.mod"\n'
f"ignore_missing_imports = true\n"
f"disallow_untyped_defs = false"
for i in range(n)
)
return "[tool.mypy]\nstrict = true\n\n" + records
def _use_fresh_store(monkeypatch): # noqa: ANN001, ANN201
from headroom.cache.compression_store import CompressionStore
store = CompressionStore()
monkeypatch.setattr("headroom.cache.compression_store.get_compression_store", lambda: store)
return store
def test_tier3_folds_toml_array_of_tables(monkeypatch) -> None:
store = _use_fresh_store(monkeypatch)
toml = _mypy_overrides_toml(25)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
result = comp.compress(toml)
assert result.was_modified
assert result.strategy == "config_schema_fold"
assert result.flavor == "toml"
assert len(result.compressed) < len(toml) // 2 # keys folded → big win
assert 0.0 < result.compression_ratio < 0.5
# The repeated per-record key appears once in the schema, not 25 times.
assert result.compressed.count("ignore_missing_imports") == 1
assert CCR_RETRIEVAL_MARKER_RE.search(result.compressed)
# The marker hash resolves to the byte-exact original.
assert result.ccr_hash is not None
assert store.retrieve(result.ccr_hash).original_content == toml
def test_tier3_disabled_in_lossless_mode() -> None:
# enable_ccr off (lossless) → no fold, no marker, only reversible tiers.
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=False))
result = comp.compress(_mypy_overrides_toml(25))
assert result.strategy != "config_schema_fold"
assert not CCR_RETRIEVAL_MARKER_RE.search(result.compressed)
def test_tier3_flag_off_keeps_text_tiers(monkeypatch) -> None:
_use_fresh_store(monkeypatch)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True, enable_schema_fold=False))
result = comp.compress(_mypy_overrides_toml(25))
assert result.strategy != "config_schema_fold"
def test_tier3_skips_non_toml_flavor(monkeypatch) -> None:
_use_fresh_store(monkeypatch)
# A YAML list-of-mappings is structurally similar but must not be bridged
# (no stdlib YAML parser is a dependency).
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
result = comp.compress(K8S_MANIFEST)
assert result.strategy != "config_schema_fold"
def test_tier3_skips_toml_without_array_of_tables(monkeypatch) -> None:
_use_fresh_store(monkeypatch)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
# Valid TOML flavor, but no `[[ ]]` → SmartCrusher has no array to fold.
result = comp.compress(PYPROJECT_TOML)
assert result.strategy != "config_schema_fold"
def test_tier3_declines_small_array_as_passthrough(monkeypatch) -> None:
_use_fresh_store(monkeypatch)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
# Three long-valued records: SmartCrusher returns `passthrough` (the folded
# keys don't outweigh the unique values), so Tier 3 declines.
toml = "version = 3\n\n" + "\n\n".join(
f'[[package]]\nname = "crate-{i}"\nversion = "1.2.{i}"\nchecksum = "{i:064x}"'
for i in range(3)
)
result = comp.compress(toml)
assert result.strategy != "config_schema_fold"
def test_tier3_store_failure_keeps_text_tiers(monkeypatch) -> None:
class _BrokenStore:
def store(self, *a, **kw): # noqa: ANN002, ANN003
raise RuntimeError("disk full")
monkeypatch.setattr(
"headroom.cache.compression_store.get_compression_store",
lambda: _BrokenStore(),
)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
result = comp.compress(_mypy_overrides_toml(25))
# Can't store the original → never emit the lossy fold.
assert result.strategy != "config_schema_fold"
assert not CCR_RETRIEVAL_MARKER_RE.search(result.compressed)
def test_tier3_rejected_when_marker_overhead_exceeds_savings(monkeypatch) -> None:
_use_fresh_store(monkeypatch)
# An unreachable-in-practice savings floor forces the final gate to reject
# even a genuine fold, exercising the guard.
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True, min_savings_chars=10**6))
result = comp.compress(_mypy_overrides_toml(25))
assert result.strategy != "config_schema_fold"
def test_schema_fold_bails_on_unparseable_toml(monkeypatch) -> None:
_use_fresh_store(monkeypatch)
import headroom.transforms.config_compressor as mod
monkeypatch.setattr(mod, "_load_toml", lambda _content: None)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
# Detector still flags it TOML, but the fold parser bails → text tiers.
assert comp._schema_fold(_mypy_overrides_toml(25), "toml", "", 1.0) is None
def test_schema_fold_bails_on_non_serializable_value(monkeypatch) -> None:
_use_fresh_store(monkeypatch)
import headroom.transforms.config_compressor as mod
monkeypatch.setattr(mod, "_load_toml", lambda _content: {"t": [{"v": {1, 2}}, {"v": {3, 4}}]})
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
# A value tomllib could never emit but that we can't represent faithfully.
assert comp._schema_fold("[[t]]\n", "toml", "", 1.0) is None
def test_load_toml_returns_none_on_invalid() -> None:
from headroom.transforms.config_compressor import _load_toml
assert _load_toml('[[x]]\nk = "unterminated') is None
assert _load_toml('[[x]]\nk = "ok"') == {"x": [{"k": "ok"}]}
def test_json_default_renders_dates_and_rejects_others() -> None:
import datetime as dt
from headroom.transforms.config_compressor import _json_default
assert _json_default(dt.date(2026, 7, 4)) == "2026-07-04"
assert _json_default(dt.datetime(2026, 7, 4, 9, 30)) == "2026-07-04T09:30:00"
with pytest.raises(TypeError):
_json_default(object())
def test_tier3_toml_with_datetime_folds(monkeypatch) -> None:
_use_fresh_store(monkeypatch)
# TOML datetimes serialize through _json_default; the fold still applies.
toml = "\n\n".join(
f'[[event]]\nname = "e{i}"\nwhen = 2026-07-04T09:30:00\nactive = true' for i in range(20)
)
comp = ConfigCompressor(ConfigCompressorConfig(enable_ccr=True))
result = comp.compress(toml)
assert result.strategy == "config_schema_fold"
assert "2026-07-04T09:30:00" in result.compressed
# Router wiring ---------------------------------------------------------------
def test_router_maps_structured_config_to_config_strategy() -> None:
router = ContentRouter()
assert (
router._strategy_from_detection_type(ContentType.STRUCTURED_CONFIG)
is CompressionStrategy.CONFIG
)
assert (
router._content_type_from_strategy(CompressionStrategy.CONFIG)
is ContentType.STRUCTURED_CONFIG
)
def test_router_lazy_getter_mirrors_ccr_setting() -> None:
router = ContentRouter(ContentRouterConfig(lossless=True))
compressor = router._get_config_compressor()
assert compressor is not None
assert compressor.config.enable_ccr is False # lossless forces markers off
def test_router_compresses_k8s_manifest_end_to_end() -> None:
router = ContentRouter()
result = router.compress(K8S_MANIFEST)
assert result.compressed != K8S_MANIFEST or result.strategy_used in (
CompressionStrategy.CONFIG,
CompressionStrategy.PASSTHROUGH,
)
# Whatever the gates decide, the output must never be a lossy mangle:
# either untouched or a reversible fold of the manifest.
if result.compressed != K8S_MANIFEST:
assert expand_runs(unfold_repeated_blocks(result.compressed)) == K8S_MANIFEST
def test_router_lossless_mode_uses_lossless_config_label() -> None:
router = ContentRouter(ContentRouterConfig(lossless=True))
compressed, _tokens, chain = router._apply_strategy_to_content(
K8S_MANIFEST, CompressionStrategy.CONFIG, context=""
)
assert chain == ["lossless_config"]
assert expand_runs(unfold_repeated_blocks(compressed)) == K8S_MANIFEST
assert not CCR_RETRIEVAL_MARKER_RE.search(compressed)
def test_router_disabled_flag_skips_config_compressor() -> None:
router = ContentRouter(ContentRouterConfig(enable_config_compressor=False))
compressed, _tokens, _chain = router._apply_strategy_to_content(
K8S_MANIFEST, CompressionStrategy.CONFIG, context=""
)
# Falls through to the unified fallback path; must not crash and must
# not claim CONFIG did work it didn't do.
assert isinstance(compressed, str)