mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
## Description Headroom's lossy compression drops rows/lines using statistical heuristics but **never checks that meaning survived** — a dropped `OOM killed worker 3` line can silently flip a model's answer with no signal that compression caused it. The repo already ships a quality-metric toolkit (`headroom/evals/metrics.py`) and a `weekly-suite` eval job, but neither gates the compression path on a PR. This adds a **per-PR fidelity regression gate**: compress vendored golden tool-outputs through SmartCrusher's lossy path and assert the evidence that answers each case's question survives. It is the first of a planned trio (this is the "offline gate" half of the fidelity work); query-aware retention and a hard token-budget API are documented follow-ups. Closes # ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **Blocking gate** (`tests/test_compression_fidelity_regression.py`): compresses each golden case via `smart_crush_tool_output(..., with_compaction=False)` and scores with `compute_information_recall`. Two assertions: - **Per-case critical recall == 1.0** — every `answer_evidence` string (placed in error/anomaly rows, the documented SmartCrusher retention guarantee) must survive. - **Aggregate recall ≥ committed baseline** (`baseline.json`, tol 0.02) — catches softer regressions. - **Vendored fixtures** (`tests/fixtures/fidelity_golden/`): deterministic `_generate.py` emits `cases.json` (4 cases: OOM crash, payment exception, latency anomaly, CI failure) + `baseline.json`. - **Non-blocking weekly report** (`.github/workflows/eval.yml`): one step in the existing `weekly-suite` job (schedule/manual only) reuses the existing `evaluate_information_retention` runner for a recall report on the production routing path. - **Pure reuse**: scoring (`evals/metrics.py`), compressor (`smart_crush_tool_output`), and the weekly runner (`evaluate_information_retention`) all already existed. ### Design notes - **Zero new CI setup.** The blocking gate runs in the existing `[dev]` test shard — no new workflow, no new deps, **no model, no network, no secrets** (verified under `HF_HUB_OFFLINE=1`). It deliberately uses small hand-made structured fixtures rather than the repo's HuggingFace dataset loaders, which would require a network download + ModernBERT and don't belong in a fast PR gate. - **Scope:** structured JSON tool-output (the dominant, deterministic, model-free case). Real-dataset (HotpotQA/BFCL) recall — which needs `[all]` + a local model — is a **documented follow-up PR**, and the `weekly-suite` job (which genuinely runs every Monday) is its natural home. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ HF_HUB_OFFLINE=1 python -m pytest tests/test_compression_fidelity_regression.py -v tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[logs_oom] PASSED tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[payment_exception] PASSED tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[latency_anomaly] PASSED tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[ci_test_failures] PASSED tests/test_compression_fidelity_regression.py::test_aggregate_recall_not_regressed PASSED ============================== 5 passed in 0.18s =============================== ``` ## Real Behavior Proof - **Environment:** local checkout of `feat/fidelity-regression-gate`, `pip install -e ".[dev]"`, `HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1` (proves no model/network). - **Exact command / steps:** `HF_HUB_OFFLINE=1 python -m pytest tests/test_compression_fidelity_regression.py -q` → `5 passed in 0.14s`. - **Negative control (proves the gate has teeth):** compressing `logs_oom` and probing for a benign row that compression legitimately drops returns `recall = 0.00, lost = ['heartbeat ping 25']` — i.e. the gate fires when critical evidence is dropped, so it is not trivially green. - **Weekly (non-blocking) step verified locally:** ```text Information retention: 50/50 cases >=0.9 recall, avg compression 65.7% ``` - **Not tested:** real-dataset (HotpotQA/BFCL) recall and prose/ModernBERT compression — intentionally deferred to a follow-up PR targeting the weekly job. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable ## Additional Notes - CHANGELOG/version intentionally untouched: repo uses **release-please**. - **Follow-up PR (planned):** wire the real HotpotQA/BFCL loaders (`headroom/evals/datasets.py`) into the `weekly-suite` job for genuine benchmark-scale recall coverage (model-allowed, non-blocking). Further follow-ups from the same design: a live per-request fidelity guardrail, query-aware lossy retention, and a hard `target_tokens` budget API. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
72 lines
3.1 KiB
Python
72 lines
3.1 KiB
Python
"""Offline fidelity regression gate (recall-based, zero-model).
|
|
|
|
Compresses vendored golden tool-output fixtures through SmartCrusher's lossy
|
|
path and asserts that the evidence a model needs to answer each case's question
|
|
survives compression. Scoring is pure stdlib (``headroom.evals.metrics``) — no
|
|
ML model, no network, no API keys — so this runs in the standard ``[dev]`` CI
|
|
shard as a blocking PR check.
|
|
|
|
A failure here means a code change made lossy compression silently drop
|
|
information that answers a known question. Fixtures and the committed baseline
|
|
are generated by ``tests/fixtures/fidelity_golden/_generate.py``.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
from pathlib import Path
|
|
|
|
import pytest
|
|
|
|
from headroom.evals.metrics import compute_information_recall
|
|
from headroom.transforms.smart_crusher import SmartCrusherConfig, smart_crush_tool_output
|
|
|
|
FIXTURE_DIR = Path(__file__).parent / "fixtures" / "fidelity_golden"
|
|
CASES: list[dict] = json.loads((FIXTURE_DIR / "cases.json").read_text())
|
|
BASELINE: dict = json.loads((FIXTURE_DIR / "baseline.json").read_text())
|
|
|
|
|
|
def _compress(case: dict) -> tuple[str, str]:
|
|
"""Compress a case's tool output via the lossy SmartCrusher path (no model)."""
|
|
original = json.dumps(case["content"])
|
|
cfg = SmartCrusherConfig(max_items_after_crush=case["compress"]["max_items_after_crush"])
|
|
crushed, _modified, _info = smart_crush_tool_output(
|
|
original, cfg, with_compaction=case["compress"]["with_compaction"]
|
|
)
|
|
return original, crushed
|
|
|
|
|
|
@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])
|
|
def test_critical_evidence_survives_compression(case: dict) -> None:
|
|
"""Every ``answer_evidence`` string MUST survive lossy compression (recall == 1.0).
|
|
|
|
Critical evidence lives in error/anomaly rows, which SmartCrusher formally
|
|
guarantees to retain (see ``tests/test_quality_retention.py``).
|
|
"""
|
|
original, crushed = _compress(case)
|
|
result = compute_information_recall(original, crushed, case["answer_evidence"])
|
|
assert result["recall"] == 1.0, (
|
|
f"FIDELITY REGRESSION in '{case['id']}': compression dropped evidence "
|
|
f"needed to answer {case['question']!r}. Lost: {result['facts_lost']}"
|
|
)
|
|
|
|
|
|
def test_aggregate_recall_not_regressed() -> None:
|
|
"""Mean recall over all evidence must not fall below the committed baseline.
|
|
|
|
Catches softer regressions (e.g. relevant-but-non-critical context being
|
|
dropped more aggressively) that the per-case critical gate would not.
|
|
"""
|
|
recalls = []
|
|
for case in CASES:
|
|
original, crushed = _compress(case)
|
|
probes = case["answer_evidence"] + case["supporting_facts"]
|
|
recalls.append(compute_information_recall(original, crushed, probes)["recall"])
|
|
|
|
mean_recall = sum(recalls) / len(recalls)
|
|
floor = BASELINE["aggregate_recall"] - BASELINE["tolerance"]
|
|
assert mean_recall >= floor, (
|
|
f"FIDELITY REGRESSION: mean recall {mean_recall:.4f} fell below baseline "
|
|
f"floor {floor:.4f} (baseline {BASELINE['aggregate_recall']} - tolerance "
|
|
f"{BASELINE['tolerance']}). If this drop is intended, regenerate the baseline."
|
|
)
|