mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
test(evals): add offline fidelity regression gate (recall-based, zero-model) (#1187)
## Description Headroom's lossy compression drops rows/lines using statistical heuristics but **never checks that meaning survived** — a dropped `OOM killed worker 3` line can silently flip a model's answer with no signal that compression caused it. The repo already ships a quality-metric toolkit (`headroom/evals/metrics.py`) and a `weekly-suite` eval job, but neither gates the compression path on a PR. This adds a **per-PR fidelity regression gate**: compress vendored golden tool-outputs through SmartCrusher's lossy path and assert the evidence that answers each case's question survives. It is the first of a planned trio (this is the "offline gate" half of the fidelity work); query-aware retention and a hard token-budget API are documented follow-ups. Closes # ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **Blocking gate** (`tests/test_compression_fidelity_regression.py`): compresses each golden case via `smart_crush_tool_output(..., with_compaction=False)` and scores with `compute_information_recall`. Two assertions: - **Per-case critical recall == 1.0** — every `answer_evidence` string (placed in error/anomaly rows, the documented SmartCrusher retention guarantee) must survive. - **Aggregate recall ≥ committed baseline** (`baseline.json`, tol 0.02) — catches softer regressions. - **Vendored fixtures** (`tests/fixtures/fidelity_golden/`): deterministic `_generate.py` emits `cases.json` (4 cases: OOM crash, payment exception, latency anomaly, CI failure) + `baseline.json`. - **Non-blocking weekly report** (`.github/workflows/eval.yml`): one step in the existing `weekly-suite` job (schedule/manual only) reuses the existing `evaluate_information_retention` runner for a recall report on the production routing path. - **Pure reuse**: scoring (`evals/metrics.py`), compressor (`smart_crush_tool_output`), and the weekly runner (`evaluate_information_retention`) all already existed. ### Design notes - **Zero new CI setup.** The blocking gate runs in the existing `[dev]` test shard — no new workflow, no new deps, **no model, no network, no secrets** (verified under `HF_HUB_OFFLINE=1`). It deliberately uses small hand-made structured fixtures rather than the repo's HuggingFace dataset loaders, which would require a network download + ModernBERT and don't belong in a fast PR gate. - **Scope:** structured JSON tool-output (the dominant, deterministic, model-free case). Real-dataset (HotpotQA/BFCL) recall — which needs `[all]` + a local model — is a **documented follow-up PR**, and the `weekly-suite` job (which genuinely runs every Monday) is its natural home. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ HF_HUB_OFFLINE=1 python -m pytest tests/test_compression_fidelity_regression.py -v tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[logs_oom] PASSED tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[payment_exception] PASSED tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[latency_anomaly] PASSED tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[ci_test_failures] PASSED tests/test_compression_fidelity_regression.py::test_aggregate_recall_not_regressed PASSED ============================== 5 passed in 0.18s =============================== ``` ## Real Behavior Proof - **Environment:** local checkout of `feat/fidelity-regression-gate`, `pip install -e ".[dev]"`, `HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1` (proves no model/network). - **Exact command / steps:** `HF_HUB_OFFLINE=1 python -m pytest tests/test_compression_fidelity_regression.py -q` → `5 passed in 0.14s`. - **Negative control (proves the gate has teeth):** compressing `logs_oom` and probing for a benign row that compression legitimately drops returns `recall = 0.00, lost = ['heartbeat ping 25']` — i.e. the gate fires when critical evidence is dropped, so it is not trivially green. - **Weekly (non-blocking) step verified locally:** ```text Information retention: 50/50 cases >=0.9 recall, avg compression 65.7% ``` - **Not tested:** real-dataset (HotpotQA/BFCL) recall and prose/ModernBERT compression — intentionally deferred to a follow-up PR targeting the weekly job. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable ## Additional Notes - CHANGELOG/version intentionally untouched: repo uses **release-please**. - **Follow-up PR (planned):** wire the real HotpotQA/BFCL loaders (`headroom/evals/datasets.py`) into the `weekly-suite` job for genuine benchmark-scale recall coverage (model-allowed, non-blocking). Further follow-ups from the same design: a live per-request fidelity guardrail, query-aware lossy retention, and a hard `target_tokens` budget API. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
6d116b15f1
commit
23d73ae070
5 changed files with 1032 additions and 0 deletions
18
.github/workflows/eval.yml
vendored
18
.github/workflows/eval.yml
vendored
|
|
@ -126,6 +126,24 @@ jobs:
|
|||
env:
|
||||
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
||||
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
|
||||
|
||||
# Recall-based fidelity report on the production routing path. Zero cost
|
||||
# (synthetic structured cases -> Rust compressors; no model, no API, no
|
||||
# secrets). Non-blocking: surfaces recall trends weekly without gating.
|
||||
# The blocking per-PR fidelity gate lives in
|
||||
# tests/test_compression_fidelity_regression.py (runs in the [dev] shard).
|
||||
- name: Information-retention recall report (zero cost, non-blocking)
|
||||
run: |
|
||||
python -c "
|
||||
from headroom.evals.runners.compression_only import CompressionOnlyRunner
|
||||
runner = CompressionOnlyRunner()
|
||||
cases = runner.generate_info_retention_cases(n=50)
|
||||
result = runner.evaluate_information_retention(cases)
|
||||
print(f'Information retention: {result.passed_cases}/{result.total_cases} cases >=0.9 recall, avg compression {result.avg_compression_ratio:.1%}')
|
||||
if not result.passed:
|
||||
print(f'::warning title=Fidelity recall::{result.failed_cases} case(s) fell below 0.9 recall: {result.errors[:3]}')
|
||||
"
|
||||
|
||||
- name: Upload results
|
||||
if: always()
|
||||
uses: actions/upload-artifact@v7
|
||||
|
|
|
|||
141
tests/fixtures/fidelity_golden/_generate.py
vendored
Normal file
141
tests/fixtures/fidelity_golden/_generate.py
vendored
Normal file
|
|
@ -0,0 +1,141 @@
|
|||
"""Deterministic generator for the fidelity-regression golden fixtures.
|
||||
|
||||
Run to (re)create ``cases.json``::
|
||||
|
||||
python tests/fixtures/fidelity_golden/_generate.py
|
||||
|
||||
Each case is a JSON tool output (a list of records) plus the question it should
|
||||
let a model answer and the evidence that must survive compression. Critical
|
||||
``answer_evidence`` is always placed in an *error* or *anomaly* row, because
|
||||
those are the rows SmartCrusher formally guarantees to retain
|
||||
(see ``tests/test_quality_retention.py``) — so the hard ``recall == 1.0`` gate
|
||||
holds on a healthy tree. ``supporting_facts`` are softer details that may live
|
||||
in droppable rows; they feed the aggregate-recall baseline, not the hard gate.
|
||||
|
||||
No models, no network: content is built from fixed strings only.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
OUT = Path(__file__).parent / "cases.json"
|
||||
|
||||
|
||||
def _filler(n: int, msg: str) -> list[dict]:
|
||||
"""n benign, near-identical INFO rows (droppable noise)."""
|
||||
return [{"seq": i, "level": "INFO", "status": "ok", "msg": f"{msg} {i}"} for i in range(n)]
|
||||
|
||||
|
||||
def _cases() -> list[dict]:
|
||||
cases: list[dict] = []
|
||||
|
||||
# 1. OOM crash buried in benign logs — error-keyword retention.
|
||||
rows = _filler(28, "heartbeat ping")
|
||||
rows.insert(
|
||||
17,
|
||||
{
|
||||
"seq": 17,
|
||||
"level": "ERROR",
|
||||
"status": "failed",
|
||||
"msg": "OOM killed worker 3 exceeded memory limit 8GiB",
|
||||
},
|
||||
)
|
||||
rows.insert(
|
||||
9, {"seq": 9, "level": "INFO", "status": "ok", "msg": "checkpoint saved at step 4500"}
|
||||
)
|
||||
cases.append(
|
||||
{
|
||||
"id": "logs_oom",
|
||||
"question": "Why did the job fail?",
|
||||
"content_type": "json_array",
|
||||
"compress": {"with_compaction": False, "max_items_after_crush": 10},
|
||||
"answer_evidence": ["OOM killed worker 3", "exceeded memory limit"],
|
||||
"supporting_facts": ["checkpoint saved at step 4500"],
|
||||
"content": rows,
|
||||
}
|
||||
)
|
||||
|
||||
# 2. Payment exception — error-keyword retention.
|
||||
rows = _filler(30, "GET /healthz 200")
|
||||
rows.insert(
|
||||
21,
|
||||
{
|
||||
"seq": 21,
|
||||
"level": "ERROR",
|
||||
"status": "failed",
|
||||
"msg": "exception: NullPointerException at PaymentService.charge line 88",
|
||||
},
|
||||
)
|
||||
cases.append(
|
||||
{
|
||||
"id": "payment_exception",
|
||||
"question": "Which service threw an exception and where?",
|
||||
"content_type": "json_array",
|
||||
"compress": {"with_compaction": False, "max_items_after_crush": 8},
|
||||
"answer_evidence": ["NullPointerException", "PaymentService.charge line 88"],
|
||||
"supporting_facts": [],
|
||||
"content": rows,
|
||||
}
|
||||
)
|
||||
|
||||
# 3. Latency anomaly — anomaly (>2 sigma) retention. Values cluster near 100ms,
|
||||
# one row spikes to 99999ms.
|
||||
rows = [
|
||||
{"seq": i, "region": "us-west-2", "latency_ms": 95 + (i % 11), "status": "ok"}
|
||||
for i in range(30)
|
||||
]
|
||||
rows[19] = {
|
||||
"seq": 19,
|
||||
"region": "us-east-1",
|
||||
"latency_ms": 99999,
|
||||
"status": "ok",
|
||||
"note": "latency spike us-east-1",
|
||||
}
|
||||
cases.append(
|
||||
{
|
||||
"id": "latency_anomaly",
|
||||
"question": "Which region had the latency spike, and how high?",
|
||||
"content_type": "json_array",
|
||||
"compress": {"with_compaction": False, "max_items_after_crush": 8},
|
||||
"answer_evidence": ["us-east-1", "99999"],
|
||||
"supporting_facts": [],
|
||||
"content": rows,
|
||||
}
|
||||
)
|
||||
|
||||
# 4. CI test failure among many passes — "failed" keyword retention.
|
||||
rows = [
|
||||
{"seq": i, "test": f"test_module_{i}", "outcome": "passed", "duration_ms": 5 + i}
|
||||
for i in range(30)
|
||||
]
|
||||
rows[12] = {
|
||||
"seq": 12,
|
||||
"test": "test_auth_token_refresh",
|
||||
"outcome": "failed",
|
||||
"duration_ms": 41,
|
||||
"error": "AssertionError: expected 200 got 401 in test_auth_token_refresh",
|
||||
}
|
||||
cases.append(
|
||||
{
|
||||
"id": "ci_test_failures",
|
||||
"question": "Which test failed and why?",
|
||||
"content_type": "json_array",
|
||||
"compress": {"with_compaction": False, "max_items_after_crush": 9},
|
||||
"answer_evidence": ["test_auth_token_refresh", "expected 200 got 401"],
|
||||
"supporting_facts": [],
|
||||
"content": rows,
|
||||
}
|
||||
)
|
||||
|
||||
return cases
|
||||
|
||||
|
||||
def main() -> None:
|
||||
OUT.write_text(json.dumps(_cases(), indent=2) + "\n")
|
||||
print(f"wrote {OUT} ({len(_cases())} cases)")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
5
tests/fixtures/fidelity_golden/baseline.json
vendored
Normal file
5
tests/fixtures/fidelity_golden/baseline.json
vendored
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
{
|
||||
"aggregate_recall": 0.9167,
|
||||
"tolerance": 0.02,
|
||||
"note": "Mean information-recall across all golden cases (answer_evidence + supporting_facts) after lossy SmartCrusher compression. Deterministic. Regenerate fixtures via `python tests/fixtures/fidelity_golden/_generate.py`; update this value only when a recall change is intended and understood."
|
||||
}
|
||||
796
tests/fixtures/fidelity_golden/cases.json
vendored
Normal file
796
tests/fixtures/fidelity_golden/cases.json
vendored
Normal file
|
|
@ -0,0 +1,796 @@
|
|||
[
|
||||
{
|
||||
"id": "logs_oom",
|
||||
"question": "Why did the job fail?",
|
||||
"content_type": "json_array",
|
||||
"compress": {
|
||||
"with_compaction": false,
|
||||
"max_items_after_crush": 10
|
||||
},
|
||||
"answer_evidence": [
|
||||
"OOM killed worker 3",
|
||||
"exceeded memory limit"
|
||||
],
|
||||
"supporting_facts": [
|
||||
"checkpoint saved at step 4500"
|
||||
],
|
||||
"content": [
|
||||
{
|
||||
"seq": 0,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 0"
|
||||
},
|
||||
{
|
||||
"seq": 1,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 1"
|
||||
},
|
||||
{
|
||||
"seq": 2,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 2"
|
||||
},
|
||||
{
|
||||
"seq": 3,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 3"
|
||||
},
|
||||
{
|
||||
"seq": 4,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 4"
|
||||
},
|
||||
{
|
||||
"seq": 5,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 5"
|
||||
},
|
||||
{
|
||||
"seq": 6,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 6"
|
||||
},
|
||||
{
|
||||
"seq": 7,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 7"
|
||||
},
|
||||
{
|
||||
"seq": 8,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 8"
|
||||
},
|
||||
{
|
||||
"seq": 9,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "checkpoint saved at step 4500"
|
||||
},
|
||||
{
|
||||
"seq": 9,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 9"
|
||||
},
|
||||
{
|
||||
"seq": 10,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 10"
|
||||
},
|
||||
{
|
||||
"seq": 11,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 11"
|
||||
},
|
||||
{
|
||||
"seq": 12,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 12"
|
||||
},
|
||||
{
|
||||
"seq": 13,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 13"
|
||||
},
|
||||
{
|
||||
"seq": 14,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 14"
|
||||
},
|
||||
{
|
||||
"seq": 15,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 15"
|
||||
},
|
||||
{
|
||||
"seq": 16,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 16"
|
||||
},
|
||||
{
|
||||
"seq": 17,
|
||||
"level": "ERROR",
|
||||
"status": "failed",
|
||||
"msg": "OOM killed worker 3 exceeded memory limit 8GiB"
|
||||
},
|
||||
{
|
||||
"seq": 17,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 17"
|
||||
},
|
||||
{
|
||||
"seq": 18,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 18"
|
||||
},
|
||||
{
|
||||
"seq": 19,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 19"
|
||||
},
|
||||
{
|
||||
"seq": 20,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 20"
|
||||
},
|
||||
{
|
||||
"seq": 21,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 21"
|
||||
},
|
||||
{
|
||||
"seq": 22,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 22"
|
||||
},
|
||||
{
|
||||
"seq": 23,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 23"
|
||||
},
|
||||
{
|
||||
"seq": 24,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 24"
|
||||
},
|
||||
{
|
||||
"seq": 25,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 25"
|
||||
},
|
||||
{
|
||||
"seq": 26,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 26"
|
||||
},
|
||||
{
|
||||
"seq": 27,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "heartbeat ping 27"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "payment_exception",
|
||||
"question": "Which service threw an exception and where?",
|
||||
"content_type": "json_array",
|
||||
"compress": {
|
||||
"with_compaction": false,
|
||||
"max_items_after_crush": 8
|
||||
},
|
||||
"answer_evidence": [
|
||||
"NullPointerException",
|
||||
"PaymentService.charge line 88"
|
||||
],
|
||||
"supporting_facts": [],
|
||||
"content": [
|
||||
{
|
||||
"seq": 0,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 0"
|
||||
},
|
||||
{
|
||||
"seq": 1,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 1"
|
||||
},
|
||||
{
|
||||
"seq": 2,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 2"
|
||||
},
|
||||
{
|
||||
"seq": 3,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 3"
|
||||
},
|
||||
{
|
||||
"seq": 4,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 4"
|
||||
},
|
||||
{
|
||||
"seq": 5,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 5"
|
||||
},
|
||||
{
|
||||
"seq": 6,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 6"
|
||||
},
|
||||
{
|
||||
"seq": 7,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 7"
|
||||
},
|
||||
{
|
||||
"seq": 8,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 8"
|
||||
},
|
||||
{
|
||||
"seq": 9,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 9"
|
||||
},
|
||||
{
|
||||
"seq": 10,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 10"
|
||||
},
|
||||
{
|
||||
"seq": 11,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 11"
|
||||
},
|
||||
{
|
||||
"seq": 12,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 12"
|
||||
},
|
||||
{
|
||||
"seq": 13,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 13"
|
||||
},
|
||||
{
|
||||
"seq": 14,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 14"
|
||||
},
|
||||
{
|
||||
"seq": 15,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 15"
|
||||
},
|
||||
{
|
||||
"seq": 16,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 16"
|
||||
},
|
||||
{
|
||||
"seq": 17,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 17"
|
||||
},
|
||||
{
|
||||
"seq": 18,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 18"
|
||||
},
|
||||
{
|
||||
"seq": 19,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 19"
|
||||
},
|
||||
{
|
||||
"seq": 20,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 20"
|
||||
},
|
||||
{
|
||||
"seq": 21,
|
||||
"level": "ERROR",
|
||||
"status": "failed",
|
||||
"msg": "exception: NullPointerException at PaymentService.charge line 88"
|
||||
},
|
||||
{
|
||||
"seq": 21,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 21"
|
||||
},
|
||||
{
|
||||
"seq": 22,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 22"
|
||||
},
|
||||
{
|
||||
"seq": 23,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 23"
|
||||
},
|
||||
{
|
||||
"seq": 24,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 24"
|
||||
},
|
||||
{
|
||||
"seq": 25,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 25"
|
||||
},
|
||||
{
|
||||
"seq": 26,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 26"
|
||||
},
|
||||
{
|
||||
"seq": 27,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 27"
|
||||
},
|
||||
{
|
||||
"seq": 28,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 28"
|
||||
},
|
||||
{
|
||||
"seq": 29,
|
||||
"level": "INFO",
|
||||
"status": "ok",
|
||||
"msg": "GET /healthz 200 29"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "latency_anomaly",
|
||||
"question": "Which region had the latency spike, and how high?",
|
||||
"content_type": "json_array",
|
||||
"compress": {
|
||||
"with_compaction": false,
|
||||
"max_items_after_crush": 8
|
||||
},
|
||||
"answer_evidence": [
|
||||
"us-east-1",
|
||||
"99999"
|
||||
],
|
||||
"supporting_facts": [],
|
||||
"content": [
|
||||
{
|
||||
"seq": 0,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 95,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 1,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 96,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 2,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 97,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 3,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 98,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 4,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 99,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 5,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 100,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 6,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 101,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 7,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 102,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 8,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 103,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 9,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 104,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 10,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 105,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 11,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 95,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 12,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 96,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 13,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 97,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 14,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 98,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 15,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 99,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 16,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 100,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 17,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 101,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 18,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 102,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 19,
|
||||
"region": "us-east-1",
|
||||
"latency_ms": 99999,
|
||||
"status": "ok",
|
||||
"note": "latency spike us-east-1"
|
||||
},
|
||||
{
|
||||
"seq": 20,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 104,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 21,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 105,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 22,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 95,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 23,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 96,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 24,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 97,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 25,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 98,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 26,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 99,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 27,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 100,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 28,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 101,
|
||||
"status": "ok"
|
||||
},
|
||||
{
|
||||
"seq": 29,
|
||||
"region": "us-west-2",
|
||||
"latency_ms": 102,
|
||||
"status": "ok"
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "ci_test_failures",
|
||||
"question": "Which test failed and why?",
|
||||
"content_type": "json_array",
|
||||
"compress": {
|
||||
"with_compaction": false,
|
||||
"max_items_after_crush": 9
|
||||
},
|
||||
"answer_evidence": [
|
||||
"test_auth_token_refresh",
|
||||
"expected 200 got 401"
|
||||
],
|
||||
"supporting_facts": [],
|
||||
"content": [
|
||||
{
|
||||
"seq": 0,
|
||||
"test": "test_module_0",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 5
|
||||
},
|
||||
{
|
||||
"seq": 1,
|
||||
"test": "test_module_1",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 6
|
||||
},
|
||||
{
|
||||
"seq": 2,
|
||||
"test": "test_module_2",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 7
|
||||
},
|
||||
{
|
||||
"seq": 3,
|
||||
"test": "test_module_3",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 8
|
||||
},
|
||||
{
|
||||
"seq": 4,
|
||||
"test": "test_module_4",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 9
|
||||
},
|
||||
{
|
||||
"seq": 5,
|
||||
"test": "test_module_5",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 10
|
||||
},
|
||||
{
|
||||
"seq": 6,
|
||||
"test": "test_module_6",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 11
|
||||
},
|
||||
{
|
||||
"seq": 7,
|
||||
"test": "test_module_7",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 12
|
||||
},
|
||||
{
|
||||
"seq": 8,
|
||||
"test": "test_module_8",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 13
|
||||
},
|
||||
{
|
||||
"seq": 9,
|
||||
"test": "test_module_9",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 14
|
||||
},
|
||||
{
|
||||
"seq": 10,
|
||||
"test": "test_module_10",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 15
|
||||
},
|
||||
{
|
||||
"seq": 11,
|
||||
"test": "test_module_11",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 16
|
||||
},
|
||||
{
|
||||
"seq": 12,
|
||||
"test": "test_auth_token_refresh",
|
||||
"outcome": "failed",
|
||||
"duration_ms": 41,
|
||||
"error": "AssertionError: expected 200 got 401 in test_auth_token_refresh"
|
||||
},
|
||||
{
|
||||
"seq": 13,
|
||||
"test": "test_module_13",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 18
|
||||
},
|
||||
{
|
||||
"seq": 14,
|
||||
"test": "test_module_14",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 19
|
||||
},
|
||||
{
|
||||
"seq": 15,
|
||||
"test": "test_module_15",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 20
|
||||
},
|
||||
{
|
||||
"seq": 16,
|
||||
"test": "test_module_16",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 21
|
||||
},
|
||||
{
|
||||
"seq": 17,
|
||||
"test": "test_module_17",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 22
|
||||
},
|
||||
{
|
||||
"seq": 18,
|
||||
"test": "test_module_18",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 23
|
||||
},
|
||||
{
|
||||
"seq": 19,
|
||||
"test": "test_module_19",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 24
|
||||
},
|
||||
{
|
||||
"seq": 20,
|
||||
"test": "test_module_20",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 25
|
||||
},
|
||||
{
|
||||
"seq": 21,
|
||||
"test": "test_module_21",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 26
|
||||
},
|
||||
{
|
||||
"seq": 22,
|
||||
"test": "test_module_22",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 27
|
||||
},
|
||||
{
|
||||
"seq": 23,
|
||||
"test": "test_module_23",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 28
|
||||
},
|
||||
{
|
||||
"seq": 24,
|
||||
"test": "test_module_24",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 29
|
||||
},
|
||||
{
|
||||
"seq": 25,
|
||||
"test": "test_module_25",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 30
|
||||
},
|
||||
{
|
||||
"seq": 26,
|
||||
"test": "test_module_26",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 31
|
||||
},
|
||||
{
|
||||
"seq": 27,
|
||||
"test": "test_module_27",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 32
|
||||
},
|
||||
{
|
||||
"seq": 28,
|
||||
"test": "test_module_28",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 33
|
||||
},
|
||||
{
|
||||
"seq": 29,
|
||||
"test": "test_module_29",
|
||||
"outcome": "passed",
|
||||
"duration_ms": 34
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
72
tests/test_compression_fidelity_regression.py
Normal file
72
tests/test_compression_fidelity_regression.py
Normal file
|
|
@ -0,0 +1,72 @@
|
|||
"""Offline fidelity regression gate (recall-based, zero-model).
|
||||
|
||||
Compresses vendored golden tool-output fixtures through SmartCrusher's lossy
|
||||
path and asserts that the evidence a model needs to answer each case's question
|
||||
survives compression. Scoring is pure stdlib (``headroom.evals.metrics``) — no
|
||||
ML model, no network, no API keys — so this runs in the standard ``[dev]`` CI
|
||||
shard as a blocking PR check.
|
||||
|
||||
A failure here means a code change made lossy compression silently drop
|
||||
information that answers a known question. Fixtures and the committed baseline
|
||||
are generated by ``tests/fixtures/fidelity_golden/_generate.py``.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
from headroom.evals.metrics import compute_information_recall
|
||||
from headroom.transforms.smart_crusher import SmartCrusherConfig, smart_crush_tool_output
|
||||
|
||||
FIXTURE_DIR = Path(__file__).parent / "fixtures" / "fidelity_golden"
|
||||
CASES: list[dict] = json.loads((FIXTURE_DIR / "cases.json").read_text())
|
||||
BASELINE: dict = json.loads((FIXTURE_DIR / "baseline.json").read_text())
|
||||
|
||||
|
||||
def _compress(case: dict) -> tuple[str, str]:
|
||||
"""Compress a case's tool output via the lossy SmartCrusher path (no model)."""
|
||||
original = json.dumps(case["content"])
|
||||
cfg = SmartCrusherConfig(max_items_after_crush=case["compress"]["max_items_after_crush"])
|
||||
crushed, _modified, _info = smart_crush_tool_output(
|
||||
original, cfg, with_compaction=case["compress"]["with_compaction"]
|
||||
)
|
||||
return original, crushed
|
||||
|
||||
|
||||
@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])
|
||||
def test_critical_evidence_survives_compression(case: dict) -> None:
|
||||
"""Every ``answer_evidence`` string MUST survive lossy compression (recall == 1.0).
|
||||
|
||||
Critical evidence lives in error/anomaly rows, which SmartCrusher formally
|
||||
guarantees to retain (see ``tests/test_quality_retention.py``).
|
||||
"""
|
||||
original, crushed = _compress(case)
|
||||
result = compute_information_recall(original, crushed, case["answer_evidence"])
|
||||
assert result["recall"] == 1.0, (
|
||||
f"FIDELITY REGRESSION in '{case['id']}': compression dropped evidence "
|
||||
f"needed to answer {case['question']!r}. Lost: {result['facts_lost']}"
|
||||
)
|
||||
|
||||
|
||||
def test_aggregate_recall_not_regressed() -> None:
|
||||
"""Mean recall over all evidence must not fall below the committed baseline.
|
||||
|
||||
Catches softer regressions (e.g. relevant-but-non-critical context being
|
||||
dropped more aggressively) that the per-case critical gate would not.
|
||||
"""
|
||||
recalls = []
|
||||
for case in CASES:
|
||||
original, crushed = _compress(case)
|
||||
probes = case["answer_evidence"] + case["supporting_facts"]
|
||||
recalls.append(compute_information_recall(original, crushed, probes)["recall"])
|
||||
|
||||
mean_recall = sum(recalls) / len(recalls)
|
||||
floor = BASELINE["aggregate_recall"] - BASELINE["tolerance"]
|
||||
assert mean_recall >= floor, (
|
||||
f"FIDELITY REGRESSION: mean recall {mean_recall:.4f} fell below baseline "
|
||||
f"floor {floor:.4f} (baseline {BASELINE['aggregate_recall']} - tolerance "
|
||||
f"{BASELINE['tolerance']}). If this drop is intended, regenerate the baseline."
|
||||
)
|
||||
Loading…
Add table
Add a link
Reference in a new issue