From 46d4378cf7796836f99933c8f614f84d0a08bf4a Mon Sep 17 00:00:00 2001 From: Ashish Date: Wed, 15 Jul 2026 14:40:55 -0700 Subject: [PATCH] feat(evals): weekly HotpotQA answer-recall report on the prose path (#1188) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ## Description Follow-up to **#1187** (the offline fidelity gate). That gate is hermetic and **structured-only** (JSON tool outputs via Rust compressors) so it can block every PR with zero setup. This PR adds the genuinely-uncovered piece: **prose answer-recall on a real dataset (HotpotQA)** in the **model-allowed weekly job**, where compression routes through Kompress (ModernBERT). > **Stacked on #1187.** Until that merges, this PR's diff shows its commit too; it reduces to just `c71cc0cb` once #1187 lands. Please review/merge #1187 first. Closes # ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **`CompressionOnlyRunner.evaluate_dataset_recall(suite)`**: for each QA case, compress the supporting `context` via the production routing path (`ContentRouter`) and check the `ground_truth` answer survives (`compute_information_recall`). Counts only **probeable** cases — answer literally present in the context and non-trivial (skips `yes/no`, too-short) — so the aggregate is meaningful rather than inflated by un-measurable cases. - **`.github/workflows/eval.yml`**: a non-blocking step in the existing `weekly-suite` job (schedule/manual only) drives it with `load_hotpotqa(n=50)`. Defensive: a dataset download or model failure emits `::warning::` and `|| true`, never failing the job. - **Hermetic unit test** (`tests/test_dataset_recall_runner.py`): exercises the method with synthetic JSON-array contexts (SmartCrusher / Rust — no model, no network), so it runs in the standard `[dev]` shard. ### Scope notes - **Prose path only.** BFCL / tool-schema integrity is already covered by the existing `evaluate_tool_schema_compaction` eval (which runs in the PR smoke-test), so this targets the previously-uncovered prose recall path. NQ is an easy further extension using the same method + `load_natural_questions`. - **Why weekly, not per-PR.** Real datasets need a network download + the ModernBERT model. The `weekly-suite` job already installs `[all]` and genuinely runs every Monday (verified: 5 consecutive successful scheduled runs), so it's the correct home — keeping PR CI fast and hermetic. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ HF_HUB_OFFLINE=1 python -m pytest tests/test_dataset_recall_runner.py -q .. [100%] 2 passed in 0.20s ``` ## Real Behavior Proof - Environment: local checkout of `feat/weekly-dataset-recall`, `pip install -e ".[dev]"`, `HF_HUB_OFFLINE=1` (proves the unit tests need no model/network) - Exact command / steps: `HF_HUB_OFFLINE=1 python -m pytest tests/test_dataset_recall_runner.py -q` -> `6 passed in 0.36s`; coverage JSON confirms the runner's per-case exception handler and both `warm_kompress_model` outcomes are exercised - Observed result: with a synthetic suite of 3 cases (one probeable answer in an error row, one trivial `yes`, one absent answer), `evaluate_dataset_recall` counts only the 1 probeable case (`passed=1`, `accuracy_rate=1.0`, `benchmark="dataset_recall:synthetic"`); a monkeypatched compressor crash records the error and counts the case failed instead of aborting; the new weekly-suite YAML step parses via `yaml.safe_load` and sits under the `schedule || workflow_dispatch` guard - Not tested: the live HotpotQA download + ModernBERT compression -- exercised only by the weekly job (or `workflow_dispatch`), by design ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable ## Additional Notes - CHANGELOG/version intentionally untouched: repo uses **release-please**. - The weekly job can be triggered on demand via **workflow_dispatch** to see the HotpotQA recall numbers without waiting for Monday. --------- Co-authored-by: Claude Opus 4.8 Co-authored-by: JD Davis --- .github/workflows/eval.yml | 26 +++++ headroom/evals/runners/compression_only.py | 90 +++++++++++++++- headroom/transforms/kompress_compressor.py | 27 +++++ tests/test_dataset_recall_runner.py | 119 +++++++++++++++++++++ 4 files changed, 261 insertions(+), 1 deletion(-) create mode 100644 tests/test_dataset_recall_runner.py diff --git a/.github/workflows/eval.yml b/.github/workflows/eval.yml index 11dcc124b..92bd0c159 100644 --- a/.github/workflows/eval.yml +++ b/.github/workflows/eval.yml @@ -144,6 +144,32 @@ jobs: print(f'::warning title=Fidelity recall::{result.failed_cases} case(s) fell below 0.9 recall: {result.errors[:3]}') " + # Real-dataset recall on the prose path (HotpotQA): does the ground-truth + # answer survive compressing the supporting context? Uses the production + # routing path, so prose flows through Kompress (ModernBERT) — allowed here + # because the weekly job installs [all]. Non-blocking and defensive: a + # dataset download or model failure warns rather than fails the job. + - name: Dataset recall report — HotpotQA (model-allowed, non-blocking) + run: | + python -c " + try: + from headroom.transforms.kompress_compressor import warm_kompress_model + from headroom.evals.datasets import load_hotpotqa + from headroom.evals.runners.compression_only import CompressionOnlyRunner + # Block until the Kompress model is loaded; otherwise prose passes + # through uncompressed and the recall number is meaningless. + warmed = warm_kompress_model() + suite = load_hotpotqa(n=50) + result = CompressionOnlyRunner().evaluate_dataset_recall(suite) + print(f'HotpotQA answer recall: {result.passed_cases}/{result.total_cases} probeable cases >=0.9, avg compression {result.avg_compression_ratio:.1%} (model_warmed={warmed})') + if result.avg_compression_ratio < 0.01: + print('::warning title=Dataset recall::compression did not engage (~0%); recall is not a meaningful fidelity signal — check Kompress model availability') + elif result.failed_cases: + print(f'::warning title=Dataset recall::{result.failed_cases} HotpotQA case(s) lost the answer under compression') + except Exception as e: + print(f'::warning title=Dataset recall::skipped (dataset/model unavailable): {e}') + " || true + - name: Upload results if: always() uses: actions/upload-artifact@v7 diff --git a/headroom/evals/runners/compression_only.py b/headroom/evals/runners/compression_only.py index e3ad17888..51a8a8f3d 100644 --- a/headroom/evals/runners/compression_only.py +++ b/headroom/evals/runners/compression_only.py @@ -14,7 +14,10 @@ import json import logging import time from dataclasses import dataclass, field -from typing import Any +from typing import TYPE_CHECKING, Any + +if TYPE_CHECKING: + from headroom.evals.core import EvalSuite logger = logging.getLogger(__name__) @@ -242,6 +245,91 @@ class CompressionOnlyRunner: errors=errors, ) + def evaluate_dataset_recall( + self, + suite: EvalSuite, + recall_threshold: float = 0.9, + min_answer_chars: int = 4, + ) -> CompressionOnlyResult: + """Compress each QA case's context and check its answer survives. + + The probe is the case's ``ground_truth`` answer. A case only counts when + the answer literally appears in the original context (otherwise survival + is not measurable); trivial answers (too short, or yes/no) are skipped. + Compression uses the production routing path, so prose flows through + Kompress (ModernBERT) — intended for the model-allowed weekly job, not + the hermetic per-PR gate. + + Args: + suite: An EvalSuite of QA cases (e.g. from ``load_hotpotqa``). + recall_threshold: Minimum answer recall for a case to pass. + min_answer_chars: Answers shorter than this are skipped as un-probeable. + """ + from headroom.evals.metrics import compute_information_recall + from headroom.transforms.content_router import ContentRouter + + trivial = {"yes", "no", "true", "false"} + start_time = time.time() + router = ContentRouter() + passed = 0 + failed = 0 + total_original = 0 + total_compressed = 0 + details: list[dict[str, Any]] = [] + errors: list[str] = [] + + for case in suite.cases: + answer = (case.ground_truth or "").strip() + if len(answer) < min_answer_chars or answer.lower() in trivial: + continue # un-probeable: survival of this answer carries no signal + if answer.lower() not in case.context.lower(): + continue # answer not literally in context; nothing to measure + + original_tokens = self._estimate_tokens(case.context) + try: + compressed = router.compress(case.context).compressed + compressed_tokens = self._estimate_tokens(compressed) + recall = compute_information_recall(case.context, compressed, [answer])["recall"] + is_pass = recall >= recall_threshold + + total_original += original_tokens + total_compressed += compressed_tokens + passed += is_pass + failed += not is_pass + details.append( + { + "id": case.id, + "passed": is_pass, + "recall": recall, + "answer": answer, + "compression_ratio": 1 - (compressed_tokens / original_tokens) + if original_tokens > 0 + else 0, + } + ) + except Exception as e: + failed += 1 + errors.append(f"Dataset recall error for {case.id}: {e}") + details.append({"id": case.id, "passed": False, "error": str(e)}) + + total_cases = passed + failed + ratios = [d["compression_ratio"] for d in details if "compression_ratio" in d] + + return CompressionOnlyResult( + benchmark=f"dataset_recall:{suite.name}", + total_cases=total_cases, + passed_cases=passed, + failed_cases=failed, + accuracy_rate=passed / total_cases if total_cases > 0 else 0.0, + avg_compression_ratio=sum(ratios) / len(ratios) if ratios else 0.0, + total_original_tokens=total_original, + total_compressed_tokens=total_compressed, + total_tokens_saved=total_original - total_compressed, + duration_seconds=time.time() - start_time, + details=details, + errors=errors, + ) + def generate_ccr_test_cases(self, n: int = 50) -> list[dict[str, Any]]: """Generate synthetic test cases for CCR needle-retention testing. diff --git a/headroom/transforms/kompress_compressor.py b/headroom/transforms/kompress_compressor.py index fc9b6d212..596ec1b44 100644 --- a/headroom/transforms/kompress_compressor.py +++ b/headroom/transforms/kompress_compressor.py @@ -946,6 +946,33 @@ def ensure_background_download(model_id: str = HF_MODEL_ID, device: str = "auto" thread.start() +def warm_kompress_model( + model_id: str = HF_MODEL_ID, + device: str = "cpu", + *, + allow_download: bool = True, +) -> bool: + """Synchronously load the Kompress model, blocking until it is ready. + + Unlike :func:`ensure_background_download` (which loads in a daemon thread and + lets early requests pass through uncompressed while the download runs), this + blocks the caller so the *next* compression uses the model rather than + passing through. Intended for batch/eval contexts that must measure real + compression, not passthrough. + + Returns ``True`` if the model is loaded and ready, ``False`` if Kompress is + unavailable or the load failed. + """ + if not is_kompress_available(): + return False + try: + _load_kompress(model_id, device, allow_download=allow_download) + return model_id in _kompress_cache + except Exception as exc: # pragma: no cover - network/model load failure + logger.warning("Kompress: synchronous warm failed for %s: %s", model_id, exc) + return False + + # ── Compressor ──────────────────────────────────────────────────────── diff --git a/tests/test_dataset_recall_runner.py b/tests/test_dataset_recall_runner.py new file mode 100644 index 000000000..3b205b762 --- /dev/null +++ b/tests/test_dataset_recall_runner.py @@ -0,0 +1,119 @@ +"""Hermetic unit tests for CompressionOnlyRunner.evaluate_dataset_recall. + +Exercises the dataset-recall plumbing with synthetic JSON-array contexts (which +route through SmartCrusher / Rust — no model, no network) so it runs in the +standard [dev] shard. The weekly job drives the same method with real prose +datasets (HotpotQA), which is intentionally not exercised here. +""" + +from __future__ import annotations + +import json + +from headroom.evals.core import EvalCase, EvalSuite +from headroom.evals.runners.compression_only import CompressionOnlyRunner + + +def _array_context_with(answer: str) -> str: + """A JSON-array tool output whose error row embeds ``answer`` (a kept row).""" + rows = [{"seq": i, "level": "INFO", "status": "ok", "msg": f"heartbeat {i}"} for i in range(30)] + rows[14] = {"seq": 14, "level": "ERROR", "status": "failed", "msg": answer} + return json.dumps(rows) + + +def _suite() -> EvalSuite: + answer = "PaymentService NullPointerException at charge line 88" + return EvalSuite( + name="synthetic", + cases=[ + # Probeable: answer is in an error row -> retained -> recall 1.0. + EvalCase( + id="probeable", + context=_array_context_with(answer), + query="what failed?", + ground_truth=answer, + ), + # Skipped: trivial yes/no answer. + EvalCase( + id="trivial", + context=_array_context_with(answer), + query="did it fail?", + ground_truth="yes", + ), + # Skipped: answer not present in the context at all. + EvalCase( + id="absent", + context=_array_context_with(answer), + query="?", + ground_truth="totally-absent-token-xyz", + ), + ], + ) + + +def test_dataset_recall_counts_only_probeable_cases() -> None: + result = CompressionOnlyRunner().evaluate_dataset_recall(_suite()) + # Only the "probeable" case is measurable; trivial + absent are skipped. + assert result.total_cases == 1 + assert result.passed_cases == 1 + assert result.accuracy_rate == 1.0 + assert result.benchmark == "dataset_recall:synthetic" + + +def test_dataset_recall_empty_suite_is_safe() -> None: + result = CompressionOnlyRunner().evaluate_dataset_recall(EvalSuite(name="empty", cases=[])) + assert result.total_cases == 0 + assert result.accuracy_rate == 0.0 + assert result.errors == [] + + +def test_dataset_recall_records_compression_errors(monkeypatch) -> None: + # A compressor crash on one case must not abort the run: the case counts + # as failed, the error is recorded, and the detail row carries it. + from headroom.transforms.content_router import ContentRouter + + def _boom(self, content, context="", question=None, bias=1.0): + raise RuntimeError("router exploded") + + monkeypatch.setattr(ContentRouter, "compress", _boom) + result = CompressionOnlyRunner().evaluate_dataset_recall(_suite()) + assert result.total_cases == 1 + assert result.failed_cases == 1 + assert result.passed_cases == 0 + assert result.errors and "router exploded" in result.errors[0] + assert result.details[0]["passed"] is False + assert "router exploded" in result.details[0]["error"] + + +def test_warm_kompress_model_returns_false_when_unavailable(monkeypatch) -> None: + # Guard path: no Kompress backend -> no download attempt, returns False. + import headroom.transforms.kompress_compressor as kc + + monkeypatch.setattr(kc, "is_kompress_available", lambda: False) + assert kc.warm_kompress_model() is False + + +def test_warm_kompress_model_true_when_load_populates_cache(monkeypatch) -> None: + # Success path: the synchronous load lands the model in the cache. + import headroom.transforms.kompress_compressor as kc + + cache: dict[str, object] = {} + monkeypatch.setattr(kc, "_kompress_cache", cache) + monkeypatch.setattr(kc, "is_kompress_available", lambda: True) + monkeypatch.setattr( + kc, + "_load_kompress", + lambda model_id, device, allow_download: cache.setdefault(model_id, object()), + ) + assert kc.warm_kompress_model("test-model") is True + + +def test_warm_kompress_model_false_when_load_leaves_cache_empty(monkeypatch) -> None: + # The loader returned without raising but the model never landed in the + # cache (e.g. download disallowed and not cached locally). + import headroom.transforms.kompress_compressor as kc + + monkeypatch.setattr(kc, "_kompress_cache", {}) + monkeypatch.setattr(kc, "is_kompress_available", lambda: True) + monkeypatch.setattr(kc, "_load_kompress", lambda model_id, device, allow_download: None) + assert kc.warm_kompress_model("test-model", allow_download=False) is False