Commit graph

3 commits

Author SHA1 Message Date
Ashish
23d73ae070
test(evals): add offline fidelity regression gate (recall-based, zero-model) (#1187)
## Description

Headroom's lossy compression drops rows/lines using statistical
heuristics but **never checks that meaning survived** — a dropped `OOM
killed worker 3` line can silently flip a model's answer with no signal
that compression caused it. The repo already ships a quality-metric
toolkit (`headroom/evals/metrics.py`) and a `weekly-suite` eval job, but
neither gates the compression path on a PR.

This adds a **per-PR fidelity regression gate**: compress vendored
golden tool-outputs through SmartCrusher's lossy path and assert the
evidence that answers each case's question survives. It is the first of
a planned trio (this is the "offline gate" half of the fidelity work);
query-aware retention and a hard token-budget API are documented
follow-ups.

Closes #

## Type of Change

- [ ] Bug fix (non-breaking change that fixes an issue)
- [x] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- **Blocking gate** (`tests/test_compression_fidelity_regression.py`):
compresses each golden case via `smart_crush_tool_output(...,
with_compaction=False)` and scores with `compute_information_recall`.
Two assertions:
- **Per-case critical recall == 1.0** — every `answer_evidence` string
(placed in error/anomaly rows, the documented SmartCrusher retention
guarantee) must survive.
- **Aggregate recall ≥ committed baseline** (`baseline.json`, tol 0.02)
— catches softer regressions.
- **Vendored fixtures** (`tests/fixtures/fidelity_golden/`):
deterministic `_generate.py` emits `cases.json` (4 cases: OOM crash,
payment exception, latency anomaly, CI failure) + `baseline.json`.
- **Non-blocking weekly report** (`.github/workflows/eval.yml`): one
step in the existing `weekly-suite` job (schedule/manual only) reuses
the existing `evaluate_information_retention` runner for a recall report
on the production routing path.
- **Pure reuse**: scoring (`evals/metrics.py`), compressor
(`smart_crush_tool_output`), and the weekly runner
(`evaluate_information_retention`) all already existed.

### Design notes

- **Zero new CI setup.** The blocking gate runs in the existing `[dev]`
test shard — no new workflow, no new deps, **no model, no network, no
secrets** (verified under `HF_HUB_OFFLINE=1`). It deliberately uses
small hand-made structured fixtures rather than the repo's HuggingFace
dataset loaders, which would require a network download + ModernBERT and
don't belong in a fast PR gate.
- **Scope:** structured JSON tool-output (the dominant, deterministic,
model-free case). Real-dataset (HotpotQA/BFCL) recall — which needs
`[all]` + a local model — is a **documented follow-up PR**, and the
`weekly-suite` job (which genuinely runs every Monday) is its natural
home.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ HF_HUB_OFFLINE=1 python -m pytest tests/test_compression_fidelity_regression.py -v
tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[logs_oom] PASSED
tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[payment_exception] PASSED
tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[latency_anomaly] PASSED
tests/test_compression_fidelity_regression.py::test_critical_evidence_survives_compression[ci_test_failures] PASSED
tests/test_compression_fidelity_regression.py::test_aggregate_recall_not_regressed PASSED
============================== 5 passed in 0.18s ===============================
```

## Real Behavior Proof

- **Environment:** local checkout of `feat/fidelity-regression-gate`,
`pip install -e ".[dev]"`, `HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1`
(proves no model/network).
- **Exact command / steps:** `HF_HUB_OFFLINE=1 python -m pytest
tests/test_compression_fidelity_regression.py -q` → `5 passed in 0.14s`.
- **Negative control (proves the gate has teeth):** compressing
`logs_oom` and probing for a benign row that compression legitimately
drops returns `recall = 0.00, lost = ['heartbeat ping 25']` — i.e. the
gate fires when critical evidence is dropped, so it is not trivially
green.
- **Weekly (non-blocking) step verified locally:**
  ```text
  Information retention: 50/50 cases >=0.9 recall, avg compression 65.7%
  ```
- **Not tested:** real-dataset (HotpotQA/BFCL) recall and
prose/ModernBERT compression — intentionally deferred to a follow-up PR
targeting the weekly job.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable

## Additional Notes

- CHANGELOG/version intentionally untouched: repo uses
**release-please**.
- **Follow-up PR (planned):** wire the real HotpotQA/BFCL loaders
(`headroom/evals/datasets.py`) into the `weekly-suite` job for genuine
benchmark-scale recall coverage (model-allowed, non-blocking). Further
follow-ups from the same design: a live per-request fidelity guardrail,
query-aware lossy retention, and a hard `target_tokens` budget API.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 22:53:59 -05:00
chopratejas
c62d45eea8 fix(memory): expose memory IDs in auto-tail + memory_list tool + ID-usage guidance
Pre-this-PR the auto-injected memory block rendered rows as `1. <content>`
with no addressable handle. To UPDATE or DELETE a row the model first had
to call memory_search to discover its ID — two round trips, against the
model-as-judge architecture.

This PR adds three tightly-coupled affordances so the model can act on
memory directly:

1. Auto-tail rows now carry the memory ID:
     `1. [mem_alpha_001] User prefers Python`
   The bracketed token is the canonical ID — same identifier accepted by
   memory_update and memory_delete.

2. New `memory_list` tool — chronological browse (vs `memory_search`'s
   semantic lookup). Returns recent memories with their IDs. Backend
   dispatches to `Backend.list_memories` if available, else falls back
   to an empty-query `search_memories`. Caps at 100 entries.

3. ID-usage guidance text appended to the auto-tail block. Tells the
   model that bracketed IDs can go straight to memory_update /
   memory_delete with no intervening search. The guidance lives in the
   user-message tail (never system) — preserves cache-prefix byte
   stability (invariant I2).

`memory_update` and `memory_delete` tool descriptions also point at the
[id] block as a valid ID source — keeps tool docs consistent with the
new affordance.

Verification:
- 10/10 tests pass in tests/test_memory_auto_tail.py (incl. 2 new
  guidance tests + 2 new ID-format tests)
- 31/31 tests pass in tests/test_memory_handler_native_ops.py (incl. 4
  new memory_list dispatch tests + existing assertions updated for the
  [id] format change)
- Golden fixtures regenerated for the tool-description copy changes
  (tests/fixtures/memory_tool_definitions/{anthropic,openai}.json)
- Live end-to-end test against real Anthropic API
  (tests/test_proxy_memory_integration.py::TestMemoryIdAutoTailAndUpdate):
  seeded memory → auto-tail → Claude → memory_update with exact ID.
  PASSED.
2026-05-19 22:10:42 -05:00
chopratejas
8dcd474aca fix: A7 — memory tool injection session-sticky for both Anthropic and OpenAI
Closes the second half of P0-6: once memory injects memory_save / memory_search
into body["tools"] for a session, every subsequent turn injects the byte-equal
same definitions — even if memory is disabled mid-session. Toggling tool list
mid-session busts Anthropic prefix cache per guide §6.3 #2.

Adds in headroom/proxy/helpers.py:

  * SessionToolTracker — bounded LRU keyed by (provider, session_id) storing
    GOLDEN tool-definition bytes from the first injection. Tracker is
    provider-aware so the same session_id under Anthropic and OpenAI keeps
    independent state. Reentrant lock for concurrent access; LRU eviction at
    HEADROOM_TOOL_TRACKER_MAX_SESSIONS (default 1000).
  * apply_session_sticky_memory_tools — single coordination point with three
    paths: first-time inject (record golden bytes), sticky replay (always
    inject golden bytes regardless of inject_this_turn), and skip. Honors
    HEADROOM_TOOL_INJECTION_STICKY=disabled as a loud operator opt-in for
    rollback (NOT a fallback).
  * serialize_tool_definition_canonical — deterministic byte serialization
    via the same separators=(",",":")/ensure_ascii=False rules as
    serialize_body_canonical.
  * log_tool_injection_decision — structured per-decision log line; never
    logs the tool definition contents.

Wires the helper into all four memory tool injection sites:
  * handlers/anthropic.py — /v1/messages
  * handlers/openai.py — /v1/chat/completions
  * handlers/openai.py — /v1/responses
  * handlers/openai.py — Codex WS path

memory_handler.MemoryHandler gains compute_memory_tool_definitions(provider) —
a pure builder that returns the tool definitions without mutating a tools
list, so the proxy can route through the sticky tracker. The legacy
inject_tools(...) is preserved for callers without a session_id.

Tests: tests/test_memory_tool_session_sticky.py — 29 unit + integration
cases covering: turn-1→turn-2 byte-equality (Anthropic + OpenAI), sticky
replay after memory disabled, golden-fixture pin, LRU eviction, provider
isolation under shared session_id, thread-safe concurrent access, env-var
contract, disabled-mode passthrough, dedupe with client tools.

Golden fixtures pin canonical bytes:
  * tests/fixtures/memory_tool_definitions/anthropic.json
  * tests/fixtures/memory_tool_definitions/openai.json

No regex. No hardcodes (env-configurable: HEADROOM_TOOL_INJECTION_STICKY,
HEADROOM_TOOL_TRACKER_MAX_SESSIONS). No silent fallbacks. Per-decision
structured logging. Realignment build constraints satisfied.
2026-05-02 10:11:27 -07:00