headroom/tests/test_agent_savings.py
Tejas Chopra 250ede2f7f
fix(reporting): show net vs gross savings, real skip thresholds, and the effective profile (#3123)
## Description

Six reporting/config defects found while investigating a user reporting
~1% savings on Claude Code. **None of these changes how much Headroom
compresses** — all of them change whether an operator can tell what it
did. Every one was found by reading that user's own 227,777 lines of
proxy logs against the code.

## Changes Made

- **`perf/analyzer`: parse and render `tok_inflated`.** Every PERF line
carried it; nothing downstream read it. The report could print
`321,239,562 -> 313,274,727` directly above `8,455,763 saved` — two
figures that differ by exactly the 490,928 tokens of inflation it
omitted.
- **`content_router`: report the real skip thresholds.** The routing
summary hardcoded `skipped (<50 words)` regardless of what was in force.
Wrong number (the message gate is `min_tokens`, 10–250 by profile),
wrong unit (tokens and characters, never words), and it merged two
different gates under one label.
- **`perf/analyzer`: disclose that Transform Effectiveness is partial.**
It is built only from `pipeline.py`'s `Transform NAME:` lines.
`compression_units.py` / `compression_batches.py` contain zero logging
calls, so the table read `content_router: 189,783 saved` against a PERF
total 44x larger. Reports the divergence rather than a coverage ratio —
the two are different populations and neither contains the other (those
lines carry no request_id, fire per stage, and are emitted before the
forwarder decides).
- **`perf/analyzer`: disclose the routing denominator.** Percentages
were taken over 4 of the router's 17 outcome buckets, silently dropping
buckets larger than several it displayed.
- **`savings_tracker`: stop dropping tool-schema dollars.**
`estimate_request_savings_usd` prices four buckets; `record_request`
read three. `tool_schema` was computed and discarded, so a quarter of
the token headline never reached "Cost saved". The two inputs are
disjoint (verified at the call site), so this is additive, not
double-counting.
- **`agent_savings`: an unknown profile no longer degrades to
`balanced`.** `balanced` is a different product posture from the default
`coding`: cache→token mode, dedup off, tool-search off, user messages
uncompressed, message floor 25x higher, block floor 20x higher. A typo
in `HEADROOM_SAVINGS_PROFILE` silently reconfigured the whole proxy. Now
degrades to `DEFAULT_PROFILE` and names the resolved profile in the
warning.
- **`agent_savings`: give `min_chars_for_block` a config-object path.**
Every other router pipeline kwarg travels on the config object; this one
alone was env-only, so an unseeded proxy applied every sibling `coding`
knob while this floor stayed at 500 instead of 25.
- **`server`: log the resolved compression posture at startup**, reading
cross-turn dedup off the constructed router rather than the environment
(the router resolves it as `config OR env`, so reading env alone would
be a guess).

## Testing

- [x] Unit tests pass, [x] ruff, [x] mypy, [x] new tests added

```text
uv run pytest tests/ -k "content_router or agent_savings or perf or analyzer or savings or proxy_server or cli_perf or prometheus"
620 passed, 25 skipped
uv run mypy headroom  # Success
uv run ruff check . && ruff format --check .  # clean
```

## Real behavior proof

- **Setup:** macOS arm64, Python 3.12, this branch. Input: 60 MB /
227,777 lines of real proxy logs from the reporting user (6 rotated
files, 2,792 PERF lines, 2026-08-17 → 2026-08-19).
- **Steps:** pointed `headroom.perf.analyzer.LOG_DIR` at that directory
and rendered the report before and after the patch.
- **After-fix output (real data, unmodified):**

```text
Requests:     2792
Tokens:       321,288,161 -> 313,323,326 (2.6% messages)
Tokens saved: 11,158,901 (3.4% reduction)
  · inflated      490,928 (net message reduction 7,964,835)
  · messages       8,455,763
  · tool schemas   2,703,138
  ! stage-level total 190,641 != PERF message total 8,455,763 — this table sees only
    engines that emit a Transform line, counts per stage, and does not check whether
    the mutation shipped
  Skipped:     44641 (77%) — below size floor
  (shares are of these 4 buckets only, n=58319; see `[router] route_counts=` for the
   full outcome space)
```

The arithmetic now closes on the page: `8,455,763 - 490,928 =
7,964,835`, matching the token delta exactly. Before the patch none of
the three annotated lines existed and the `Skipped` line claimed `<50
words`.

- **Profile resolution verified by execution**, not inspection —
subprocesses with controlled env:

```text
vanilla (nothing set)          mode=cache dedupe=1 tool_search=1 min_tokens=10  min_chars=25
HEADROOM_SAVINGS_PROFILE=coding  mode=cache dedupe=1 tool_search=1 min_tokens=10  min_chars=25
unknown profile name (before)  mode=token dedupe=0 tool_search=0 min_tokens=250 min_chars=500
unknown profile name (after)   -> resolves to `coding`, warning names it
coding, seeding never runs     min_chars=25 (was 500 before this patch)
```

- **Not tested:** live paid Anthropic traffic. These are
reporting/config surfaces; the wire path is untouched by this PR.

## Review readiness

- [x] Self-reviewed. Three overclaims in my own first draft were
corrected before this PR: a false subset claim in the Transform
Effectiveness note, a comment asserting `min_chars_for_block` was the
*only* env-only field (it is the only env-only *router pipeline kwarg*;
`cross_turn_dedup`, `tool_search`, `protect_reads`, `code_aware`,
`effort_router`, `lossless` remain env-only via a different mechanism
and are **not** fixed here), and a money-path expression that relied on
`a + b if c else d` grouping.

## Known remaining (deliberately out of scope)

- `Requests: N` still overcounts: the Codex WS forwarder reuses one
`request_id` across every turn (one observed 156x), plus ~18 duplicate
PERF emissions.
- `compression_units.py` / `compression_batches.py` remain unlogged —
this PR *discloses* the blind spot rather than closing it.
- The headline stays **gross**. True net is `11,158,901 - 490,928 =
10,667,973` (3.3%, not 3.4%). Making net the headline lowers every
user's reported savings ~4.4%; that is a product call, not mine, so the
inflation is surfaced beside it instead.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Tejas Chopra <tejas@Tejass-MacBook-Pro.local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 00:16:37 -07:00

842 lines
29 KiB
Python

from __future__ import annotations
import json
import logging
from importlib import import_module
from types import SimpleNamespace
import pytest
from click.testing import CliRunner
from headroom.agent_savings import (
AGENT_90_PROFILE,
DEFAULT_PROFILE,
apply_agent_savings_env_defaults,
apply_agent_savings_profile,
get_agent_savings_profile,
proxy_pipeline_kwargs,
with_target_savings,
)
from headroom.cli import wrap as wrap_module
from headroom.cli.main import main
from headroom.compress import CompressConfig, compress
from headroom.proxy.models import ProxyConfig
from headroom.transforms.compression_units import (
CompressionUnit,
compress_unit_with_router,
)
from headroom.transforms.content_router import (
CompressionStrategy,
ContentRouter,
ContentRouterConfig,
RouterCompressionResult,
)
compress_module = import_module("headroom.compress")
def test_agent_90_profile_sets_accuracy_preserving_compress_config() -> None:
cfg = CompressConfig()
apply_agent_savings_profile(cfg, AGENT_90_PROFILE)
assert cfg.compress_user_messages is True
assert cfg.compress_system_messages is True
assert cfg.protect_recent == 2
assert cfg.protect_analysis_context is True
assert cfg.target_ratio == 0.10
assert cfg.min_tokens_to_compress == 120
def test_agent_90_profile_exports_cross_agent_proxy_env() -> None:
profile = get_agent_savings_profile(AGENT_90_PROFILE)
env = profile.proxy_env()
assert env["HEADROOM_MODE"] == "token"
assert env["HEADROOM_SAVINGS_PROFILE"] == "agent-90"
assert env["HEADROOM_SAVINGS_TARGET"] == "0.90"
assert env["HEADROOM_TARGET_RATIO"] == "0.10"
assert env["HEADROOM_COMPRESS_USER_MESSAGES"] == "1"
assert env["HEADROOM_COMPRESS_SYSTEM_MESSAGES"] == "1"
assert env["HEADROOM_MAX_ITEMS"] == "8"
assert env["HEADROOM_SMART_CRUSHER_COMPACTION"] == "0"
assert env["HEADROOM_FORCE_KOMPRESS"] == "1"
assert env["HEADROOM_ACCURACY_GUARD"] == "strict"
def test_coding_persona_compresses_recent_delta_and_stays_visible() -> None:
profile = get_agent_savings_profile("coding")
env = profile.proxy_env()
assert env["HEADROOM_SAVINGS_PROFILE"] == "coding"
assert env["HEADROOM_MODE"] == "cache" # delta-only compression at ~0 prefix-cache busts
assert env["HEADROOM_PROTECT_RECENT"] == "0" # reads guarded by type, not position
assert env["HEADROOM_MIN_TOKENS"] == "10" # low → even modest deltas are eligible
# Cache mode compresses the newest observation delta → compress_user must be ON.
assert env["HEADROOM_COMPRESS_USER_MESSAGES"] == "1"
assert env["HEADROOM_COMPRESS_SYSTEM_MESSAGES"] == "0" # system prompt is the hottest cache
assert env["HEADROOM_ACCURACY_GUARD"] == "strict"
assert "HEADROOM_TARGET_RATIO" not in env # unset → Kompress / ambient default decides
# Coding posture toggles seeded through the profile.
assert env["HEADROOM_TOOL_SEARCH"] == "1"
assert env["HEADROOM_DEDUPE"] == "1"
assert env["HEADROOM_LOSSLESS_THEN_LOSSY"] == "1"
assert env["HEADROOM_PROTECT_READS"] == "1"
assert env["HEADROOM_CODE_AWARE_ENABLED"] == "1"
assert env["HEADROOM_EFFORT_ROUTER"] == "0"
assert env["HEADROOM_LOSSLESS"] == "0" # lossy enabled (CCR keeps it recoverable)
assert env["HEADROOM_MIN_CHARS_FOR_BLOCK"] == "25"
def test_coding_profile_couples_zero_protect_recent_with_type_read_guard() -> None:
"""Fidelity invariant behind #2145's protect_recent 2->0.
Dropping positional protection (protect_recent=0) is only safe because the
code working set stays byte-exact via the TYPE-based read guard
(protect_reads=True), and any *other* recent delta that does get compressed
stays losslessly recoverable via CCR (lossless=0 means lossy-with-CCR, not
silent loss) while the frozen prefix is left untouched (cache mode). This is
the honest boundary: recent file reads are verbatim, recent non-read deltas
are recoverable — not "nothing is ever touched". If a future edit drops the
read guard while keeping protect_recent=0, recent reads would silently
degrade, so this test fails closed on that pairing.
"""
profile = get_agent_savings_profile("coding")
# Positional protection is off ...
assert profile.protect_recent == 0
# ... so the byte-exact guarantee for reads MUST come from the type guard.
assert profile.protect_reads is True
env = profile.proxy_env()
assert env["HEADROOM_PROTECT_RECENT"] == "0"
assert env["HEADROOM_PROTECT_READS"] == "1"
# Lossy compression is on, but CCR keeps compressed deltas recoverable, and
# cache mode compresses only the newest delta (frozen prefix stays byte-stable).
assert env["HEADROOM_LOSSLESS"] == "0"
assert env["HEADROOM_MODE"] == "cache"
def test_general_persona_has_no_positional_code_protection() -> None:
profile = get_agent_savings_profile("general")
env = profile.proxy_env()
assert env["HEADROOM_PROTECT_RECENT"] == "0"
assert env["HEADROOM_MIN_TOKENS"] == "25"
assert "HEADROOM_TARGET_RATIO" not in env
def test_personas_omit_target_ratio_in_pipeline_kwargs() -> None:
# coding compresses the delta observation (cache mode) → compress_user True;
# general has no positional code working set and leaves user turns intact.
for name, expected_protect, expected_compress_user, expected_min_tokens in (
("coding", 0, True, 10),
("general", 0, False, 25),
):
kwargs = proxy_pipeline_kwargs(ProxyConfig(savings_profile=name))
assert kwargs["protect_recent"] == expected_protect
assert kwargs["read_protection_window"] == expected_protect
assert kwargs["min_tokens_to_compress"] == expected_min_tokens
assert kwargs["compress_user_messages"] is expected_compress_user
assert kwargs["compress_system_messages"] is False
assert kwargs["force_kompress"] is False
assert "target_ratio" not in kwargs # persona never pins a keep-ratio
def test_persona_apply_profile_leaves_target_ratio_untouched() -> None:
cfg = CompressConfig(target_ratio=0.42)
apply_agent_savings_profile(cfg, "coding")
assert cfg.protect_recent == 0
assert cfg.min_tokens_to_compress == 10
assert cfg.target_ratio == 0.42 # persona did not override an explicit ratio
def test_agent_savings_env_defaults_preserve_user_overrides() -> None:
env = {
"HEADROOM_TARGET_RATIO": "0.25",
"HEADROOM_MAX_ITEMS": "12",
}
apply_agent_savings_env_defaults(env, AGENT_90_PROFILE)
assert env["HEADROOM_SAVINGS_PROFILE"] == "agent-90"
assert env["HEADROOM_TARGET_RATIO"] == "0.25"
assert env["HEADROOM_MAX_ITEMS"] == "12"
assert env["HEADROOM_SMART_CRUSHER_COMPACTION"] == "0"
def test_unknown_agent_savings_profile_falls_back_to_default(
caplog: pytest.LogCaptureFixture,
) -> None:
# An unknown profile must NOT raise: it's resolved during proxy startup, so
# raising takes the whole proxy down before it opens its port (desktop asked
# for a profile a fallback runtime predates).
#
# It must degrade to the DEFAULT profile, not to "balanced". The two are not
# interchangeable: balanced flips cache->token mode, turns cross-turn dedup
# and tool-search off, stops compressing user messages, and raises the
# message floor 25x (250 vs 10) and the block floor 20x (500 vs 25). A typo
# in HEADROOM_SAVINGS_PROFILE used to silently reconfigure the entire proxy
# into that posture, which is strictly worse than behaving as if the
# variable were unset.
with caplog.at_level(logging.WARNING):
profile = get_agent_savings_profile("missing")
assert profile is get_agent_savings_profile(None)
assert profile.name == DEFAULT_PROFILE
assert profile is not get_agent_savings_profile("balanced")
assert "unknown savings profile" in caplog.text
assert "missing" in caplog.text
# The warning has to name the resolved profile, so an operator reading it
# knows what they actually got rather than only what they asked for.
assert DEFAULT_PROFILE in caplog.text
def test_with_target_savings_recomputes_target_ratio() -> None:
profile = with_target_savings(get_agent_savings_profile("balanced"), 0.85)
assert profile.target_savings == 0.85
assert profile.target_ratio == 0.15
def test_agent_savings_cli_renders_shell_exports() -> None:
result = CliRunner().invoke(main, ["agent-savings", "--profile", "agent-90"])
assert result.exit_code == 0
assert 'export HEADROOM_SAVINGS_PROFILE="agent-90"' in result.output
assert 'export HEADROOM_SAVINGS_TARGET="0.90"' in result.output
assert 'export HEADROOM_ACCURACY_GUARD="strict"' in result.output
def test_agent_savings_cli_renders_json() -> None:
result = CliRunner().invoke(
main,
["agent-savings", "--profile", "agent-90", "--format", "json"],
)
assert result.exit_code == 0
assert '"HEADROOM_TARGET_RATIO": "0.10"' in result.output
def test_compress_applies_agent_savings_profile_to_pipeline(monkeypatch) -> None:
captured: dict[str, object] = {}
messages = [{"role": "user", "content": "x" * 500}]
class Pipeline:
def apply(self, **kwargs):
captured.update(kwargs)
return SimpleNamespace(
messages=messages,
tokens_before=1000,
tokens_after=100,
transforms_applied=["test"],
)
monkeypatch.setattr(compress_module, "_get_pipeline", lambda: Pipeline())
config = CompressConfig()
apply_agent_savings_profile(config, AGENT_90_PROFILE)
result = compress(messages, config=config)
assert result.compression_ratio == 0.9
assert captured["compress_user_messages"] is True
assert captured["compress_system_messages"] is True
assert captured["protect_recent"] == 2
assert captured["protect_analysis_context"] is True
assert captured["target_ratio"] == 0.10
assert captured["min_tokens_to_compress"] == 120
def test_compress_savings_profile_does_not_mutate_supplied_config(monkeypatch) -> None:
captured: dict[str, object] = {}
messages = [{"role": "user", "content": "x" * 500}]
config = CompressConfig(
compress_user_messages=False,
compress_system_messages=False,
protect_recent=9,
protect_analysis_context=False,
target_ratio=None,
min_tokens_to_compress=999,
)
class Pipeline:
def apply(self, **kwargs):
captured.update(kwargs)
return SimpleNamespace(
messages=messages,
tokens_before=1000,
tokens_after=100,
transforms_applied=["test"],
)
monkeypatch.setattr(compress_module, "_get_pipeline", lambda: Pipeline())
compress(messages, config=config, savings_profile=AGENT_90_PROFILE)
assert captured["target_ratio"] == 0.10
assert captured["min_tokens_to_compress"] == 120
assert config.compress_user_messages is False
assert config.compress_system_messages is False
assert config.protect_recent == 9
assert config.protect_analysis_context is False
assert config.target_ratio is None
assert config.min_tokens_to_compress == 999
def test_wrap_agent_savings_profile_is_opt_in(monkeypatch) -> None:
monkeypatch.delenv("HEADROOM_SAVINGS_PROFILE", raising=False)
assert wrap_module._wrap_agent_savings_profile("codex") is None
monkeypatch.setenv("HEADROOM_SAVINGS_PROFILE", AGENT_90_PROFILE)
assert wrap_module._wrap_agent_savings_profile("codex") == AGENT_90_PROFILE
def test_agent_savings_config_mismatches_requires_explicit_profile(monkeypatch) -> None:
monkeypatch.delenv("HEADROOM_SAVINGS_PROFILE", raising=False)
assert wrap_module._agent_savings_config_mismatches({}, "claude") == []
def test_start_proxy_does_not_inject_agent_savings_by_default(monkeypatch, tmp_path) -> None:
captured_env: dict[str, str] = {}
class Proc:
returncode = None
def poll(self) -> None:
return None
def popen(cmd, **kwargs): # noqa: ANN001
captured_env.update(kwargs["env"])
return Proc()
monkeypatch.delenv("HEADROOM_SAVINGS_PROFILE", raising=False)
monkeypatch.setattr(wrap_module.subprocess, "Popen", popen)
monkeypatch.setattr(wrap_module.time, "sleep", lambda seconds: None)
monkeypatch.setattr(wrap_module, "_check_proxy", lambda port: True)
monkeypatch.setattr(wrap_module, "_get_log_path", lambda: tmp_path / "proxy.log")
wrap_module._start_proxy(8787, agent_type="codex")
assert "HEADROOM_SAVINGS_PROFILE" not in captured_env
assert "HEADROOM_TARGET_RATIO" not in captured_env
def test_start_proxy_injects_explicit_agent_savings_profile(monkeypatch, tmp_path) -> None:
captured_env: dict[str, str] = {}
class Proc:
returncode = None
def poll(self) -> None:
return None
def popen(cmd, **kwargs): # noqa: ANN001
captured_env.update(kwargs["env"])
return Proc()
monkeypatch.setenv("HEADROOM_SAVINGS_PROFILE", AGENT_90_PROFILE)
monkeypatch.setattr(wrap_module.subprocess, "Popen", popen)
monkeypatch.setattr(wrap_module.time, "sleep", lambda seconds: None)
monkeypatch.setattr(wrap_module, "_check_proxy", lambda port: True)
monkeypatch.setattr(wrap_module, "_get_log_path", lambda: tmp_path / "proxy.log")
wrap_module._start_proxy(8787, agent_type="codex")
assert captured_env["HEADROOM_SAVINGS_PROFILE"] == AGENT_90_PROFILE
assert captured_env["HEADROOM_TARGET_RATIO"] == "0.10"
def test_agent_savings_config_mismatches_returns_specific_labels(monkeypatch) -> None:
monkeypatch.setenv("HEADROOM_SAVINGS_PROFILE", AGENT_90_PROFILE)
profile = get_agent_savings_profile(AGENT_90_PROFILE)
running_config = {
"savings_profile": profile.name,
"target_ratio": 0.20,
"compress_user_messages": profile.compress_user_messages,
"compress_system_messages": profile.compress_system_messages,
"protect_recent": profile.protect_recent,
"protect_analysis_context": profile.protect_analysis_context,
"min_tokens_to_crush": profile.min_tokens_to_compress,
"max_items_after_crush": profile.max_items_after_crush,
"smart_crusher_with_compaction": profile.smart_crusher_with_compaction,
"accuracy_guard": profile.accuracy_guard,
}
assert wrap_module._agent_savings_config_mismatches(running_config, "codex") == ["target-ratio"]
def test_agent_savings_config_mismatches_ignores_non_target_agents() -> None:
assert wrap_module._agent_savings_config_mismatches({}, "openhands") == []
def test_agent_savings_config_mismatches_accepts_matching_runtime_config(monkeypatch) -> None:
monkeypatch.setenv("HEADROOM_SAVINGS_PROFILE", AGENT_90_PROFILE)
profile = get_agent_savings_profile(AGENT_90_PROFILE)
running_config = {
"savings_profile": profile.name,
"target_ratio": "0.10",
"compress_user_messages": True,
"compress_system_messages": True,
"protect_recent": "2",
"protect_analysis_context": True,
"min_tokens_to_crush": "120",
"max_items_after_crush": "8",
"smart_crusher_with_compaction": False,
"accuracy_guard": "strict",
}
assert wrap_module._agent_savings_config_mismatches(running_config, "cursor") == []
def test_agent_savings_config_mismatches_reports_unparseable_values(monkeypatch) -> None:
monkeypatch.setenv("HEADROOM_SAVINGS_PROFILE", AGENT_90_PROFILE)
running_config = {
"savings_profile": None,
"target_ratio": "not-a-float",
"compress_user_messages": None,
"compress_system_messages": None,
"protect_recent": "not-an-int",
"protect_analysis_context": None,
"min_tokens_to_crush": object(),
"max_items_after_crush": object(),
"smart_crusher_with_compaction": None,
"accuracy_guard": None,
}
assert wrap_module._agent_savings_config_mismatches(running_config, "claude") == [
"savings-profile",
"target-ratio",
"compress-user-messages",
"compress-system-messages",
"protect-recent",
"protect-analysis-context",
"min-tokens",
"max-items",
"smart-crusher-compaction",
"accuracy-guard",
]
def test_agent_90_profile_applies_to_proxy_config_runtime_kwargs() -> None:
config = ProxyConfig(savings_profile="agent-90")
kwargs = proxy_pipeline_kwargs(config)
assert kwargs["compress_user_messages"] is True
assert kwargs["compress_system_messages"] is True
assert kwargs["protect_recent"] == 2
assert kwargs["protect_analysis_context"] is True
assert kwargs["target_ratio"] == 0.10
assert kwargs["min_tokens_to_compress"] == 120
assert kwargs["max_items_after_crush"] == 8
assert kwargs["smart_crusher_with_compaction"] is False
assert kwargs["force_kompress"] is True
assert kwargs["read_protection_window"] == 2
def test_proxy_explicit_config_overrides_agent_90_profile() -> None:
config = ProxyConfig(
savings_profile="agent-90",
target_ratio=0.25,
protect_recent=5,
min_tokens_to_crush=300,
)
kwargs = proxy_pipeline_kwargs(config)
assert kwargs["target_ratio"] == 0.25
assert kwargs["protect_recent"] == 5
assert kwargs["min_tokens_to_compress"] == 300
def test_agent_90_router_uses_ccr_sampling_not_lossless_table() -> None:
router = ContentRouter(
ContentRouterConfig(
smart_crusher_max_items_after_crush=8,
smart_crusher_with_compaction=False,
)
)
crusher = router._get_smart_crusher()
assert crusher is not None
assert crusher.config.max_items_after_crush == 8
assert crusher._with_compaction is False
def test_router_lossless_only_flag_reaches_crusher() -> None:
# HEADROOM_LOSSLESS_ONLY=1 sets this field on the proxy router; it
# must flow through to the SmartCrusher so a real proxy session runs
# strict marker-free mode.
router = ContentRouter(ContentRouterConfig(smart_crusher_lossless_only=True))
crusher = router._get_smart_crusher()
assert crusher is not None
assert crusher._lossless_only is True
def test_router_lossless_only_defaults_off() -> None:
# Unset (None) must not force the flag — default crushers stay in
# the marker-emitting mode.
router = ContentRouter(ContentRouterConfig())
crusher = router._get_smart_crusher()
assert crusher is not None
assert crusher._lossless_only is False
def test_agent_90_router_json_tool_output_reaches_target_with_needle() -> None:
needle = "CRITICAL_NEEDLE_42"
rows = [
{
"id": i,
"status": "ok",
"message": "normal repeated telemetry payload",
"value": i % 7,
}
for i in range(1000)
]
rows.append(
{
"id": 99999,
"status": "error",
"message": f"{needle} root cause disk full",
"value": 999.99,
}
)
router = ContentRouter(
ContentRouterConfig(
smart_crusher_max_items_after_crush=8,
smart_crusher_with_compaction=False,
)
)
result = router.compress(json.dumps(rows), question=f"Find {needle}")
before = len(result.original.split())
after = len(result.compressed.split())
assert 1 - after / before >= 0.90
assert needle in result.compressed
assert "<<ccr:" in result.compressed
def test_proxy_cli_reads_agent_90_profile_env() -> None:
captured_config: dict[str, ProxyConfig] = {}
def mock_run_server(config: ProxyConfig, **kwargs: object) -> None:
captured_config["config"] = config
runner = CliRunner()
with pytest.MonkeyPatch.context() as mp:
mp.setattr("headroom.proxy.server.run_server", mock_run_server)
result = runner.invoke(
main,
["proxy"],
env={"HEADROOM_SAVINGS_PROFILE": "agent-90"},
catch_exceptions=False,
)
assert result.exit_code == 0, result.output
config = captured_config["config"]
assert config.savings_profile == "agent-90"
assert proxy_pipeline_kwargs(config)["target_ratio"] == 0.10
def test_unit_router_receives_agent_target_ratio() -> None:
seen: dict[str, object] = {}
class Tokenizer:
def count_text(self, text: str) -> int:
return len(text.split())
class Router:
_runtime_target_ratio = None
def compress(self, text: str, **kwargs: object) -> RouterCompressionResult:
seen["target_ratio"] = self._runtime_target_ratio
return RouterCompressionResult(
compressed="short text",
original=text,
strategy_used=CompressionStrategy.KOMPRESS,
strategy_chain=["kompress"],
)
unit = CompressionUnit(
text=("long text " * 40) + "\nRetrieve more: hash=abc123\n",
provider="openai",
endpoint="responses",
role="assistant",
item_type="message",
cache_zone="live",
mutable=True,
min_bytes=1,
metadata={"compress_assistant": "true"},
)
result = compress_unit_with_router(
unit,
router=Router(),
tokenizer=Tokenizer(),
target_ratio=0.10,
)
assert result.modified is True
assert seen["target_ratio"] == 0.10
def test_agent_savings_check_perf_and_accuracy_report_passes(
monkeypatch,
tmp_path,
) -> None:
from headroom.perf import analyzer
monkeypatch.setattr(analyzer, "parse_log_files", lambda last_n_hours: object())
monkeypatch.setattr(
analyzer,
"build_perf_summary",
lambda report: {"savings_pct": 92.0},
)
report = tmp_path / "eval.json"
report.write_text(json.dumps({"totals": {"accuracy_rate": 1.0}}))
result = CliRunner().invoke(
main,
[
"agent-savings",
"--profile",
"agent-90",
"--check-perf",
"--accuracy-report",
str(report),
],
)
assert result.exit_code == 0, result.output
assert "92.0% savings meets 90.0%" in result.output
assert "100.0% accuracy meets 90.0%" in result.output
def test_agent_savings_accuracy_report_below_threshold_fails(
monkeypatch,
tmp_path,
) -> None:
from headroom.perf import analyzer
monkeypatch.setattr(analyzer, "parse_log_files", lambda last_n_hours: object())
monkeypatch.setattr(
analyzer,
"build_perf_summary",
lambda report: {"savings_pct": 92.0},
)
report = tmp_path / "eval.json"
report.write_text(json.dumps({"totals": {"accuracy_rate": 0.89}}))
result = CliRunner().invoke(
main,
[
"agent-savings",
"--profile",
"agent-90",
"--check-perf",
"--accuracy-report",
str(report),
],
)
assert result.exit_code != 0
assert "89.0% accuracy below 90.0%" in result.output
def test_agent_savings_requires_each_agent_to_meet_target(monkeypatch) -> None:
from headroom.perf import analyzer
from headroom.perf.analyzer import PerfRecord, PerfReport
report = PerfReport(
perf_records=[
PerfRecord(
timestamp="2026-06-10 10:00:00,000",
request_id="claude-1",
model="claude-sonnet",
client="claude",
tokens_before=1000,
tokens_after=80,
tokens_saved=920,
),
PerfRecord(
timestamp="2026-06-10 10:01:00,000",
request_id="codex-1",
model="gpt-5",
client="codex",
tokens_before=1000,
tokens_after=90,
tokens_saved=910,
),
PerfRecord(
timestamp="2026-06-10 10:02:00,000",
request_id="cursor-1",
model="gpt-5",
client="cursor",
tokens_before=1000,
tokens_after=70,
tokens_saved=930,
),
]
)
monkeypatch.setattr(analyzer, "parse_log_files", lambda last_n_hours: report)
result = CliRunner().invoke(
main,
[
"agent-savings",
"--check-perf",
"--require-agents",
"claude,codex,cursor",
],
)
assert result.exit_code == 0, result.output
assert "claude: 92.0% savings meets 90.0%" in result.output
assert "codex: 91.0% savings meets 90.0%" in result.output
assert "cursor: 93.0% savings meets 90.0%" in result.output
def test_agent_savings_required_agent_missing_fails(monkeypatch) -> None:
from headroom.perf import analyzer
from headroom.perf.analyzer import PerfRecord, PerfReport
report = PerfReport(
perf_records=[
PerfRecord(
timestamp="2026-06-10 10:00:00,000",
request_id="claude-1",
model="claude-sonnet",
client="claude",
tokens_before=1000,
tokens_after=80,
tokens_saved=920,
),
PerfRecord(
timestamp="2026-06-10 10:01:00,000",
request_id="codex-1",
model="gpt-5",
client="codex",
tokens_before=1000,
tokens_after=90,
tokens_saved=910,
),
]
)
monkeypatch.setattr(analyzer, "parse_log_files", lambda last_n_hours: report)
result = CliRunner().invoke(
main,
[
"agent-savings",
"--check-perf",
"--require-agents",
"claude,codex,cursor",
],
)
assert result.exit_code != 0
assert "missing required agent traffic: cursor" in result.output
def test_agent_savings_writes_three_agent_smoke_fixture(tmp_path) -> None:
workspace = tmp_path / "workspace"
result = CliRunner().invoke(
main,
["agent-savings", "--write-smoke-fixture", str(workspace)],
)
assert result.exit_code == 0, result.output
assert (workspace / "logs" / "proxy.log").exists()
eval_report = workspace / "agent-90-eval.json"
assert eval_report.exists()
assert "--require-agents claude,codex,cursor" in result.output
def test_agent_savings_smoke_fixture_passes_real_gate(tmp_path) -> None:
workspace = tmp_path / "workspace"
runner = CliRunner()
write_result = runner.invoke(
main,
["agent-savings", "--write-smoke-fixture", str(workspace)],
)
assert write_result.exit_code == 0, write_result.output
gate_result = runner.invoke(
main,
[
"agent-savings",
"--check-perf",
"--hours",
"0",
"--require-agents",
"claude,codex,cursor",
"--accuracy-report",
str(workspace / "agent-90-eval.json"),
],
env={"HEADROOM_WORKSPACE_DIR": str(workspace)},
)
assert gate_result.exit_code == 0, gate_result.output
assert "claude: 92.0% savings meets 90.0%" in gate_result.output
assert "codex: 91.0% savings meets 90.0%" in gate_result.output
assert "cursor: 93.0% savings meets 90.0%" in gate_result.output
assert "100.0% accuracy meets 90.0%" in gate_result.output
def test_coding_profile_min_chars_block_reaches_router_without_env_seeding(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The block-char floor must travel on the config object, not env only.
Every other router pipeline kwarg this function builds travels on the config
object; ``min_chars_for_block`` alone was populated only from
``HEADROOM_MIN_CHARS_FOR_BLOCK`` (emitted by ``proxy_env()``). A proxy whose
config carried ``savings_profile="coding"`` but whose process env was never
seeded applied every sibling coding knob while this floor silently stayed at
``ContentRouterConfig.min_chars_for_block_compression`` (500) instead of the
profile's 25 — a 20x gap on the gate that governs tool_result blocks.
Note this does not make the profile fully config-deliverable: fields whose
consumers read ``os.environ`` directly (``cross_turn_dedup`` via
ContentRouter, ``tool_search`` via the Anthropic handler) never pass through
this function and remain seed-dependent.
"""
monkeypatch.delenv("HEADROOM_MIN_CHARS_FOR_BLOCK", raising=False)
class _Config:
savings_profile = "coding"
min_tokens_to_crush = 500
kwargs = proxy_pipeline_kwargs(_Config())
assert kwargs["min_chars_for_block_compression"] == 25
assert kwargs["min_tokens_to_compress"] == 10
def test_explicit_min_chars_block_env_overrides_the_profile(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""An explicit operator override still wins over the profile value."""
monkeypatch.setenv("HEADROOM_MIN_CHARS_FOR_BLOCK", "120")
class _Config:
savings_profile = "coding"
min_tokens_to_crush = 500
assert proxy_pipeline_kwargs(_Config())["min_chars_for_block_compression"] == 120