headroom/tests/test_content_router_exclude_tools.py
Ingmar Krusch 3d0e59e518
fix(content-router): protect_tool_results must not be weakened by profile-derived read_protection_window (#2105)
## Description

`ContentRouter.apply()` computes `read_protection_window` from
`protect_recent_reads_fraction`, where `0.0` (the sentinel
`--protect-tool-results` sets, per #1374's documented contract) means
"protect all excluded-tool output regardless of conversation depth." The
method then unconditionally overwrote that window with a per-request
`read_protection_window` kwarg whenever one was present.
`proxy_pipeline_kwargs()` supplies that kwarg on every request from the
active `AgentSavingsProfile.protect_recent` (the default `coding`
profile sets `protect_recent=2`), so in practice only the last 2
messages ever kept read-protection regardless of
`--protect-tool-results` — older excluded-tool output (`Read`, `Glob`,
`Grep`, `Write`, `Edit` results) silently fell through to lossy Kompress
compression.

Closes #

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- `headroom/transforms/content_router.py`: the runtime
`read_protection_window` kwarg may now only *narrow* the window when
`self.config.protect_recent_reads_fraction > 0`. It can no longer
override the `0.0` ("protect everything") sentinel that
`--protect-tool-results` sets.
- `tests/test_content_router_exclude_tools.py`: regression coverage that
`--protect-tool-results`-equivalent config
(`protect_recent_reads_fraction=0.0`) stays fully protected even when a
savings-profile kwarg would otherwise shrink the window.
- `tests/test_transforms/test_content_router.py`: unit coverage of the
precedence logic itself (kwarg narrows when fraction > 0, kwarg is
ignored when fraction == 0.0).
- `CHANGELOG.md`: added an `### Bug Fixes` entry under `Unreleased`.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ uv run pytest tests/test_content_router_exclude_tools.py tests/test_transforms/test_content_router.py -q
============================= test session starts ==============================
platform darwin -- Python 3.13.13, pytest-9.0.3, pluggy-1.6.0
collected 64 items

tests/test_content_router_exclude_tools.py ......                        [  9%]
tests/test_transforms/test_content_router.py ........................... [ 51%]
...............................                                          [100%]

============================== 64 passed in 2.77s ==============================

$ uv run ruff check headroom/transforms/content_router.py tests/test_content_router_exclude_tools.py tests/test_transforms/test_content_router.py
All checks passed!

$ uv run mypy headroom/transforms/content_router.py
Success: no issues found in 1 source file
```

## Real Behavior Proof

- **Environment:** personal fork deployed as a real proxy (macOS launchd
service, `headroom install apply`) with `--backend bedrock --mode token
--code-aware --protect-tool-results Bash`,
`HEADROOM_SAVINGS_PROFILE=coding` (library default `protect_recent=2`),
fronting a live Claude Code session.
- **Exact command / steps:** in a long-running Claude Code session
against this deployment, `Read` a source file, continue the conversation
past 2 more assistant turns (so the file's `Read` result ages past the
profile's `protect_recent=2` window), then have the agent re-read or
reference the same file.
- **Observed result:** before the fix, the aged `Read` output for a
plain (non-code) file came back as `[N items compressed to M. Retrieve
more: hash=...]` despite `--protect-tool-results` being set and `Read`
sitting in `DEFAULT_EXCLUDE_TOOLS` — confirmed by direct proxy log
inspection (`content_router.py`'s override silently winning over the
`0.0` sentinel) and by byte-diffing the installed pipx package against
this same fork's git source to rule out a stale build. After applying
the fix, the same sequence leaves the aged `Read` output intact (no
compression marker) — verified via `pytest` regression tests plus a
fresh live-session check post-deploy.
- **Not tested:** this deployment has since switched to `--mode cache`
(upstream's tested/benchmarked default for the `coding` profile as of
`68676daa`), where the whole `read_protection_window` mechanism this bug
lives in is structurally unreachable for anything inside the frozen
prefix — so the precedence fix in this PR is primarily relevant to
`token`-mode deployments (or any deployment where cache mode's
frozen-prefix boundary hasn't yet advanced past the affected message).
It has not been independently re-verified live under `--mode token`
after the most recent rebase onto `main` (only the automated test suite
was rerun post-rebase).

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable

## Screenshots (if applicable)

N/A — backend logic change, no UI surface.

## Additional Notes

- "I have made corresponding changes to the documentation" is unchecked:
no file in `docs/`, `README.md`, or `CONTRIBUTING.md` documents
`read_protection_window`, `protect_recent_reads_fraction`, or
`--protect-tool-results` precedence at all, so there was no existing
section to update, and no new section was added either. This is arguably
a pre-existing documentation gap this PR doesn't close.
- No linked issue number: this was found via independent investigation
of a personal deployment, not filed as a `headroomlabs-ai/headroom`
issue first.

Co-authored-by: Ingmar Krusch <ingmar.krusch@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Tejas Chopra <chopratejas@gmail.com>
2026-07-13 14:01:18 -04:00

226 lines
8.6 KiB
Python

"""Tests for --protect-tool-results / HEADROOM_PROTECT_TOOL_RESULTS.
Three behavioral tests:
1. protect_tool_results merges into the exclude set so named tools are never
lossy-compressed.
2. _parse_csv_tools parses CSV strings without merging HEADROOM_EXCLUDE_TOOLS.
3. ContentRouter with Bash in exclude_tools passes Bash tool_result verbatim.
"""
from __future__ import annotations
import pytest
from headroom.config import DEFAULT_EXCLUDE_TOOLS
from headroom.proxy.server import (
HeadroomProxy,
ProxyConfig,
_parse_csv_tools,
)
def _build(**overrides: object) -> HeadroomProxy:
config = ProxyConfig(
optimize=False,
cache_enabled=False,
rate_limit_enabled=False,
cost_tracking_enabled=False,
code_aware_enabled=False,
**overrides,
)
return HeadroomProxy(config)
def _router(proxy: HeadroomProxy):
# ContentRouter is the last transform in the Anthropic pipeline.
return proxy.anthropic_pipeline.transforms[-1]
# ---------------------------------------------------------------------------
# Test 1: protect_tool_results merges into exclude set
# ---------------------------------------------------------------------------
def test_protect_tool_results_merges_into_exclude_set() -> None:
"""Bash added via protect_tool_results must appear in exclude_tools alongside
the built-in defaults (e.g. Read), base-fails / head-passes.
The frozenset is merged as-is; lowercase normalization is handled by
_parse_exclude_tools on the CLI/env path (tested separately).
"""
proxy = _build(protect_tool_results=frozenset({"Bash", "bash"}))
exclude = _router(proxy).config.exclude_tools
assert exclude is not None, "exclude_tools must be set when protect_tool_results is non-empty"
assert "Bash" in exclude, "Bash must be in exclude_tools after protect_tool_results merges"
assert "bash" in exclude, (
"lowercase bash must be in exclude_tools after protect_tool_results merges"
)
assert "Read" in exclude, "Read (built-in default) must still be in exclude_tools"
def test_protect_tool_results_disables_age_decay_in_token_mode() -> None:
"""In token mode, protect_tool_results forces protect_recent_reads_fraction to 0.0
so protected tools are never compressed by age-decay."""
proxy = _build(protect_tool_results=frozenset({"Bash", "bash"}), mode="token")
assert _router(proxy).config.protect_recent_reads_fraction == 0.0
# ---------------------------------------------------------------------------
# Test 2: CSV env var / CLI string parsing
# ---------------------------------------------------------------------------
def test_protect_tool_results_env_var_csv() -> None:
"""_parse_csv_tools parses a comma-separated value into both original-case
and lowercase entries without merging HEADROOM_EXCLUDE_TOOLS."""
result = _parse_csv_tools("Bash,WebFetch")
assert "Bash" in result
assert "bash" in result
assert "WebFetch" in result
assert "webfetch" in result
# ---------------------------------------------------------------------------
# Test 3: Bash tool_result passthrough when protected
# ---------------------------------------------------------------------------
def test_bash_tool_result_passthrough_when_protected() -> None:
"""When Bash is in exclude_tools (via protect_tool_results), its tool_result
content passes through the ContentRouter verbatim without lossy compression.
Base-fails / head-passes."""
pytest.importorskip("tiktoken") # needed for OpenAI tokenizer
from headroom.providers import OpenAIProvider
from headroom.tokenizer import Tokenizer
from headroom.transforms.content_router import ContentRouter, ContentRouterConfig
provider = OpenAIProvider()
token_counter = provider.get_token_counter("gpt-4o")
tokenizer = Tokenizer(token_counter, "gpt-4o")
# Build router with Bash explicitly in exclude_tools
config = ContentRouterConfig(
min_section_tokens=10,
exclude_tools=set(DEFAULT_EXCLUDE_TOOLS) | {"Bash", "bash"},
)
router = ContentRouter(config)
bash_output = "\n".join(
f"line {i}: some output from a bash command that is long enough to compress"
for i in range(80)
)
messages = [
{
"role": "assistant",
"content": None,
"tool_calls": [
{
"id": "call_bash_1",
"type": "function",
"function": {"name": "Bash", "arguments": "{}"},
}
],
},
{
"role": "tool",
"tool_call_id": "call_bash_1",
"content": bash_output,
},
]
result = router.apply(messages, tokenizer)
# Bash tool_result must pass through unchanged
tool_msg = next(m for m in result.messages if m.get("tool_call_id") == "call_bash_1")
assert tool_msg["content"] == bash_output, (
"Bash tool_result must be verbatim when Bash is in exclude_tools"
)
assert "router:excluded:tool" in result.transforms_applied
# ---------------------------------------------------------------------------
# Test 4: protect_tool_results sentinel survives a profile-derived
# read_protection_window kwarg, even when the protected output is old
# ---------------------------------------------------------------------------
def test_protect_tool_results_survives_runtime_read_protection_window_kwarg() -> None:
"""A profile-derived `read_protection_window` kwarg (e.g. from
AgentSavingsProfile.protect_recent=2, threaded in via
proxy_pipeline_kwargs()) must not shrink protection below what
protect_recent_reads_fraction == 0.0 (the --protect-tool-results
sentinel) already guarantees for the whole conversation.
Regression test for the precedence bug: content_router.py used to apply
the runtime kwarg unconditionally, so a Bash tool_result more than
`read_protection_window` messages old fell through to lossy compression
even though --protect-tool-results promised it would never compress
"regardless of conversation depth" (see PR #1374)."""
pytest.importorskip("tiktoken") # needed for OpenAI tokenizer
from headroom.providers import OpenAIProvider
from headroom.tokenizer import Tokenizer
provider = OpenAIProvider()
token_counter = provider.get_token_counter("gpt-4o")
tokenizer = Tokenizer(token_counter, "gpt-4o")
proxy = _build(protect_tool_results=frozenset({"Bash", "bash"}), mode="token")
router = _router(proxy)
bash_output = "\n".join(
f"line {i}: some output from a bash command that is long enough to compress"
for i in range(80)
)
messages: list[dict[str, object]] = [
{
"role": "assistant",
"content": None,
"tool_calls": [
{
"id": "call_bash_1",
"type": "function",
"function": {"name": "Bash", "arguments": "{}"},
}
],
},
{
"role": "tool",
"tool_call_id": "call_bash_1",
"content": bash_output,
},
]
# Pad with enough intervening turns that the Bash tool_result above
# falls outside a read_protection_window=2 (it's ~9-10 messages from
# the end once padding is added).
for i in range(8):
messages.append({"role": "user", "content": f"follow-up turn {i}"})
messages.append({"role": "assistant", "content": f"reply {i}"})
# Simulate the profile-derived kwarg the proxy threads into every
# request via proxy_pipeline_kwargs() (AgentSavingsProfile("coding")
# sets protect_recent=2).
result = router.apply(messages, tokenizer, read_protection_window=2)
tool_msg = next(m for m in result.messages if m.get("tool_call_id") == "call_bash_1")
assert tool_msg["content"] == bash_output, (
"Bash tool_result must stay verbatim: protect_recent_reads_fraction == 0.0 "
"(set by --protect-tool-results) must not be weakened by a profile-derived "
"read_protection_window kwarg"
)
assert "router:excluded:tool" in result.transforms_applied
# ---------------------------------------------------------------------------
# Baseline: Bash NOT in DEFAULT_EXCLUDE_TOOLS (unchanged by this PR)
# ---------------------------------------------------------------------------
def test_bash_not_in_default_exclude_tools() -> None:
"""Bash must remain absent from DEFAULT_EXCLUDE_TOOLS; protect_tool_results
is the opt-in path."""
assert "Bash" not in DEFAULT_EXCLUDE_TOOLS
assert "bash" not in DEFAULT_EXCLUDE_TOOLS