fix(proxy): stop re-compressing headroom_retrieve output and emitting unredeemable markers (#1323)

## Description

Two related CCR problems that both end in unreadable content.

The first one (#1077) is an infinite loop. Any tool output over ~500
bytes gets replaced with a `<<ccr:hash>>` marker, and you call
`headroom_retrieve` to get the original back. But the proxy then
compresses the *retrieve response too*, so what comes back is a brand
new marker. Retrieve that one and you get another marker.

The second one (#1006), the proxy makes two independent decisions per
request: SmartCrusher compresses, and the `headroom_retrieve` tool gets
injected. The injection is deferred when there's a frozen message prefix
(`frozen_message_count > 0`), but compression keeps running anyway. So
the agent receives `[... compressed to N. Retrieve more: hash=...]`
markers with no `headroom_retrieve` tool to redeem them.

For #1077, SmartCrusher now skips `headroom_retrieve` results. Before
crushing a tool message (OpenAI `role=tool`) or tool-result block
(Anthropic `type=tool_result`), it checks whether that tool id maps to
the CCR tool, and if so leaves it alone. Retrieved content stays
readable.

For #1006, compression and injection are no longer decided in isolation.
The injection decision is extracted into `should_inject_ccr_tool`, which
the Anthropic handler calls: when injection was deferred because of a
frozen prefix but compression just emitted new markers, it injects the
tool anyway, so a marker is never handed to an agent that can't act on
it. The existing session-sticky dedup means sessions that already have
the tool don't get it re-injected and don't lose their cache.

Closes #1077
Closes #1006

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- `headroom/transforms/smart_crusher.py`: exempt `headroom_retrieve`
results from compression on both the OpenAI `role=tool` and Anthropic
`type=tool_result` paths.
- `headroom/proxy/helpers.py`: add `should_inject_ccr_tool`, the
deferral-plus-override decision the handler used to inline, so the #1006
behaviour is testable at the decision point.
- `headroom/proxy/handlers/anthropic.py`: call `should_inject_ccr_tool`
to couple injection with compression; rename the misleading
`frozen_prefix=` log key to `frozen_message_count=`.
- `tests/test_transforms/test_smart_crusher_ccr_retrieve_exemption.py`
and `tests/test_proxy/test_ccr_frozen_prefix_coupling.py`: new tests;
the frozen-prefix test now drives `should_inject_ccr_tool` so it would
fail if the override were removed.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed

### Test Output

```text
$ uv run --extra dev python -m pytest tests/test_proxy/test_ccr_frozen_prefix_coupling.py tests/test_transforms/test_smart_crusher_ccr_retrieve_exemption.py -q
5 passed, 1 skipped
ruff: All checks passed!
mypy: Success: no issues found
```

The SmartCrusher test skips locally because the Rust extension `.so` is
built for a different OS, the same skip the existing SmartCrusher tests
take locally. It runs in CI where the extension is built.

## Real Behavior Proof

- Environment: macOS, Python 3.13, this branch.
- Exact command / steps: `uv run --extra dev python -m pytest
tests/test_proxy/test_ccr_frozen_prefix_coupling.py
tests/test_transforms/test_smart_crusher_ccr_retrieve_exemption.py -q`.
The frozen-prefix test calls `should_inject_ccr_tool` (the function the
Anthropic handler now uses) with a frozen prefix and freshly emitted
markers, then drives `apply_session_sticky_ccr_tool` end to end and
asserts `headroom_retrieve` lands in the outbound tools. The exemption
test runs a `headroom_retrieve` tool result through SmartCrusher on both
the OpenAI and Anthropic shapes.
- Observed result: 5 passed, 1 skipped. The retrieve tool is injected
even under a frozen prefix once markers exist, and is not injected when
no markers were emitted. Removing the handler override flips
`should_inject_ccr_tool` and fails the test.
- Not tested: a full live proxy session. The behaviours are covered at
the decision, transform, and handler-call level by the new tests.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable

## Additional Notes

This one touches compression gating, so it's worth a careful read on the
injection coupling, that's the part where a wrong call would
re-introduce data loss.

1. Tool results with no id mapping still compress, marked with `#
ponytail:` comments. Only ids we can positively identify as the CCR tool
are exempted.
2. The injection coupling keys off `injector.has_compressed_content`, so
the tool only shows up when there's actually something to retrieve.

---------

Co-authored-by: JD Davis <mxjerrett@gmail.com>
This commit is contained in:
Lucas Santos 2026-06-25 17:11:42 +02:00 committed by GitHub
parent 9758817979
commit 43494ff526
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
6 changed files with 445 additions and 15 deletions

View file

@ -8,20 +8,14 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## Unreleased
### Features
* **wrap:** `headroom wrap claude --1m` preserves the 1M context window. Behind a custom `ANTHROPIC_BASE_URL` (the proxy) Claude Code drops the `context-1m` beta header and caps the window at 200k for entitled subscription users; the opt-in flag sets `ANTHROPIC_MODEL=<opus>[1m]` on the launched process so the 1M window activates through Headroom. A model already selected via `ANTHROPIC_MODEL` is preserved (only the `[1m]` suffix is appended) ([#1158](https://github.com/chopratejas/headroom/issues/1158)).
### Changed
* **telemetry:** anonymous usage telemetry is now **opt-in** (off by default) instead of opt-out. Nothing is collected or sent unless you set `HEADROOM_TELEMETRY=on` or pass `--telemetry` to `headroom proxy` / `headroom install apply`. `is_telemetry_enabled()` is fail-closed — only explicit on-values (`on`/`true`/`1`/`yes`/`enable`/`enabled`) enable it; unset, empty, or unrecognized values stay disabled. The existing `--no-telemetry` flag and `HEADROOM_TELEMETRY=off` remain accepted for back-compat, and install manifests now write the `HEADROOM_TELEMETRY` value explicitly so generated deployments are unambiguous.
### Fixed
* **rtk:** stop `rtk` hook registration from spuriously timing out during `headroom wrap`. Output is captured to a temp file instead of pipes, and `stdin` is closed, so a background process forked by `rtk init` can no longer hold the pipe open and block `subprocess.run` past its 10s timeout after the hooks were already registered.
### Features
* **wrap:** `headroom wrap claude --1m` preserves the 1M context window. Behind a custom `ANTHROPIC_BASE_URL` (the proxy) Claude Code drops the `context-1m` beta header and caps the window at 200k for entitled subscription users; the opt-in flag sets `ANTHROPIC_MODEL=<opus>[1m]` on the launched process so the 1M window activates through Headroom. A model already selected via `ANTHROPIC_MODEL` is preserved (only the `[1m]` suffix is appended) ([#1158](https://github.com/chopratejas/headroom/issues/1158)).
* **learn:** weight loops in `headroom learn`. A new loop detector (`headroom/learn/loops.py`) recognizes repeated tool-call patterns — including RTK re-fetch loops, where RTK's output truncation makes the agent re-run larger-limit variants of a *successful* command — collapses output-limit variants to one signature, measures the wasted tokens, surfaces loops as a highest-priority digest section, and weights loop guardrails above one-off rules by their measured waste. Previously loops had no special weight and a no-failure re-fetch loop was skipped entirely. Adds an RTK-loop eval (`benchmarks/rtk_loop_learn_eval.py`) that reproduces a loop, runs it through Learn, and asserts the generated guardrail ranks first and prevents re-triggering.
* **learn:** write per-project learnings to the personal, gitignored `CLAUDE.local.md` by default instead of the team-shared `CLAUDE.md`, matching Claude Code's memory convention so machine-specific paths and tool-discovery byproducts no longer pollute the shared file. Adds a `--target` flag to override the destination (e.g. `--target CLAUDE.md` to opt back into the shared file, or any custom path), and auto-migrates a stale learned-patterns block out of an existing `CLAUDE.md` into `CLAUDE.local.md` with a warning ([#1072](https://github.com/chopratejas/headroom/issues/1072)).
* **proxy/transforms:** take large cold-start contexts off the synchronous kompress path — the root cause behind the `compression_first_stage` 30s-timeout + leaked-thread → executor-saturation cascade ([#1171](https://github.com/chopratejas/headroom/issues/1171)). A token size-gate inside the ML boundary routes oversized text away from ModernBERT (`HEADROOM_KOMPRESS_MAX_TOKENS`); a cooperative chunk-deadline bounds any kompress run that does proceed (`HEADROOM_COMPRESSION_DEADLINE_MS`); an opt-in off-path mode forwards uncompressed immediately and compresses in a single per-process background drain so the request never blocks on ML (`HEADROOM_BACKGROUND_COMPRESSION`); and a new native `TextCrusher` — a fast deterministic extractive prose compressor in `headroom._core` that reuses the shared BM25 relevance scorer — is the fast alternative to ModernBERT for large plain text (`HEADROOM_TEXT_CRUSHER`). All default off and fail-open. On a SQuAD answer-retention eval (requires the SQuAD dev set) TextCrusher keeps ~94% of buried answers at 30% size vs ~36% for truncate/random, and runs in one O(n) pass -- sub-second where ModernBERT takes minutes (self-contained speed benchmark in `benchmarks/text_crusher_quality_eval.py`).
@ -29,15 +23,14 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
* **vibe:** add Mistral Vibe CLI support with `headroom wrap vibe`.
* **proxy:** per-project savings breakdown on the dashboard for all wrapped agents — Claude Code, Codex, aider, Copilot, and Cursor ([#802](https://github.com/chopratejas/headroom/issues/802)). `headroom wrap claude`/`codex` tag requests with an `X-Headroom-Project` header (launch-directory name); `wrap aider`/`copilot`/`cursor` — whose clients cannot send custom headers — use a `/p/<name>` base-URL prefix the proxy strips. Savings are aggregated per project (persisted, schema v3 with transparent v2 migration), exposed as `savings.per_project` in `/stats` and `projects` in `/stats-history`, and shown in a Per-Project Savings dashboard table.
* **memory:** opt-in Apple-GPU (MPS) embedding offload via `HEADROOM_EMBEDDER_RUNTIME=pytorch_mps`. When set (and Apple MPS is available), the memory embedder runs on the torch sentence-transformers backend on the Apple GPU instead of the default ONNX CPU embedder, freeing the CPU under load. If MPS or the dependencies are unavailable, Headroom logs a warning and uses the existing default embedder selection path (ONNX when available, then the pre-existing local fallback). MPS encode calls are serialized internally (torch-MPS is not thread-safe). Adds the new `[pytorch-mps]` extra (`pip install 'headroom-ai[pytorch-mps]'`). Default behavior is unchanged.
### Features
* **proxy:** cross-region Bedrock inference-profile detection — geo-prefixed model IDs (`eu.`/`us.`/`apac.`/`global.`) are now resolved to their canonical vendor, so Anthropic cross-region profiles (e.g. `eu.anthropic.claude-haiku-4-5-20251001-v1:0`) receive live-zone compression instead of being silently skipped ([#999](https://github.com/chopratejas/headroom/pull/999)).
* **proxy:** Converse-body compression on the native Bedrock route — the live-zone dispatcher now recognizes Bedrock Converse content blocks (typeless `{"text": …}`, not only Anthropic `{"type":"text", …}`), so Converse user-message text compresses; `run_anthropic_compression` no longer bails to passthrough when the body lacks an InvokeModel `anthropic_version` envelope, and envelope re-emit stays gated on successful parse ([#999](https://github.com/chopratejas/headroom/pull/999)).
* **docker:** bundle `headroom-proxy` binary in published `runtime` and `runtime-slim` images — closes [#976](https://github.com/chopratejas/headroom/issues/976) ([#999](https://github.com/chopratejas/headroom/pull/999)).
### Bug Fixes
* **rtk:** stop `rtk` hook registration from spuriously timing out during `headroom wrap`. Output is captured to a temp file instead of pipes, and `stdin` is closed, so a background process forked by `rtk init` can no longer hold the pipe open and block `subprocess.run` past its 10s timeout after the hooks were already registered.
* **ccr:** stop re-compressing `headroom_retrieve` output, which created an infinite retrieval loop, and stop emitting retrieval markers when the `headroom_retrieve` tool is not injected, which silently dropped data ([#1077](https://github.com/chopratejas/headroom/issues/1077), [#1006](https://github.com/chopratejas/headroom/issues/1006)).
* **dashboard:** include RTK stats in the Historical tab; `/stats-history` now attaches live RTK/CLI-filtering stats the same way the Session tab does, so they survive a proxy restart ([#1177](https://github.com/chopratejas/headroom/issues/1177)).
* **proxy:** stop discarding a finished compression on very large requests. After the transform pipeline completed, a telemetry-only waste-signal re-parse of the *original* messages ran on the critical path; on huge Claude Code transcripts (~400k tokens) that parse could exceed the Anthropic compression timeout, so the proxy failed open and forwarded the uncompressed request despite "Pipeline complete" logging real savings (`tokens_saved: 0`, `transforms_applied: []`, ~31s latency). Waste-signal detection is now skipped above `MAX_WASTE_SIGNAL_DETECTION_TOKENS` (100k) so the compression result stays on the critical path ([#296](https://github.com/chopratejas/headroom/issues/296)).
* **codex:** retag existing Codex threads when `headroom init` injects the `headroom` provider, so Codex Desktop history stays visible. Codex filters its sidebar/search by the active `model_provider`; the init path set `model_provider = "headroom"` without retagging, so existing native `openai` threads disappeared from the menu (data was never deleted, only hidden). `_ensure_codex_provider` now reconciles thread tags openai→headroom, matching what the install and `wrap` paths already do; `headroom unwrap codex` handles the revert direction ([#961](https://github.com/chopratejas/headroom/issues/961)).

View file

@ -1456,13 +1456,12 @@ class AnthropicHandlerMixin:
f"(frozen prefix={frozen_message_count}) to preserve cache"
)
inject_system_instructions = False
inject_tool = self.config.ccr_inject_tool
if inject_tool and frozen_message_count > 0:
configured_inject_tool = self.config.ccr_inject_tool
if configured_inject_tool and frozen_message_count > 0:
logger.info(
f"[{request_id}] CCR: deferring tool injection "
f"(frozen prefix={frozen_message_count}) to preserve cache"
f"(frozen_message_count={frozen_message_count}) to preserve cache"
)
inject_tool = False
# Scan for compression markers + maybe inject system instructions.
# Tool-list injection is handled separately via the sticky helper.
injector = CCRToolInjector(
@ -1477,7 +1476,31 @@ class AnthropicHandlerMixin:
# Sticky-on tool registration (PR-B7): always inject the
# retrieval tool once a session has done CCR, regardless
# of whether THIS turn produced compressed content.
if inject_tool:
#
# #1006: if tool injection was deferred (frozen prefix) but
# compression just emitted NEW markers this turn, override the
# deferral — the agent has no other way to redeem those markers.
# The cache miss on this one request is preferable to silent
# data loss. If the session has already done CCR the tool is
# already in the client's tool list, so sticky replay is a
# no-op and the cache is unaffected.
# ponytail: ceiling is one extra cache miss on the first CCR
# turn in a frozen-prefix session.
from headroom.proxy.helpers import should_inject_ccr_tool
should_inject, is_marker_override = should_inject_ccr_tool(
configured_inject_tool=configured_inject_tool,
frozen_message_count=frozen_message_count,
has_compressed_content=injector.has_compressed_content,
)
if should_inject:
if is_marker_override:
logger.info(
f"[{request_id}] CCR: overriding injection deferral — "
f"new markers emitted but headroom_retrieve unavailable "
f"(frozen_message_count={frozen_message_count}); injecting to "
"prevent unredeemable markers (#1006)"
)
from headroom.proxy.helpers import apply_session_sticky_ccr_tool
tools, ccr_tool_injected = apply_session_sticky_ccr_tool(

View file

@ -2546,6 +2546,35 @@ def _reset_session_ccr_tracker_for_test() -> None:
_session_ccr_tracker = None
def should_inject_ccr_tool(
*,
configured_inject_tool: bool,
frozen_message_count: int,
has_compressed_content: bool,
) -> tuple[bool, bool]:
"""Decide whether the ``headroom_retrieve`` tool must be injected this turn.
This is the decision the Anthropic handler used to inline. It is extracted
so the #1006 regression can be pinned at the decision point itself.
Tool injection is normally deferred when there is a frozen message prefix
(``frozen_message_count > 0``) to preserve the prompt cache. But if
compression emitted fresh markers this turn, deferring would hand the agent
a ``<<ccr:hash>>`` marker with no tool to redeem it silent data loss. In
that case we override the deferral and inject anyway (one cache miss is
cheaper than dropped content).
Returns ``(should_inject, is_marker_override)``. ``is_marker_override`` is
True only when injection happens *because* of new markers despite a deferral,
so the caller can log the override distinctly.
"""
inject_tool = configured_inject_tool
if inject_tool and frozen_message_count > 0:
inject_tool = False # defer to preserve cache
is_marker_override = not inject_tool and has_compressed_content
return (inject_tool or is_marker_override), is_marker_override
def apply_session_sticky_ccr_tool(
*,
provider: Literal["anthropic", "openai", "google"],

View file

@ -49,6 +49,7 @@ import os
from dataclasses import dataclass
from typing import Any
from ..ccr.tool_injection import CCR_TOOL_NAME
from ..config import CCRConfig, TransformResult
from ..tokenizer import Tokenizer
from ..utils import compute_short_hash, create_tool_digest_marker, deep_copy_messages
@ -1008,6 +1009,13 @@ class SmartCrusher(Transform):
# OpenAI-style: top-level role=tool with string content.
if msg.get("role") == "tool":
# #1077: never re-compress headroom_retrieve results — they ARE
# already-retrieved CCR content; compressing them again creates an
# unresolvable retrieval loop.
# ponytail: ceiling is tool_call_id lookup; if the id is missing we
# compress (conservative: unknown tool names don't get a free pass).
if tool_names_by_id.get(msg.get("tool_call_id") or "") == CCR_TOOL_NAME:
continue
content = msg.get("content", "")
if isinstance(content, str):
tokens = tokenizer.count_text(content)
@ -1032,6 +1040,12 @@ class SmartCrusher(Transform):
for i, block in enumerate(content):
if not isinstance(block, dict) or block.get("type") != "tool_result":
continue
# #1077: skip headroom_retrieve results — compressing them
# would produce a new <<ccr:hash>> marker the agent cannot
# redeem (infinite retrieval loop).
# ponytail: ceiling is tool_use_id lookup; unknown ids pass through.
if tool_names_by_id.get(block.get("tool_use_id") or "") == CCR_TOOL_NAME:
continue
tool_content = block.get("content", "")
if not isinstance(tool_content, str):
continue

View file

@ -0,0 +1,137 @@
"""Regression test for #1006: the proxy must not emit unredeemable CCR markers.
When frozen_message_count > 0, the old code deferred headroom_retrieve tool
injection unconditionally even if compression just emitted NEW <<ccr:hash>>
markers the agent has no tool to redeem.
The fix: if new markers were emitted this turn, override the deferral and inject
the tool (one cache miss is acceptable; silent data loss is not). That decision
lives in ``should_inject_ccr_tool``, which the Anthropic handler calls; this test
pins the decision at that function so removing the override would fail here.
"""
from __future__ import annotations
from unittest.mock import MagicMock, patch
from headroom.ccr.tool_injection import CCR_TOOL_NAME, CCRToolInjector
from headroom.proxy.helpers import (
apply_session_sticky_ccr_tool,
should_inject_ccr_tool,
)
class TestShouldInjectCCRTool:
"""The decision the handler used to inline. This is where #1006 lived."""
def test_overrides_deferral_when_markers_emitted(self):
"""Frozen prefix would normally defer, but fresh markers force injection."""
should_inject, is_override = should_inject_ccr_tool(
configured_inject_tool=True,
frozen_message_count=3,
has_compressed_content=True,
)
assert should_inject, "must inject to keep markers redeemable (#1006)"
assert is_override, "this is the deferral override path"
def test_defers_when_no_markers(self):
"""Frozen prefix with no new markers stays deferred — no spurious tool."""
should_inject, is_override = should_inject_ccr_tool(
configured_inject_tool=True,
frozen_message_count=3,
has_compressed_content=False,
)
assert not should_inject
assert not is_override
def test_injects_normally_without_frozen_prefix(self):
"""No frozen prefix → inject as configured, not via the override path."""
should_inject, is_override = should_inject_ccr_tool(
configured_inject_tool=True,
frozen_message_count=0,
has_compressed_content=False,
)
assert should_inject
assert not is_override
class TestCCRInjectionEndToEnd:
"""The decision feeds apply_session_sticky_ccr_tool; assert the tool lands."""
def test_marker_in_frozen_prefix_yields_injected_tool(self):
# Injector detects a fresh marker, i.e. compression ran this turn.
injector = CCRToolInjector(provider="anthropic")
injector.scan_for_markers(
[
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_bash_x",
"content": "[50 items compressed to 5. Retrieve more: hash=abc123def456abc123def456]",
}
],
}
]
)
assert injector.has_compressed_content, "test setup: injector should detect marker"
# Drive the real decision the handler makes under a frozen prefix.
should_inject, _ = should_inject_ccr_tool(
configured_inject_tool=True,
frozen_message_count=3,
has_compressed_content=injector.has_compressed_content,
)
assert should_inject
with patch("headroom.proxy.helpers.get_session_ccr_tracker") as mock_tracker_fn:
mock_tracker = MagicMock()
mock_tracker.has_done_ccr.return_value = False # first CCR ever
mock_tracker.get_golden_tool_bytes.return_value = None
mock_tracker_fn.return_value = mock_tracker
tools_out, _was_injected = apply_session_sticky_ccr_tool(
provider="anthropic",
session_id="session-frozen-test",
request_id="req-test-1",
existing_tools=[],
has_compressed_content_this_turn=injector.has_compressed_content,
)
tool_names = [t.get("name") for t in tools_out]
assert CCR_TOOL_NAME in tool_names, (
f"headroom_retrieve not injected when markers emitted and prefix frozen (#1006). "
f"tools={tool_names}"
)
def test_no_marker_in_frozen_prefix_skips_tool(self):
injector = CCRToolInjector(provider="anthropic")
injector.scan_for_markers([{"role": "user", "content": "hello"}])
assert not injector.has_compressed_content, "test setup: no markers expected"
should_inject, _ = should_inject_ccr_tool(
configured_inject_tool=True,
frozen_message_count=3,
has_compressed_content=injector.has_compressed_content,
)
assert not should_inject, "no markers → no forced injection"
with patch("headroom.proxy.helpers.get_session_ccr_tracker") as mock_tracker_fn:
mock_tracker = MagicMock()
mock_tracker.has_done_ccr.return_value = False
mock_tracker.get_golden_tool_bytes.return_value = None
mock_tracker_fn.return_value = mock_tracker
tools_out, _was_injected = apply_session_sticky_ccr_tool(
provider="anthropic",
session_id="session-frozen-no-markers",
request_id="req-test-2",
existing_tools=[],
has_compressed_content_this_turn=False,
)
tool_names = [t.get("name") for t in tools_out]
assert CCR_TOOL_NAME not in tool_names, (
"headroom_retrieve should NOT be injected when no markers and frozen prefix"
)

View file

@ -0,0 +1,234 @@
"""Regression tests for #1077: SmartCrusher must not re-compress headroom_retrieve
tool results.
When the proxy's CCR path returns content via headroom_retrieve, the client sends
it back as a tool_result. Without the fix, SmartCrusher.apply() would compress
that content again new <<ccr:hash>> marker infinite retrieval loop.
"""
from __future__ import annotations
import json
import pytest
from headroom import OpenAIProvider, Tokenizer
from headroom.ccr.tool_injection import CCR_TOOL_NAME
from headroom.config import SmartCrusherConfig
def _build_extension() -> None:
try:
from headroom._core import SmartCrusher # noqa: F401
except ImportError:
pytest.skip(
"headroom._core not built — run `bash scripts/build_rust_extension.sh`",
allow_module_level=True,
)
_build_extension()
_provider = OpenAIProvider()
def _get_tokenizer(model: str = "gpt-4o") -> Tokenizer:
return Tokenizer(_provider.get_token_counter(model), model)
from headroom.transforms.smart_crusher import SmartCrusher # noqa: E402
def _make_crusher(min_tokens: int = 0) -> SmartCrusher:
return SmartCrusher(config=SmartCrusherConfig(min_tokens_to_crush=min_tokens))
def _big_content() -> str:
"""Return a JSON array large enough to trigger compression."""
return json.dumps([{"id": i, "value": "x" * 20} for i in range(60)])
class TestHeadroomRetrieveExemptionOpenAI:
"""OpenAI-style role=tool messages from headroom_retrieve must not be crushed."""
def test_retrieve_result_not_compressed(self):
"""headroom_retrieve tool result is skipped even when content is large."""
content = _big_content()
messages = [
# The assistant turn that called headroom_retrieve
{
"role": "assistant",
"tool_calls": [
{
"id": "call_abc",
"function": {"name": CCR_TOOL_NAME, "arguments": '{"hash":"abc"}'},
}
],
},
# The tool result — this must NOT be re-compressed
{"role": "tool", "tool_call_id": "call_abc", "content": content},
]
crusher = _make_crusher(min_tokens=0)
tokenizer = _get_tokenizer()
result = crusher.apply(messages, tokenizer)
# Content must be byte-for-byte unchanged
tool_msg = result.messages[1]
assert tool_msg["content"] == content, (
"headroom_retrieve tool result was re-compressed (infinite loop bug #1077)"
)
# No CCR transforms should have fired
assert not any("smart_crush" in t for t in result.transforms_applied)
def test_non_retrieve_tool_still_compressed(self):
"""Normal tool results are still compressed — exemption is narrow."""
content = _big_content()
messages = [
{
"role": "assistant",
"tool_calls": [
{
"id": "call_xyz",
"function": {"name": "Bash", "arguments": '{"cmd":"ls"}'},
}
],
},
{"role": "tool", "tool_call_id": "call_xyz", "content": content},
]
crusher = _make_crusher(min_tokens=0)
tokenizer = _get_tokenizer()
result = crusher.apply(messages, tokenizer)
tool_msg = result.messages[1]
# Content should have been modified (compressed)
assert tool_msg["content"] != content or result.tokens_after <= result.tokens_before
class TestHeadroomRetrieveExemptionAnthropic:
"""Anthropic-style tool_result content blocks from headroom_retrieve must not be crushed."""
def test_retrieve_result_not_compressed(self):
"""Anthropic tool_result block for headroom_retrieve is skipped."""
content = _big_content()
messages = [
# assistant turn with tool_use block calling headroom_retrieve
{
"role": "assistant",
"content": [
{
"type": "tool_use",
"id": "toolu_ccr_1",
"name": CCR_TOOL_NAME,
"input": {"hash": "abc123def456"},
}
],
},
# user turn with tool_result — must NOT be re-compressed
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_ccr_1",
"content": content,
}
],
},
]
crusher = _make_crusher(min_tokens=0)
tokenizer = _get_tokenizer()
result = crusher.apply(messages, tokenizer)
tool_result_block = result.messages[1]["content"][0]
assert tool_result_block["content"] == content, (
"headroom_retrieve Anthropic tool_result was re-compressed (#1077)"
)
assert not any("smart_crush" in t for t in result.transforms_applied)
def test_non_retrieve_anthropic_tool_still_compressed(self):
"""Normal Anthropic tool_result blocks are still compressed."""
content = _big_content()
messages = [
{
"role": "assistant",
"content": [
{
"type": "tool_use",
"id": "toolu_bash_1",
"name": "Bash",
"input": {"cmd": "ls"},
}
],
},
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_bash_1",
"content": content,
}
],
},
]
crusher = _make_crusher(min_tokens=0)
tokenizer = _get_tokenizer()
result = crusher.apply(messages, tokenizer)
# Either the content changed (compressed) or tokens went down
tool_result_block = result.messages[1]["content"][0]
assert (
tool_result_block["content"] != content or result.tokens_after <= result.tokens_before
)
def test_mixed_retrieve_and_normal_only_normal_compressed(self):
"""With two tool_results in one user turn, only the non-CCR one is compressed."""
content = _big_content()
messages = [
{
"role": "assistant",
"content": [
{
"type": "tool_use",
"id": "toolu_ccr_2",
"name": CCR_TOOL_NAME,
"input": {"hash": "abc123def456"},
},
{
"type": "tool_use",
"id": "toolu_bash_2",
"name": "Bash",
"input": {"cmd": "ls"},
},
],
},
{
"role": "user",
"content": [
{
"type": "tool_result",
"tool_use_id": "toolu_ccr_2",
"content": content, # must NOT be compressed
},
{
"type": "tool_result",
"tool_use_id": "toolu_bash_2",
"content": content, # MAY be compressed
},
],
},
]
crusher = _make_crusher(min_tokens=0)
tokenizer = _get_tokenizer()
result = crusher.apply(messages, tokenizer)
blocks = result.messages[1]["content"]
ccr_block = blocks[0]
bash_block = blocks[1]
# headroom_retrieve result must be untouched
assert ccr_block["content"] == content, (
"headroom_retrieve result was compressed — infinite loop bug #1077"
)
# The bash result should differ (or at least the crush count > 0)
assert bash_block["content"] != content or result.tokens_after < result.tokens_before