Extract tool schema savings policy (#1971)
## Description
Extracts the pure tool-schema savings attribution logic from `server.py`
into `headroom.proxy.tool_schema_savings_policy`. The server keeps the
`_tool_schema_saved_from_tags` compatibility alias used by the existing
stats payload path.
Closes #
## Type of Change
- [ ] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [x] Code refactoring (no functional changes)
## Changes Made
- Added `tool_schema_savings_policy.py` with stable savings tag names
and pure summation behavior.
- Replaced the inline `server.py` helper body with a compatibility alias
to the extracted policy.
- Added direct tests for valid tag summing, invalid values, non-mapping
input, and stable tag names.
- Carried forward the LiteLLM callback compatibility shim needed for
current mypy on `main`.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
### Test Output
```text
python -m pytest tests\test_tool_schema_savings_policy.py
4 passed in 0.14s
python -m ruff check .
All checks passed!
python -m ruff format --check .
1095 files already formatted
python -m mypy headroom --ignore-missing-imports
Success: no issues found in 409 source files
gitleaks protect --staged --no-banner --redact
no leaks found
```
## Real Behavior Proof
- Environment: Windows, Python 3.13.13, branch
`jd/architecture-slice-24`.
- Exact command / steps: ran focused tool-schema savings policy tests,
ruff, ruff format check, mypy, and staged gitleaks scan.
- Observed result: pure policy behavior is directly covered and local
lint/type/security checks pass.
- Not tested: full proxy runtime; this slice only moves pure stats
attribution logic while preserving the server alias.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable
## Screenshots (if applicable)
N/A
## Additional Notes
Documentation and changelog updates are N/A for this internal
architecture-only refactor. The push reported existing default-branch
Dependabot alerts; no staged secret leaks were found for this PR.
2026-07-11 15:10:21 +00:00
|
|
|
from __future__ import annotations
|
|
|
|
|
|
|
|
|
|
from headroom.proxy.tool_schema_savings_policy import (
|
|
|
|
|
TOOL_SCHEMA_SAVINGS_TAGS,
|
fix(stats): report one "Tokens Saved" headline across every harness (#2737)
## Description
The "tokens saved" figure a user sees depended on which harness they
ran. Headroom saves tool-definition tokens in two accounting shapes,
both legitimate, but the rule was never written down — so two harnesses
silently dropped savings and three surfaces open-coded the sum
differently.
- **Compaction** rewrites the tool array, so both endpoints are
countable → handlers fold the delta into
`original_tokens`/`optimized_tokens`, keeping `tok_before - tok_after ==
tok_saved` coherent.
- **Deferral / hook shrink** removes schemas `count_messages` never sees
→ can only be recorded as a tag, additive to `tokens_saved`.
`tool_schema_savings_policy` now owns the sum via
`headline_tokens_saved()`, and every reporting surface routes through
it.
Closes #
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
## Changes Made
Producer gaps (both Anthropic — i.e. Claude Code, the primary harness):
- `anthropic:tool_schema_compaction` / `anthropic:tool_desc_compaction`
computed their savings, debug-logged them, and **discarded them**. Now
folded at the final recount, mirroring the OpenAI chat handler. A
14-tool array drops 786 tokens that previously reported `tok_saved=0`.
- Anthropic never wrote `turn_hook_tools_saved_tokens` at all, so a
turn-hook extension that shrinks tools got zero credit there while
OpenAI credited it. Now tagged.
Reporting gaps:
- `headroom perf` printed `Total saved (messages)` and `Tool saved` as
rival lines — on a tool-heavy session the headline read `0` and the real
win looked like a footnote. Now one `Tokens saved:` headline with a
messages/tool-schemas breakdown.
- `active_savings_percent` divided a **compression-only numerator** by a
denominator that already included compacted tool schema, undercounting
every tool-heavy session. Numerator is now all-layers, with deferred
schemas added to both sides.
- The headline and its percent now share a numerator. Previously the
dashboard tile showed an all-layers total next to a compression-only
percent.
- Session summary and dashboard tile relabelled to `Tokens Saved`; the
tool-schema panel is labelled as a component (`Tokens Saved · Tool
Schemas`) rather than a rival metric.
- `outcome.py` had two drifted inline copies of the tag sum; both now
call the policy module that exists for it. `total_saved=` added to the
PERF line.
- JSON: added `total_tokens_saved` / `total_savings_pct`; existing
`tokens_saved` / `tool_saved` / `savings_pct` keys unchanged for
back-compat.
Not changed by design: the Codex per-component attribution sub-line
would need a 9th positional tuple element threaded through 4 unpack
sites, and its headline is already correct without it.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ .venv/bin/ruff check headroom/ tests/test_tool_schema_savings_policy.py --exclude headroom/dashboard/templates
All checks passed!
$ .venv/bin/ruff format --check headroom/ tests/... --exclude headroom/dashboard/templates
510 files already formatted
$ .venv/bin/mypy headroom/
Success: no issues found in 508 source files
$ python -m pytest tests/test_tool_schema_savings_policy.py tests/test_cli_perf_format.py \
tests/test_request_outcome.py tests/test_savings_tool_search_aggregation.py \
tests/test_dashboard_token_savings.py tests/test_anthropic_compaction_transforms.py -q
70 passed, 1 warning in 6.15s
$ python -m pytest tests/test_handler_outcome_tag_invariant.py tests/test_cold_start_fast_pass.py \
tests/test_anthropic_ccr_workspace_unbound.py tests/test_anthropic_pre_upstream_backpressure.py \
tests/test_vertex_claude_compression.py tests/test_provider_route_specs.py -q
50 passed in 10.11s
$ python -m pytest tests/test_agent_savings.py tests/test_bundled_tools_savings.py \
tests/test_codex_ws_savings_deferral.py tests/test_savings_ledger_before_forwarded.py \
tests/test_savings_ledger_offload.py tests/test_proxy_savings_history.py \
tests/test_proxy_dashboard_stats_cache.py tests/test_output_savings_cli.py -q
97 passed, 2 skipped in 18.60s
$ python -m pytest tests/test_tool_schema_compaction.py tests/test_openai_responses_context_compaction.py \
tests/test_proxy_openai_cache_stability.py tests/test_codex_ws_compression_scheduler.py \
tests/test_proxy_streaming_request_logger.py -q
66 passed, 1 skipped in 16.69s
```
## Real Behavior Proof
- **Environment:** macOS 26.4 arm64, Python 3.12.6, repo `.venv`.
Motivated by a real user proxy log (0.33.0, `client=opencode` →
nano-gpt, 722 requests) reporting 0.12% savings.
- **Exact command / steps (1) — the Anthropic fold, real compaction +
real provider tokenizer:**
```python
tok = AnthropicProvider().get_token_counter("claude-sonnet-4-6")
payload = {"tools": [ ...14 tools with $schema/title/examples... ]}
body, modified, bb, ba = compact_tools(payload)
```
**Observed:**
```text
modified=True bytes 4503->2539 TOKENS 1650->864 delta=786
tok_before=6650 tok_after=5864 tok_saved=786 coherent=True
pre-fix: Claude Code reported tok_saved=0 and discarded 786 tokens
```
Pinned as `test_tool_schema_compaction_saves_real_tokens_not_just_bytes`
— it asserts a positive **token** delta (not just bytes), which is the
premise of folding at all.
- **Exact command / steps (2) — the report, on the reported session's
shape** (tool schemas carry the win, message compression is 0 because
everything routed to `excluded_tool`):
**Observed after:**
```text
Requests: 2
Tokens: 45,760 -> 45,760 (0.0% messages)
Tokens saved: 811 (1.7% reduction)
· messages 0
· tool schemas 811
JSON: {'total_tokens_saved': 811, 'total_savings_pct': 1.7, 'tokens_saved': 0,
'tool_saved': 811, 'savings_pct': 0.0}
```
Before, the same input printed `Total saved: 0 tokens (messages)` as the
headline with `Tool saved: 811` beneath it.
- **Not tested:** no live proxy run against a real provider — the
Anthropic fold is proven at the accounting layer (real `compact_tools` +
real provider tokenizer) and via the existing handler suites, not by an
end-to-end Claude Code session. Dashboard changes are template-label
edits verified by reading `stats.tokens.saved` / `by_layer.tool_search`
shapes, not by a browser screenshot.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md`
🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-03 09:03:54 -07:00
|
|
|
headline_tokens_saved,
|
Extract tool schema savings policy (#1971)
## Description
Extracts the pure tool-schema savings attribution logic from `server.py`
into `headroom.proxy.tool_schema_savings_policy`. The server keeps the
`_tool_schema_saved_from_tags` compatibility alias used by the existing
stats payload path.
Closes #
## Type of Change
- [ ] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [x] Code refactoring (no functional changes)
## Changes Made
- Added `tool_schema_savings_policy.py` with stable savings tag names
and pure summation behavior.
- Replaced the inline `server.py` helper body with a compatibility alias
to the extracted policy.
- Added direct tests for valid tag summing, invalid values, non-mapping
input, and stable tag names.
- Carried forward the LiteLLM callback compatibility shim needed for
current mypy on `main`.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
### Test Output
```text
python -m pytest tests\test_tool_schema_savings_policy.py
4 passed in 0.14s
python -m ruff check .
All checks passed!
python -m ruff format --check .
1095 files already formatted
python -m mypy headroom --ignore-missing-imports
Success: no issues found in 409 source files
gitleaks protect --staged --no-banner --redact
no leaks found
```
## Real Behavior Proof
- Environment: Windows, Python 3.13.13, branch
`jd/architecture-slice-24`.
- Exact command / steps: ran focused tool-schema savings policy tests,
ruff, ruff format check, mypy, and staged gitleaks scan.
- Observed result: pure policy behavior is directly covered and local
lint/type/security checks pass.
- Not tested: full proxy runtime; this slice only moves pure stats
attribution logic while preserving the server alias.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable
## Screenshots (if applicable)
N/A
## Additional Notes
Documentation and changelog updates are N/A for this internal
architecture-only refactor. The push reported existing default-branch
Dependabot alerts; no staged secret leaks were found for this PR.
2026-07-11 15:10:21 +00:00
|
|
|
tool_schema_saved_from_tags,
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_tool_schema_saved_from_tags_sums_headroom_deferral_tags() -> None:
|
|
|
|
|
assert (
|
|
|
|
|
tool_schema_saved_from_tags(
|
|
|
|
|
{
|
|
|
|
|
"tool_search_deferred_tokens": "120",
|
|
|
|
|
"turn_hook_tools_saved_tokens": 30,
|
|
|
|
|
"unrelated": 999,
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
== 150
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_tool_schema_saved_from_tags_ignores_invalid_values() -> None:
|
|
|
|
|
assert (
|
|
|
|
|
tool_schema_saved_from_tags(
|
|
|
|
|
{
|
|
|
|
|
"tool_search_deferred_tokens": "not-an-int",
|
|
|
|
|
"turn_hook_tools_saved_tokens": None,
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
== 0
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_tool_schema_saved_from_tags_rejects_non_mapping_tags() -> None:
|
|
|
|
|
assert tool_schema_saved_from_tags(None) == 0
|
|
|
|
|
assert tool_schema_saved_from_tags([("tool_search_deferred_tokens", 10)]) == 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_tool_schema_savings_tags_are_stable() -> None:
|
|
|
|
|
assert TOOL_SCHEMA_SAVINGS_TAGS == (
|
|
|
|
|
"tool_search_deferred_tokens",
|
|
|
|
|
"turn_hook_tools_saved_tokens",
|
|
|
|
|
)
|
fix(stats): report one "Tokens Saved" headline across every harness (#2737)
## Description
The "tokens saved" figure a user sees depended on which harness they
ran. Headroom saves tool-definition tokens in two accounting shapes,
both legitimate, but the rule was never written down — so two harnesses
silently dropped savings and three surfaces open-coded the sum
differently.
- **Compaction** rewrites the tool array, so both endpoints are
countable → handlers fold the delta into
`original_tokens`/`optimized_tokens`, keeping `tok_before - tok_after ==
tok_saved` coherent.
- **Deferral / hook shrink** removes schemas `count_messages` never sees
→ can only be recorded as a tag, additive to `tokens_saved`.
`tool_schema_savings_policy` now owns the sum via
`headline_tokens_saved()`, and every reporting surface routes through
it.
Closes #
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
## Changes Made
Producer gaps (both Anthropic — i.e. Claude Code, the primary harness):
- `anthropic:tool_schema_compaction` / `anthropic:tool_desc_compaction`
computed their savings, debug-logged them, and **discarded them**. Now
folded at the final recount, mirroring the OpenAI chat handler. A
14-tool array drops 786 tokens that previously reported `tok_saved=0`.
- Anthropic never wrote `turn_hook_tools_saved_tokens` at all, so a
turn-hook extension that shrinks tools got zero credit there while
OpenAI credited it. Now tagged.
Reporting gaps:
- `headroom perf` printed `Total saved (messages)` and `Tool saved` as
rival lines — on a tool-heavy session the headline read `0` and the real
win looked like a footnote. Now one `Tokens saved:` headline with a
messages/tool-schemas breakdown.
- `active_savings_percent` divided a **compression-only numerator** by a
denominator that already included compacted tool schema, undercounting
every tool-heavy session. Numerator is now all-layers, with deferred
schemas added to both sides.
- The headline and its percent now share a numerator. Previously the
dashboard tile showed an all-layers total next to a compression-only
percent.
- Session summary and dashboard tile relabelled to `Tokens Saved`; the
tool-schema panel is labelled as a component (`Tokens Saved · Tool
Schemas`) rather than a rival metric.
- `outcome.py` had two drifted inline copies of the tag sum; both now
call the policy module that exists for it. `total_saved=` added to the
PERF line.
- JSON: added `total_tokens_saved` / `total_savings_pct`; existing
`tokens_saved` / `tool_saved` / `savings_pct` keys unchanged for
back-compat.
Not changed by design: the Codex per-component attribution sub-line
would need a 9th positional tuple element threaded through 4 unpack
sites, and its headline is already correct without it.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ .venv/bin/ruff check headroom/ tests/test_tool_schema_savings_policy.py --exclude headroom/dashboard/templates
All checks passed!
$ .venv/bin/ruff format --check headroom/ tests/... --exclude headroom/dashboard/templates
510 files already formatted
$ .venv/bin/mypy headroom/
Success: no issues found in 508 source files
$ python -m pytest tests/test_tool_schema_savings_policy.py tests/test_cli_perf_format.py \
tests/test_request_outcome.py tests/test_savings_tool_search_aggregation.py \
tests/test_dashboard_token_savings.py tests/test_anthropic_compaction_transforms.py -q
70 passed, 1 warning in 6.15s
$ python -m pytest tests/test_handler_outcome_tag_invariant.py tests/test_cold_start_fast_pass.py \
tests/test_anthropic_ccr_workspace_unbound.py tests/test_anthropic_pre_upstream_backpressure.py \
tests/test_vertex_claude_compression.py tests/test_provider_route_specs.py -q
50 passed in 10.11s
$ python -m pytest tests/test_agent_savings.py tests/test_bundled_tools_savings.py \
tests/test_codex_ws_savings_deferral.py tests/test_savings_ledger_before_forwarded.py \
tests/test_savings_ledger_offload.py tests/test_proxy_savings_history.py \
tests/test_proxy_dashboard_stats_cache.py tests/test_output_savings_cli.py -q
97 passed, 2 skipped in 18.60s
$ python -m pytest tests/test_tool_schema_compaction.py tests/test_openai_responses_context_compaction.py \
tests/test_proxy_openai_cache_stability.py tests/test_codex_ws_compression_scheduler.py \
tests/test_proxy_streaming_request_logger.py -q
66 passed, 1 skipped in 16.69s
```
## Real Behavior Proof
- **Environment:** macOS 26.4 arm64, Python 3.12.6, repo `.venv`.
Motivated by a real user proxy log (0.33.0, `client=opencode` →
nano-gpt, 722 requests) reporting 0.12% savings.
- **Exact command / steps (1) — the Anthropic fold, real compaction +
real provider tokenizer:**
```python
tok = AnthropicProvider().get_token_counter("claude-sonnet-4-6")
payload = {"tools": [ ...14 tools with $schema/title/examples... ]}
body, modified, bb, ba = compact_tools(payload)
```
**Observed:**
```text
modified=True bytes 4503->2539 TOKENS 1650->864 delta=786
tok_before=6650 tok_after=5864 tok_saved=786 coherent=True
pre-fix: Claude Code reported tok_saved=0 and discarded 786 tokens
```
Pinned as `test_tool_schema_compaction_saves_real_tokens_not_just_bytes`
— it asserts a positive **token** delta (not just bytes), which is the
premise of folding at all.
- **Exact command / steps (2) — the report, on the reported session's
shape** (tool schemas carry the win, message compression is 0 because
everything routed to `excluded_tool`):
**Observed after:**
```text
Requests: 2
Tokens: 45,760 -> 45,760 (0.0% messages)
Tokens saved: 811 (1.7% reduction)
· messages 0
· tool schemas 811
JSON: {'total_tokens_saved': 811, 'total_savings_pct': 1.7, 'tokens_saved': 0,
'tool_saved': 811, 'savings_pct': 0.0}
```
Before, the same input printed `Total saved: 0 tokens (messages)` as the
headline with `Tool saved: 811` beneath it.
- **Not tested:** no live proxy run against a real provider — the
Anthropic fold is proven at the accounting layer (real `compact_tools` +
real provider tokenizer) and via the existing handler suites, not by an
end-to-end Claude Code session. Dashboard changes are template-label
edits verified by reading `stats.tokens.saved` / `by_layer.tool_search`
shapes, not by a browser screenshot.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md`
🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-03 09:03:54 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
# ── headline_tokens_saved: the one figure every surface reports ────────────────
|
|
|
|
|
# Headroom saves tool-definition tokens in two accounting shapes — compaction
|
|
|
|
|
# folds into tokens_saved, deferral is tagged and additive. Both existed before
|
|
|
|
|
# but the rule was never written down, so two harnesses dropped their compaction
|
|
|
|
|
# savings and three surfaces open-coded the sum. These cases pin the contract.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_folded_compaction_is_not_counted_twice() -> None:
|
|
|
|
|
"""Compaction is ALREADY inside tokens_saved (handlers fold both endpoints).
|
|
|
|
|
|
|
|
|
|
Adding an attribution amount back on top would inflate every tool-heavy turn.
|
|
|
|
|
"""
|
|
|
|
|
assert headline_tokens_saved(420, {}) == 420
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_deferral_tags_are_additive_to_tokens_saved() -> None:
|
|
|
|
|
"""Deferral removes schemas count_messages never saw, so it can't be folded."""
|
|
|
|
|
tags = {"tool_search_deferred_tokens": 9639}
|
|
|
|
|
assert headline_tokens_saved(0, tags) == 9639
|
|
|
|
|
assert headline_tokens_saved(1_000, tags) == 10_639
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_headline_identical_across_harnesses_for_equivalent_work() -> None:
|
|
|
|
|
"""A 500-token saving reports as 500 whichever accounting shape produced it.
|
|
|
|
|
|
|
|
|
|
Anthropic/Claude Code folds its compaction; a Codex deferral is tagged. Same
|
|
|
|
|
real saving, same headline — that equivalence is the point of the helper.
|
|
|
|
|
"""
|
|
|
|
|
anthropic_folded = headline_tokens_saved(500, {})
|
|
|
|
|
codex_tagged = headline_tokens_saved(0, {"tool_search_deferred_tokens": 500})
|
|
|
|
|
assert anthropic_folded == codex_tagged == 500
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_headline_survives_malformed_tags() -> None:
|
|
|
|
|
for tags in (None, {}, "not-a-dict", {"tool_search_deferred_tokens": None}):
|
|
|
|
|
assert headline_tokens_saved(10, tags) == 10
|
|
|
|
|
assert headline_tokens_saved(10, {"tool_search_deferred_tokens": "abc"}) == 10
|
|
|
|
|
assert headline_tokens_saved(None, None) == 0
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_headline_clamps_negative_message_savings() -> None:
|
|
|
|
|
"""Handlers revert inflation before forwarding, so a negative is a count artifact."""
|
|
|
|
|
assert headline_tokens_saved(-5, {}) == 0
|
|
|
|
|
assert headline_tokens_saved(-5, {"tool_search_deferred_tokens": 100}) == 95
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_tool_schema_compaction_saves_real_tokens_not_just_bytes() -> None:
|
|
|
|
|
"""The premise of folding compaction into tokens_saved on every handler.
|
|
|
|
|
|
|
|
|
|
Compaction strips annotation keys ($schema/title/examples). If that only moved
|
|
|
|
|
bytes that tokenize to nothing, the fold would be worthless — so pin a positive
|
|
|
|
|
TOKEN delta on a realistically-shaped tool array, and pin that folding it into
|
|
|
|
|
both endpoints keeps ``tok_before - tok_after == tok_saved`` coherent.
|
|
|
|
|
"""
|
|
|
|
|
import json
|
|
|
|
|
|
|
|
|
|
from headroom.providers.anthropic import AnthropicProvider
|
|
|
|
|
from headroom.proxy.tool_schema_compaction import compact_tools
|
|
|
|
|
|
|
|
|
|
tok = AnthropicProvider().get_token_counter("claude-sonnet-4-6")
|
|
|
|
|
payload = {
|
|
|
|
|
"tools": [
|
|
|
|
|
{
|
|
|
|
|
"name": f"tool_{i}",
|
|
|
|
|
"description": "Does a thing.\n\n Returns text.",
|
|
|
|
|
"input_schema": {
|
|
|
|
|
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
|
|
|
|
"title": f"tool_{i}_schema",
|
|
|
|
|
"examples": [{"path": "/tmp/x"}, {"path": "/tmp/y"}],
|
|
|
|
|
"type": "object",
|
|
|
|
|
"properties": {"path": {"type": "string", "description": "File path"}},
|
|
|
|
|
"required": ["path"],
|
|
|
|
|
},
|
|
|
|
|
}
|
|
|
|
|
for i in range(14)
|
|
|
|
|
]
|
|
|
|
|
}
|
|
|
|
|
before_tools = payload["tools"]
|
|
|
|
|
body, modified, _bytes_before, _bytes_after = compact_tools(payload)
|
|
|
|
|
assert modified is True
|
|
|
|
|
|
|
|
|
|
tool_before = tok.count_text(json.dumps(before_tools, default=str))
|
|
|
|
|
tool_after = tok.count_text(json.dumps(body["tools"], default=str))
|
|
|
|
|
assert tool_after < tool_before, "compaction must shrink tool TOKENS, not only bytes"
|
|
|
|
|
|
|
|
|
|
# Mirrors the fold each handler applies at its final recount, with zero message
|
|
|
|
|
# compression — the shape that used to report tok_saved=0 on Claude Code.
|
|
|
|
|
original_tokens = optimized_tokens = 5_000
|
|
|
|
|
if 0 < tool_after < tool_before:
|
|
|
|
|
original_tokens += tool_before
|
|
|
|
|
optimized_tokens += tool_after
|
|
|
|
|
tokens_saved = max(0, original_tokens - optimized_tokens)
|
|
|
|
|
|
|
|
|
|
assert tokens_saved == tool_before - tool_after
|
|
|
|
|
assert original_tokens - optimized_tokens == tokens_saved
|
|
|
|
|
assert headline_tokens_saved(tokens_saved, {}) == tokens_saved
|