headroom/tests/test_openai_chat_tool_desc_compaction.py
Tejas Chopra f9db5b5060
fix(proxy/openai): run tool-description compaction on chat-completions (#2741)
## Description

`HEADROOM_TOOL_DESC_MAX_CHARS` was wired into the Anthropic handler and
the Responses (Codex) handler, but never into **chat-completions** — so
the env var was a silent no-op for every chat client: opencode, Cline,
Aider, Roo, anything routed through LiteLLM.

Tool descriptions live on the `tools` array, which the message pipeline
never inspects, so no other pass was covering them.

Closes #

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)

## Changes Made

- Run the L2 tool-description pass on the chat-completions path,
mirroring the block the Anthropic and Responses handlers already had.
- `compact_tool_descriptions` already walks both wire shapes — nested
`{"function": {"description": ...}}` for chat, flat for Responses — so
this is wiring, not a new codec.
- Chains after the existing schema compaction, seeding the token
"before" count only when that pass didn't, so the two compose instead of
double-counting.
- Labelled `openai:chat:tool_desc_compaction`, distinct from the
Anthropic and Responses labels so `headroom perf --by-transform` can
attribute it.
- Still opt-in and off by default: an unset env var leaves the tools
array — and therefore its cache prefix — byte-identical.

### Scope note: two adjacent "gaps" that turned out not to be

While surveying handler parity I flagged three missing chat-completions
transforms. Only one was real; recording the other two so nobody
re-opens them:

- **`tool_search_deferral` — correctly absent.** `{"type":
"tool_search"}` and `defer_loading` are Responses-API constructs, and
`_model_supports_openai_tool_search` gates them to `gpt-5.4+`. Injecting
that shape into a chat-completions request would be invalid, not an
improvement.
- **`system_prompt_compaction` — not applicable.** Anthropic needs a
dedicated pass because `system` is an out-of-band top-level field the
message pipeline never sees. On chat-completions the system prompt *is*
`messages[0]`, so it already reaches ContentRouter and is governed by
the existing `compress_system_messages` / `skip_system` gate. Wiring a
second path there would change system-prefix cache behavior for no new
coverage.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ .venv/bin/ruff check headroom/ tests/test_openai_chat_tool_desc_compaction.py --exclude headroom/dashboard/templates
All checks passed!

$ .venv/bin/mypy headroom/
Success: no issues found in 508 source files

$ python -m pytest tests/test_openai_chat_tool_desc_compaction.py tests/test_tool_schema_compaction.py \
    tests/test_proxy_openai_cache_stability.py tests/test_openai_responses_context_compaction.py -q
49 passed in 16.59s
```

## Real Behavior Proof

- **Environment:** macOS 26.4 arm64, Python 3.12.6, repo `.venv`.
- **Exact command / steps:** ran `compact_tool_descriptions` at
`HEADROOM_TOOL_DESC_MAX_CHARS=30` against both wire shapes with the same
tool (a `read` tool with an 86-char description and a described `path`
param).
- **Observed result:**

```text
chat-completions (nested)    modified=True bytes 272->215
responses (flat)             modified=True bytes 259->202
```

Chat previously reported `modified=False` from the handler because the
pass was never invoked at all.

- **Not tested:** no live chat-completions request against a real
provider — the handler block is a thin adapter over
`compact_tool_descriptions`, and the regression was a missing *call*,
which the wiring test catches at source level. A full end-to-end drive
would need an upstream endpoint.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md`

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-03 10:52:05 -07:00

126 lines
4.4 KiB
Python

"""Tool-description compaction must run on chat-completions, not just Anthropic/Responses.
``HEADROOM_TOOL_DESC_MAX_CHARS`` was wired into the Anthropic handler and the
Responses (Codex) handler but never into chat-completions, so the env var was a
silent no-op for every chat client — opencode, Cline, Aider, Roo, anything routed
through LiteLLM. Tool descriptions live on the ``tools`` array, which the message
pipeline never inspects, so no other pass was covering them.
These tests pin the shape handling and the opt-in gate rather than driving the
whole handler: the handler block is a thin adapter over
``compact_tool_descriptions``, and the thing that actually broke was that nobody
called it with the chat-shaped payload.
"""
from __future__ import annotations
import pytest
import headroom.proxy.tool_schema_compaction as tsc
from headroom.proxy.tool_schema_compaction import compact_tool_descriptions, tool_desc_max_chars
_LONG_DESC = "Reads a file from disk and returns the full text content as a string, with numbers."
@pytest.fixture(autouse=True)
def _reset_desc_cache():
"""The max-chars lookup is process-cached; clear it around each test."""
tsc._TOOL_DESC_MAX_CHARS = None
yield
tsc._TOOL_DESC_MAX_CHARS = None
def _chat_tools() -> list[dict]:
"""chat-completions shape: name/description nested under "function"."""
return [
{
"type": "function",
"function": {
"name": "read",
"description": _LONG_DESC,
"parameters": {
"type": "object",
"properties": {"path": {"type": "string", "description": "File path to read"}},
},
},
}
]
def _responses_tools() -> list[dict]:
"""Responses shape: name/description flat on the tool."""
return [
{
"type": "function",
"name": "read",
"description": _LONG_DESC,
"parameters": {
"type": "object",
"properties": {"path": {"type": "string", "description": "File path to read"}},
},
}
]
def test_compacts_the_nested_chat_completions_tool_shape(monkeypatch):
"""The shape the chat handler passes — the one that was never being compacted."""
monkeypatch.setenv("HEADROOM_TOOL_DESC_MAX_CHARS", "30")
payload, modified, before, after = compact_tool_descriptions(
{"tools": _chat_tools()}, tool_desc_max_chars()
)
assert modified is True
assert after < before
desc = payload["tools"][0]["function"]["description"]
assert len(desc) <= len(_LONG_DESC)
assert desc != _LONG_DESC
def test_both_wire_shapes_are_handled(monkeypatch):
"""One helper serves both handlers, so chat needed wiring — not a new codec."""
monkeypatch.setenv("HEADROOM_TOOL_DESC_MAX_CHARS", "30")
max_chars = tool_desc_max_chars()
_, chat_modified, chat_before, chat_after = compact_tool_descriptions(
{"tools": _chat_tools()}, max_chars
)
_, resp_modified, resp_before, resp_after = compact_tool_descriptions(
{"tools": _responses_tools()}, max_chars
)
assert chat_modified is resp_modified is True
assert chat_after < chat_before
assert resp_after < resp_before
def test_disabled_by_default_leaves_tools_untouched():
"""Opt-in only: an unset env var must not perturb the tools prefix or its cache."""
tools = _chat_tools()
assert tool_desc_max_chars() == 0
payload, modified, before, after = compact_tool_descriptions(
{"tools": tools}, tool_desc_max_chars()
)
assert modified is False
assert payload["tools"] == tools
assert (before, after) == (0, 0)
def test_chat_handler_calls_the_desc_pass(monkeypatch):
"""Guard the wiring itself: the handler source must invoke the L2 pass.
ponytail: source-level check, not a live handler drive — spinning the full
chat-completions path needs an upstream, and the regression here was a missing
CALL, which is exactly what this catches.
"""
import inspect
from headroom.proxy.handlers import openai as openai_handler
source = inspect.getsource(openai_handler)
assert "openai:chat:tool_desc_compaction" in source
# The Anthropic and Responses handlers already had their own labels; make sure
# the chat one is distinct so `headroom perf --by-transform` can attribute it.
assert "openai:responses:tool_desc_compaction" in source