headroom/tests/test_proxy_openai_cache_stability.py
JD Davis 55efb1c77d
fix(proxy): keep OpenAI tool observations mutable in cache mode (#1884)
## Description

Diagnoses and fixes the low-savings OpenAI-compatible cache-mode path
reported in #1696.

OpenAI-compatible tool-calling clients can end a turn with `role:
"tool"` (or legacy `role: "function"`) rather than `role: "user"`. The
OpenAI chat handler's cache-mode freeze boundary treated those tails as
non-mutable, and because `HeadroomProxy` resolves
`_strict_previous_turn_frozen_count` from the Anthropic mixin first, the
OpenAI-specific helper was not used in production. That froze the entire
conversation before `ContentRouter` ran, leaving no live tool
observation to compress and producing near-pass-through savings on long
coding sessions.

This PR keeps final OpenAI tool/function observations mutable in cache
mode, explicitly calls the OpenAI helper to avoid the mixin-name
collision, and clamps negative token-savings artifacts at the
metrics/cost aggregation boundary so stats cannot under-report actual
forwarded savings.

Closes #1696

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- Treat final OpenAI `user`, `tool`, and `function` messages as the
mutable cache-mode live zone.
- Route OpenAI cache-boundary calls through
`OpenAIHandlerMixin._strict_previous_turn_frozen_count` explicitly so
the Anthropic mixin method cannot shadow it in `HeadroomProxy`'s MRO.
- Preserve cache-mode live-tail boundaries even when compression-cache
state would otherwise freeze the whole request.
- Clamp negative `tokens_saved` artifacts in `CostTracker.record_tokens`
and `PrometheusMetrics.record_request`.
- Add regression coverage for OpenAI final `tool`/`function` tails,
over-frozen tracker state, and non-negative savings aggregation.

## Testing

- [ ] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ maturin build --profile ci --out dist --interpreter python
Built wheel for abi3 Python >= 3.10 to dist\headroom_ai-0.29.0-cp310-abi3-win_amd64.whl

$ python -m pytest tests\test_proxy_handler_helpers.py tests\test_proxy_openai_cache_stability.py tests\test_observability_metrics.py tests\test_cost_tracker_counterfactual.py
49 passed in 10.27s

$ python -m ruff check .
All checks passed!

$ python -m mypy headroom
Success: no issues found in 407 source files

$ python -m pytest
53 failed, 7703 passed, 488 skipped, 5893 warnings, 131 errors in 595.18s (0:09:55)
```

Full-suite note: the full local `pytest` run was attempted on
Windows/Python 3.13 after building `headroom._core`. It did not complete
green due to broad pre-existing/local-environment failures outside this
change area, dominated by SQLite/memory persistence permission/path
errors plus unrelated adapter/cache/tool tests. The focused regression
suite for this PR passes, and repo-level lint/type gates pass.

## Real Behavior Proof

- Environment: Windows, Python 3.13.13, Rust/Cargo available, local
`headroom._core` wheel built with `maturin build --profile ci`.
- Exact command / steps: ran the OpenAI cache-stability tests with final
`role: "tool"` and `role: "function"` chat tails.
- Observed result:
`test_openai_cache_mode_keeps_final_tool_observation_mutable[tool]` and
`[function]` pass, proving the pipeline receives `frozen_message_count
== 2` for a 3-message request instead of freezing all 3 messages.
- Not tested: live Lemonade/KiloCode upstream session; no local Lemonade
Server was available.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [ ] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable

## Screenshots (if applicable)

N/A

## Additional Notes

Docs and CHANGELOG are N/A for this narrow proxy bug fix. The broad
local `pytest` checkbox is intentionally left unchecked because the full
suite had unrelated local-environment failures; see the test output
above. Focused regression tests, `ruff check .`, and `mypy headroom` are
green.
2026-07-09 07:51:01 -07:00

371 lines
13 KiB
Python

"""Regression tests for OpenAI cache-mode stability in proxy mode."""
from __future__ import annotations
from types import SimpleNamespace
import httpx
import pytest
pytest.importorskip("fastapi")
from fastapi.testclient import TestClient
from headroom.proxy.server import ProxyConfig, create_app
class _FakePrefixTracker:
def __init__(self, frozen_count: int):
self._frozen_count = frozen_count
def get_frozen_message_count(self) -> int:
return self._frozen_count
# Empty history → overlay_cached_prefix() is a no-op here, so these tests
# keep asserting the cache-freeze behavior they always have. The cross-turn
# overlay itself is exercised in test_cross_turn_cache_safety.py against the
# real tracker; these stubs just satisfy the handler's overlay call.
def get_last_original_messages(self): # noqa: ANN201
return []
def get_last_forwarded_messages(self): # noqa: ANN201
return []
def update_from_response(self, **kwargs): # noqa: ANN003
return None
def _make_proxy_client() -> TestClient:
config = ProxyConfig(
optimize=False,
cache_enabled=False,
rate_limit_enabled=False,
cost_tracking_enabled=False,
log_requests=False,
ccr_inject_tool=False,
ccr_handle_responses=False,
ccr_context_tracking=False,
image_optimize=False,
)
app = create_app(config)
return TestClient(app)
def test_openai_cache_mode_freezes_previous_turns() -> None:
captured = {}
with _make_proxy_client() as client:
proxy = client.app.state.proxy
proxy.config.optimize = True
proxy.config.mode = "cache"
fake_tracker = _FakePrefixTracker(frozen_count=0)
proxy.session_tracker_store.compute_session_id = lambda request, model, messages: (
"stable-session"
)
proxy.session_tracker_store.get_or_create = lambda session_id, provider: fake_tracker
def _fake_apply(**kwargs):
captured["frozen_message_count"] = kwargs.get("frozen_message_count")
return SimpleNamespace(
messages=kwargs["messages"],
transforms_applied=[],
timing={},
tokens_before=60,
tokens_after=60,
waste_signals=None,
)
proxy.openai_pipeline.apply = _fake_apply
async def _fake_retry(method, url, headers, body, stream=False, **kwargs): # noqa: ANN001
return httpx.Response(
200,
json={
"id": "chatcmpl_1",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "ok"},
"finish_reason": "stop",
}
],
"usage": {"prompt_tokens": 60, "completion_tokens": 3, "total_tokens": 63},
},
)
proxy._retry_request = _fake_retry
response = client.post(
"/v1/chat/completions",
headers={"authorization": "Bearer test-key"},
json={
"model": "gpt-4o-mini",
"messages": [
{"role": "user", "content": "turn1"},
{"role": "assistant", "content": "turn1-assistant"},
{"role": "user", "content": "current turn"},
],
},
)
assert response.status_code == 200
assert captured["frozen_message_count"] == 2
@pytest.mark.parametrize("tail_role", ["tool", "function"])
def test_openai_cache_mode_keeps_final_tool_observation_mutable(tail_role: str) -> None:
captured = {}
with _make_proxy_client() as client:
proxy = client.app.state.proxy
proxy.config.optimize = True
proxy.config.mode = "cache"
fake_tracker = _FakePrefixTracker(frozen_count=0)
proxy.session_tracker_store.compute_session_id = lambda request, model, messages: (
"stable-session"
)
proxy.session_tracker_store.get_or_create = lambda session_id, provider: fake_tracker
def _fake_apply(**kwargs):
captured.setdefault("calls", []).append(
{
"frozen_message_count": kwargs.get("frozen_message_count"),
"roles": [msg.get("role") for msg in kwargs["messages"]],
"mode": proxy.config.mode,
}
)
return SimpleNamespace(
messages=kwargs["messages"],
transforms_applied=["test:compress-tail"],
timing={},
tokens_before=120,
tokens_after=80,
waste_signals=None,
)
proxy.openai_pipeline.apply = _fake_apply
async def _fake_retry(method, url, headers, body, stream=False, **kwargs): # noqa: ANN001
return httpx.Response(
200,
json={
"id": "chatcmpl_tool_tail",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "ok"},
"finish_reason": "stop",
}
],
"usage": {"prompt_tokens": 80, "completion_tokens": 3, "total_tokens": 83},
},
)
proxy._retry_request = _fake_retry
tail = {
"role": tail_role,
"content": "large command observation " * 200,
}
if tail_role == "tool":
tail["tool_call_id"] = "call_1"
else:
tail["name"] = "bash"
response = client.post(
"/v1/chat/completions",
headers={"authorization": "Bearer test-key"},
json={
"model": "gpt-4o-mini",
"messages": [
{"role": "user", "content": "turn1"},
{"role": "assistant", "content": "run command"},
tail,
],
},
)
assert response.status_code == 200
assert any(call["frozen_message_count"] == 2 for call in captured["calls"]), captured[
"calls"
]
def test_openai_cache_mode_restores_mutated_frozen_prefix() -> None:
captured = {}
with _make_proxy_client() as client:
proxy = client.app.state.proxy
proxy.config.optimize = True
proxy.config.mode = "cache"
fake_tracker = _FakePrefixTracker(frozen_count=0)
proxy.session_tracker_store.compute_session_id = lambda request, model, messages: (
"stable-session"
)
proxy.session_tracker_store.get_or_create = lambda session_id, provider: fake_tracker
original_messages = [
{"role": "user", "content": "turn1"},
{"role": "assistant", "content": "turn1-assistant"},
{"role": "user", "content": "current turn"},
]
def _fake_apply(**kwargs):
mutated = list(kwargs["messages"])
mutated[0] = {**mutated[0], "content": "MUTATED_PREFIX"}
return SimpleNamespace(
messages=mutated,
transforms_applied=["fake:mutated"],
timing={},
tokens_before=70,
tokens_after=65,
waste_signals=None,
)
proxy.openai_pipeline.apply = _fake_apply
async def _fake_retry(method, url, headers, body, stream=False, **kwargs): # noqa: ANN001
captured["body"] = body
return httpx.Response(
200,
json={
"id": "chatcmpl_2",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "ok"},
"finish_reason": "stop",
}
],
"usage": {"prompt_tokens": 65, "completion_tokens": 3, "total_tokens": 68},
},
)
proxy._retry_request = _fake_retry
response = client.post(
"/v1/chat/completions",
headers={"authorization": "Bearer test-key"},
json={
"model": "gpt-4o-mini",
"messages": original_messages,
},
)
assert response.status_code == 200
sent_messages = captured["body"]["messages"]
assert sent_messages[0] == original_messages[0]
assert sent_messages[1] == original_messages[1]
# ─── Issue #327 cross-handler regression ────────────────────────────────
#
# The OpenAI handler was never affected by issue #327's content-keyed walker
# bug — it has only ever used `compute_frozen_count` (positional). This test
# locks that property by spying on the OpenAI traffic path and asserting that
# the buggy walker functions (`should_defer_compression`, `mark_stable`) are
# never called from the production handler. If a future refactor accidentally
# adds the same walker to OpenAI, this test fails immediately.
def test_issue_327_openai_handler_does_not_call_walker_functions() -> None:
calls: list[tuple[str, tuple, dict]] = []
class _SpyCompCache:
def apply_cached(self, messages): # noqa: ANN001
calls.append(("apply_cached", (), {}))
return list(messages)
def compute_frozen_count(self, messages): # noqa: ANN001
calls.append(("compute_frozen_count", (), {}))
return 0
def update_from_result(self, originals, compressed): # noqa: ANN001
calls.append(("update_from_result", (), {}))
def mark_stable_from_messages(self, messages, up_to): # noqa: ANN001
calls.append(("mark_stable_from_messages", (up_to,), {}))
# Methods below MUST NOT be called from OpenAI handler.
def should_defer_compression(self, *args, **kwargs): # noqa: ANN001, ANN002, ANN003
calls.append(("should_defer_compression", args, kwargs))
return False
def mark_stable(self, content_hash): # noqa: ANN001
calls.append(("mark_stable", (content_hash,), {}))
@staticmethod
def content_hash(content): # noqa: ANN001
return f"H({content[:40] if isinstance(content, str) else 'list'})"
with _make_proxy_client() as client:
proxy = client.app.state.proxy
proxy.config.optimize = True
proxy.config.mode = "token" # token mode is where Anthropic had the bug
fake_tracker = _FakePrefixTracker(frozen_count=0)
proxy.session_tracker_store.compute_session_id = lambda request, model, messages: (
"openai-spy-session"
)
proxy.session_tracker_store.get_or_create = lambda s, p: fake_tracker
proxy._get_compression_cache = lambda s: _SpyCompCache()
def _fake_apply(**kwargs): # noqa: ANN003
return SimpleNamespace(
messages=list(kwargs["messages"]),
transforms_applied=[],
timing={},
tokens_before=60,
tokens_after=60,
waste_signals=None,
)
proxy.openai_pipeline.apply = _fake_apply
async def _fake_retry(method, url, headers, body, stream=False, **kwargs): # noqa: ANN001
return httpx.Response(
200,
json={
"id": "cmpl",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "ok"},
"finish_reason": "stop",
}
],
"usage": {"prompt_tokens": 60, "completion_tokens": 3, "total_tokens": 63},
},
)
proxy._retry_request = _fake_retry
# Drive 5 turns so any walker bug would have time to fire repeatedly.
for turn in range(5):
r = client.post(
"/v1/chat/completions",
headers={"authorization": "Bearer test-key"},
json={
"model": "gpt-4o-mini",
"messages": [
{"role": "user", "content": f"turn-{turn}-q"},
{"role": "assistant", "content": f"turn-{turn}-a"},
{"role": "tool", "tool_call_id": "t1", "content": "x" * 600},
{"role": "user", "content": f"continue-{turn}"},
],
},
)
assert r.status_code == 200
method_names = [c[0] for c in calls]
assert "should_defer_compression" not in method_names, (
f"OpenAI handler unexpectedly called should_defer_compression. "
f"Calls observed: {method_names}"
)
assert "mark_stable" not in method_names, (
f"OpenAI handler unexpectedly called mark_stable (the walker side-effect). "
f"Calls observed: {method_names}"
)
# Sanity: the safe positional methods DID fire.
assert "compute_frozen_count" in method_names
assert "apply_cached" in method_names