mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-10 14:27:00 -04:00
Three logically-related sets of proxy changes ship in this branch:
1. Strands integration on the Bedrock path (HeadroomBundle + 4 OpenAI
handler fixes + LiteLLM cache stats + dep pin)
2. /stats MCP aggregation (cross-process events log → proxy summary)
3. Codex compression-failure fail-closed (WS + HTTP /v1/responses)
== 1. Strands integration on the Bedrock path ==
* HeadroomBundle (headroom/integrations/strands/bundle.py): single-helper
MCP wiring for a Strands Agent — Headroom MCP server (headroom_compress
/ headroom_retrieve / headroom_stats) plus optional Serena MCP and
optional in-process compression hook. Constructor builds unstarted
MCPClient instances per server; Strands' Agent owns the subprocess
lifecycle. Default config: MCP enabled, Serena enabled, hook OFF
(proxy is the single source of truth for compression). User-side
integration is two lines in any Strands app.
* headroom/proxy/handlers/openai.py — backend path now:
- calls PrefixCacheTracker.update_from_response (was direct-OpenAI only)
- intercepts CCR headroom_retrieve tool_calls server-side, mirroring
the Anthropic handler pattern; NO silent fallback, re-raises on
CCR errors (per feedback_no_silent_fallbacks)
- works for both non-streaming and streaming paths
* headroom/proxy/handlers/streaming.py: _stream_openai_via_backend now
accepts prefix_tracker + optimized_messages, parses cache stats from
the SSE final-usage frame (cache_creation_input_tokens added to the
state machine), records CCR retrieve feedback via a new
_record_ccr_feedback_from_openai_sse helper. Streaming CCR intercept
is intentionally out of scope (mirrors Anthropic streaming behaviour).
* headroom/backends/litellm.py: send_openai_message response usage block
now carries cache_read_input_tokens / cache_creation_input_tokens
(Anthropic/Bedrock dialect) and prompt_tokens_details.cached_tokens
(OpenAI dialect). Backwards-compatible — cold-start callers see the
same 3-key shape; cache keys appear only when the underlying provider
returns them. Pinned by test_no_cache_fields_means_no_cache_keys.
* headroom/proxy/auth_mode.py: ("strands-agents/", "strands") added to
CLIENT_UA_MAP. Production callers should also set X-Client: strands
since the default openai-python UA carries no Strands signal.
* pyproject.toml: huggingface-hub>=1.5.0,<2.0 pinned in [ml] so a sibling
install (e.g. strands-agents) can't drag the version below the floor
transformers 5.x requires (otherwise Kompress silently goes
"unavailable").
== 2. /stats MCP aggregation ==
* headroom/proxy/cost.py: _aggregate_mcp_events() reads the cross-process
shared events file the Headroom MCP server already writes to and
surfaces summary.mcp with three new keys:
- compressions (count of headroom_compress invocations)
- tokens_removed (sum of input - output across those)
- retrievals (count of headroom_retrieve — the load-bearing
over-compression alarm; if it grows linearly
with turn count, lossy compressors are
dropping info the model actually needs)
Defensive on every axis — missing MCP SDK, missing file, malformed
events, read errors — never blocks /stats.
* examples/strands_bundle_demo.py: stats panel prints the new fields so
the demo shows the full proxy-HTTP + MCP-tool story in one view.
== 3. Codex compression-failure fail-closed protection ==
Reported by Camille (2026-05-21): Codex threads were locking with
"ran out of room in the model's context window" after Headroom's
compression timed out on an oversized response.create frame and
forwarded the original ~1.7 MB frame to the upstream, which then
rejected it. Codex's auto-compact heuristic gates on the upstream-
reported total_usage_tokens (which Headroom had been shrinking on
earlier turns), so its compaction never fired and the thread locked.
Validated against open Codex issues (CLI + Desktop share codex-rs/core):
* #16068 — confirms compaction gates on total_usage_tokens,
estimated_token_count is computed but only logged
* #19806 — confirms image token estimator unbounded, contributes to
the same ContextManager.get_total_token_usage → auto-compaction chain
* headroom/proxy/helpers.py: decide_compression_failure_action() with a
unit-tested decision matrix:
- asyncio.TimeoutError → refuse, always
- non-timeout failure + frame > 256 KiB (configurable) → refuse
- non-timeout failure + small frame → forward (legacy)
Operator escape hatches:
- HEADROOM_WS_FAIL_OPEN_ON_COMPRESSION_FAILURE=1 restores legacy
- HEADROOM_WS_COMPRESSION_FAIL_THRESHOLD_BYTES tunes the threshold
* headroom/proxy/handlers/openai.py (WS /v1/responses): consults the
helper after compression failure. On refuse: close client websocket
code 1009 with "headroom: compression <reason> — please compact
context and retry" reason; set termination_cause for the outer
lifecycle finally; return.
* headroom/proxy/handlers/openai.py (HTTP /v1/responses): same helper.
On refuse: raise HTTPException(413) with a structured error body so
FastAPI's HTTPException handler emits a clean 413. The existing
`except HTTPException: raise` guard in this handler already ensures
the 413 propagates without being swallowed by the 502 catch-all.
Anthropic /v1/messages NOT changed in this branch: no equivalent bug
report on Anthropic-protocol clients, Claude Code (Anthropic-owned)
handles context overflow via its own cache_control/ephemeral
primitives, and Cursor/Aider don't maintain the local-Y estimate the
Codex bug requires. Deferred until a real report lands; the patch is
a one-liner reusing the same helper.
== Tests + verification ==
* tests/test_backends/test_litellm_cache_stats.py — 3 tests pinning
cache-stat surfacing across Anthropic/OpenAI dialects + backwards-
compat for no-cache responses.
* tests/test_proxy/test_openai_backend_path.py — 5 tests (Bedrock cache
fields, OpenAI fallback shape, CCR intercept with provider="openai",
CCR re-raise on exception, streaming signature contract).
* tests/test_proxy/test_mcp_stats_aggregation.py — 5 tests pinning the
aggregator across compress+retrieve mixes, empty events, unknown event
types, missing token fields, and read failures.
* tests/test_proxy/test_compression_failure_action.py — 12 tests pinning
the fail-closed decision matrix (timeout always refuses, small
transient passes through, oversize refuses, env override variants,
custom threshold, invalid threshold falls back, 0/negative ignored).
* examples/strands_bedrock_demo.py — model_id bumped from deprecated
Claude 3 Haiku to Sonnet 4.5 (the deprecated model now errors on
account access).
* examples/strands_via_proxy_demo.py — proxy + Bedrock cache + streaming
smoke test.
* examples/strands_mcp_dispatch_test.py — pure MCP round-trip probe.
* examples/strands_bundle_demo.py — full Strands + HeadroomBundle E2E
demo (this is the shape a real Strands user copies into their app).
Full pytest: 5327 passed, 178 skipped. The previously-failing
test_core_operations.py::TestAddBatch::test_add_batch_basic passes now
that the huggingface-hub pin in pyproject.toml unblocks transformers
imports.
E2E verified live against AWS Bedrock (Sonnet 4.5):
* cache_write=10,438 on turn A → cache_read=10,438 on turn B
* streaming SSE final usage frame carries cache_read_input_tokens
* 78.7% reduction on a 50 KB JSON tool_result via SmartCrusher (
dispatched per-content-type by ContentRouter)
* Strands Agent + HeadroomBundle: model autonomously called
headroom_compress + headroom_retrieve via MCP; CompressionStore
round-trip succeeded; final answer correct.
217 lines
8 KiB
Python
217 lines
8 KiB
Python
"""Cache-stat surfacing for `LiteLLMBackend.send_openai_message`.
|
|
|
|
LiteLLM normalizes prompt-cache statistics onto its `Usage` object from
|
|
multiple upstream dialects:
|
|
|
|
* Anthropic / Bedrock-Claude → top-level attrs `cache_read_input_tokens`
|
|
and `cache_creation_input_tokens` (also mirrored into
|
|
`prompt_tokens_details.cached_tokens` / `cache_creation_tokens`).
|
|
* OpenAI prompt-caching → only `prompt_tokens_details.cached_tokens`.
|
|
|
|
Before the fix, `send_openai_message` flattened only
|
|
`prompt_tokens / completion_tokens / total_tokens` into the response dict
|
|
and silently dropped all cache stats on the floor — breaking
|
|
`PrefixCacheTracker.update_from_response` for the entire backend-routed
|
|
path (it always saw zero cache hits, so live-zone-only compression never
|
|
engaged).
|
|
|
|
These tests pin the contract for the three relevant shapes.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from typing import Any
|
|
from unittest.mock import AsyncMock, MagicMock, patch
|
|
|
|
from tests._dotenv import importorskip_no_env_leak
|
|
|
|
importorskip_no_env_leak("litellm")
|
|
|
|
from headroom.backends.litellm import LiteLLMBackend # noqa: E402 (must follow importorskip)
|
|
|
|
|
|
class _FakeUsage:
|
|
"""Stand-in for `litellm.types.utils.Usage`.
|
|
|
|
`MagicMock` auto-creates attributes on access, which would defeat the
|
|
point of the "no cache fields → no keys added" test. A plain object
|
|
with only the attributes we explicitly set keeps `getattr(..., 0)`
|
|
honest.
|
|
"""
|
|
|
|
def __init__(
|
|
self,
|
|
*,
|
|
prompt_tokens: int,
|
|
completion_tokens: int,
|
|
total_tokens: int,
|
|
cache_read_input_tokens: int | None = None,
|
|
cache_creation_input_tokens: int | None = None,
|
|
prompt_tokens_details: Any | None = None,
|
|
) -> None:
|
|
self.prompt_tokens = prompt_tokens
|
|
self.completion_tokens = completion_tokens
|
|
self.total_tokens = total_tokens
|
|
if cache_read_input_tokens is not None:
|
|
self.cache_read_input_tokens = cache_read_input_tokens
|
|
if cache_creation_input_tokens is not None:
|
|
self.cache_creation_input_tokens = cache_creation_input_tokens
|
|
if prompt_tokens_details is not None:
|
|
self.prompt_tokens_details = prompt_tokens_details
|
|
|
|
|
|
class _FakePromptTokensDetails:
|
|
"""OpenAI-style nested cache shape stand-in."""
|
|
|
|
def __init__(
|
|
self,
|
|
*,
|
|
cached_tokens: int | None = None,
|
|
cache_creation_tokens: int | None = None,
|
|
) -> None:
|
|
if cached_tokens is not None:
|
|
self.cached_tokens = cached_tokens
|
|
if cache_creation_tokens is not None:
|
|
self.cache_creation_tokens = cache_creation_tokens
|
|
|
|
|
|
def _make_response(usage: _FakeUsage) -> MagicMock:
|
|
"""Build a minimal `ModelResponse`-shaped mock with the given usage."""
|
|
response = MagicMock()
|
|
response.id = "chatcmpl-test"
|
|
response.created = 1_700_000_000
|
|
response.choices = [
|
|
MagicMock(
|
|
index=0,
|
|
message=MagicMock(role="assistant", content="hi", tool_calls=None),
|
|
finish_reason="stop",
|
|
)
|
|
]
|
|
response.usage = usage
|
|
return response
|
|
|
|
|
|
def _make_backend() -> LiteLLMBackend:
|
|
# Patch the inference-profile fetch so `__init__` doesn't try to talk to AWS.
|
|
with patch("headroom.backends.litellm._fetch_bedrock_inference_profiles", return_value={}):
|
|
return LiteLLMBackend(provider="openrouter")
|
|
|
|
|
|
def _request_body() -> dict[str, Any]:
|
|
return {
|
|
"model": "gpt-4",
|
|
"messages": [{"role": "user", "content": "hello"}],
|
|
"max_tokens": 32,
|
|
}
|
|
|
|
|
|
# =============================================================================
|
|
# 1. Anthropic-style (top-level cache_read_input_tokens / cache_creation_input_tokens)
|
|
# =============================================================================
|
|
|
|
|
|
async def test_anthropic_style_cache_fields_surface_in_usage_block() -> None:
|
|
"""Bedrock-Claude / Anthropic responses set the top-level dialect.
|
|
|
|
LiteLLM mirrors them into `prompt_tokens_details` too. Our extractor
|
|
must prefer the explicit top-level values (cache_read=1500, cache_write=200)
|
|
and also expose the OpenAI nested shape so single-dialect callers
|
|
don't have to branch.
|
|
"""
|
|
usage = _FakeUsage(
|
|
prompt_tokens=2000,
|
|
completion_tokens=100,
|
|
total_tokens=2100,
|
|
cache_read_input_tokens=1500,
|
|
cache_creation_input_tokens=200,
|
|
prompt_tokens_details=_FakePromptTokensDetails(
|
|
cached_tokens=1500,
|
|
cache_creation_tokens=200,
|
|
),
|
|
)
|
|
response = _make_response(usage)
|
|
|
|
backend = _make_backend()
|
|
with patch("headroom.backends.litellm.acompletion", new_callable=AsyncMock) as mock_acomp:
|
|
mock_acomp.return_value = response
|
|
result = await backend.send_openai_message(_request_body(), {})
|
|
|
|
body_usage = result.body["usage"]
|
|
assert body_usage["prompt_tokens"] == 2000
|
|
assert body_usage["completion_tokens"] == 100
|
|
assert body_usage["total_tokens"] == 2100
|
|
assert body_usage["cache_read_input_tokens"] == 1500
|
|
assert body_usage["cache_creation_input_tokens"] == 200
|
|
assert body_usage["prompt_tokens_details"] == {"cached_tokens": 1500}
|
|
|
|
|
|
# =============================================================================
|
|
# 2. OpenAI-style only (prompt_tokens_details.cached_tokens, no top-level)
|
|
# =============================================================================
|
|
|
|
|
|
async def test_openai_nested_cache_fields_surface_when_top_level_absent() -> None:
|
|
"""OpenAI prompt-caching responses only populate the nested dialect.
|
|
|
|
With no top-level `cache_read_input_tokens` attribute on the Usage
|
|
object, we must fall back to `prompt_tokens_details.cached_tokens`
|
|
and mirror it into the Anthropic-style top-level keys for downstream
|
|
consumers.
|
|
"""
|
|
usage = _FakeUsage(
|
|
prompt_tokens=1200,
|
|
completion_tokens=50,
|
|
total_tokens=1250,
|
|
prompt_tokens_details=_FakePromptTokensDetails(cached_tokens=800),
|
|
)
|
|
response = _make_response(usage)
|
|
|
|
backend = _make_backend()
|
|
with patch("headroom.backends.litellm.acompletion", new_callable=AsyncMock) as mock_acomp:
|
|
mock_acomp.return_value = response
|
|
result = await backend.send_openai_message(_request_body(), {})
|
|
|
|
body_usage = result.body["usage"]
|
|
assert body_usage["prompt_tokens"] == 1200
|
|
assert body_usage["completion_tokens"] == 50
|
|
assert body_usage["total_tokens"] == 1250
|
|
assert body_usage["cache_read_input_tokens"] == 800
|
|
assert body_usage["cache_creation_input_tokens"] == 0
|
|
assert body_usage["prompt_tokens_details"] == {"cached_tokens": 800}
|
|
|
|
|
|
# =============================================================================
|
|
# 3. Cold start — no cache fields anywhere → keep usage_block shape stable
|
|
# =============================================================================
|
|
|
|
|
|
async def test_no_cache_fields_means_no_cache_keys_in_usage_block() -> None:
|
|
"""Cold-start path: no cache attributes at all on the Usage object.
|
|
|
|
We must NOT inject `cache_read_input_tokens`, `cache_creation_input_tokens`,
|
|
or `prompt_tokens_details` into `usage_block` — keep the dict shape
|
|
identical to the pre-fix behaviour so callers that key off presence
|
|
(rather than value) don't accidentally start seeing 0 as "we have
|
|
cache data, the model just didn't cache".
|
|
"""
|
|
usage = _FakeUsage(
|
|
prompt_tokens=500,
|
|
completion_tokens=25,
|
|
total_tokens=525,
|
|
)
|
|
response = _make_response(usage)
|
|
|
|
backend = _make_backend()
|
|
with patch("headroom.backends.litellm.acompletion", new_callable=AsyncMock) as mock_acomp:
|
|
mock_acomp.return_value = response
|
|
result = await backend.send_openai_message(_request_body(), {})
|
|
|
|
body_usage = result.body["usage"]
|
|
assert body_usage == {
|
|
"prompt_tokens": 500,
|
|
"completion_tokens": 25,
|
|
"total_tokens": 525,
|
|
}
|
|
assert "cache_read_input_tokens" not in body_usage
|
|
assert "cache_creation_input_tokens" not in body_usage
|
|
assert "prompt_tokens_details" not in body_usage
|