headroom/tests/test_proxy/test_mcp_stats_aggregation.py

Ignoring revisions in .git-blame-ignore-revs. Click here to bypass and see the normal blame view.

96 lines
3.3 KiB
Python
Raw Permalink Normal View History

fix(proxy): Strands MCP bundle + backend path fixes + Codex fail-closed protection Three logically-related sets of proxy changes ship in this branch: 1. Strands integration on the Bedrock path (HeadroomBundle + 4 OpenAI handler fixes + LiteLLM cache stats + dep pin) 2. /stats MCP aggregation (cross-process events log → proxy summary) 3. Codex compression-failure fail-closed (WS + HTTP /v1/responses) == 1. Strands integration on the Bedrock path == * HeadroomBundle (headroom/integrations/strands/bundle.py): single-helper MCP wiring for a Strands Agent — Headroom MCP server (headroom_compress / headroom_retrieve / headroom_stats) plus optional Serena MCP and optional in-process compression hook. Constructor builds unstarted MCPClient instances per server; Strands' Agent owns the subprocess lifecycle. Default config: MCP enabled, Serena enabled, hook OFF (proxy is the single source of truth for compression). User-side integration is two lines in any Strands app. * headroom/proxy/handlers/openai.py — backend path now: - calls PrefixCacheTracker.update_from_response (was direct-OpenAI only) - intercepts CCR headroom_retrieve tool_calls server-side, mirroring the Anthropic handler pattern; NO silent fallback, re-raises on CCR errors (per feedback_no_silent_fallbacks) - works for both non-streaming and streaming paths * headroom/proxy/handlers/streaming.py: _stream_openai_via_backend now accepts prefix_tracker + optimized_messages, parses cache stats from the SSE final-usage frame (cache_creation_input_tokens added to the state machine), records CCR retrieve feedback via a new _record_ccr_feedback_from_openai_sse helper. Streaming CCR intercept is intentionally out of scope (mirrors Anthropic streaming behaviour). * headroom/backends/litellm.py: send_openai_message response usage block now carries cache_read_input_tokens / cache_creation_input_tokens (Anthropic/Bedrock dialect) and prompt_tokens_details.cached_tokens (OpenAI dialect). Backwards-compatible — cold-start callers see the same 3-key shape; cache keys appear only when the underlying provider returns them. Pinned by test_no_cache_fields_means_no_cache_keys. * headroom/proxy/auth_mode.py: ("strands-agents/", "strands") added to CLIENT_UA_MAP. Production callers should also set X-Client: strands since the default openai-python UA carries no Strands signal. * pyproject.toml: huggingface-hub>=1.5.0,<2.0 pinned in [ml] so a sibling install (e.g. strands-agents) can't drag the version below the floor transformers 5.x requires (otherwise Kompress silently goes "unavailable"). == 2. /stats MCP aggregation == * headroom/proxy/cost.py: _aggregate_mcp_events() reads the cross-process shared events file the Headroom MCP server already writes to and surfaces summary.mcp with three new keys: - compressions (count of headroom_compress invocations) - tokens_removed (sum of input - output across those) - retrievals (count of headroom_retrieve — the load-bearing over-compression alarm; if it grows linearly with turn count, lossy compressors are dropping info the model actually needs) Defensive on every axis — missing MCP SDK, missing file, malformed events, read errors — never blocks /stats. * examples/strands_bundle_demo.py: stats panel prints the new fields so the demo shows the full proxy-HTTP + MCP-tool story in one view. == 3. Codex compression-failure fail-closed protection == Reported by Camille (2026-05-21): Codex threads were locking with "ran out of room in the model's context window" after Headroom's compression timed out on an oversized response.create frame and forwarded the original ~1.7 MB frame to the upstream, which then rejected it. Codex's auto-compact heuristic gates on the upstream- reported total_usage_tokens (which Headroom had been shrinking on earlier turns), so its compaction never fired and the thread locked. Validated against open Codex issues (CLI + Desktop share codex-rs/core): * #16068 — confirms compaction gates on total_usage_tokens, estimated_token_count is computed but only logged * #19806 — confirms image token estimator unbounded, contributes to the same ContextManager.get_total_token_usage → auto-compaction chain * headroom/proxy/helpers.py: decide_compression_failure_action() with a unit-tested decision matrix: - asyncio.TimeoutError → refuse, always - non-timeout failure + frame > 256 KiB (configurable) → refuse - non-timeout failure + small frame → forward (legacy) Operator escape hatches: - HEADROOM_WS_FAIL_OPEN_ON_COMPRESSION_FAILURE=1 restores legacy - HEADROOM_WS_COMPRESSION_FAIL_THRESHOLD_BYTES tunes the threshold * headroom/proxy/handlers/openai.py (WS /v1/responses): consults the helper after compression failure. On refuse: close client websocket code 1009 with "headroom: compression <reason> — please compact context and retry" reason; set termination_cause for the outer lifecycle finally; return. * headroom/proxy/handlers/openai.py (HTTP /v1/responses): same helper. On refuse: raise HTTPException(413) with a structured error body so FastAPI's HTTPException handler emits a clean 413. The existing `except HTTPException: raise` guard in this handler already ensures the 413 propagates without being swallowed by the 502 catch-all. Anthropic /v1/messages NOT changed in this branch: no equivalent bug report on Anthropic-protocol clients, Claude Code (Anthropic-owned) handles context overflow via its own cache_control/ephemeral primitives, and Cursor/Aider don't maintain the local-Y estimate the Codex bug requires. Deferred until a real report lands; the patch is a one-liner reusing the same helper. == Tests + verification == * tests/test_backends/test_litellm_cache_stats.py — 3 tests pinning cache-stat surfacing across Anthropic/OpenAI dialects + backwards- compat for no-cache responses. * tests/test_proxy/test_openai_backend_path.py — 5 tests (Bedrock cache fields, OpenAI fallback shape, CCR intercept with provider="openai", CCR re-raise on exception, streaming signature contract). * tests/test_proxy/test_mcp_stats_aggregation.py — 5 tests pinning the aggregator across compress+retrieve mixes, empty events, unknown event types, missing token fields, and read failures. * tests/test_proxy/test_compression_failure_action.py — 12 tests pinning the fail-closed decision matrix (timeout always refuses, small transient passes through, oversize refuses, env override variants, custom threshold, invalid threshold falls back, 0/negative ignored). * examples/strands_bedrock_demo.py — model_id bumped from deprecated Claude 3 Haiku to Sonnet 4.5 (the deprecated model now errors on account access). * examples/strands_via_proxy_demo.py — proxy + Bedrock cache + streaming smoke test. * examples/strands_mcp_dispatch_test.py — pure MCP round-trip probe. * examples/strands_bundle_demo.py — full Strands + HeadroomBundle E2E demo (this is the shape a real Strands user copies into their app). Full pytest: 5327 passed, 178 skipped. The previously-failing test_core_operations.py::TestAddBatch::test_add_batch_basic passes now that the huggingface-hub pin in pyproject.toml unblocks transformers imports. E2E verified live against AWS Bedrock (Sonnet 4.5): * cache_write=10,438 on turn A → cache_read=10,438 on turn B * streaming SSE final usage frame carries cache_read_input_tokens * 78.7% reduction on a 50 KB JSON tool_result via SmartCrusher ( dispatched per-content-type by ContentRouter) * Strands Agent + HeadroomBundle: model autonomously called headroom_compress + headroom_retrieve via MCP; CompressionStore round-trip succeeded; final answer correct.
2026-05-21 11:00:14 -07:00
"""Coverage for the MCP-events aggregator inside the proxy /stats summary.
``headroom mcp serve`` writes every compress / retrieve invocation to a
shared file-locked log (``_append_shared_event``). Before this fix,
``/stats`` only reported proxy-HTTP-path compressions and silently
ignored the MCP tool work which is exactly where Strands-style
agents spend most of their compression budget when the LLM calls
``headroom_compress`` directly.
These tests pin the aggregation logic in
:func:`headroom.proxy.cost._aggregate_mcp_events`.
"""
from __future__ import annotations
from typing import Any
from unittest.mock import patch
from headroom.proxy.cost import _aggregate_mcp_events
def _mock_events(events: list[dict[str, Any]]) -> Any:
"""Return a patch target that makes _read_shared_events yield ``events``."""
return patch(
"headroom.ccr.mcp_server._read_shared_events",
return_value=events,
)
def test_aggregates_compress_and_retrieve_events() -> None:
events = [
{"type": "compress", "input_tokens": 1000, "output_tokens": 400},
{"type": "compress", "input_tokens": 500, "output_tokens": 300},
{"type": "retrieve", "hash": "abc123"},
{"type": "retrieve", "hash": "def456"},
{"type": "retrieve", "hash": "ghi789"},
]
with _mock_events(events):
result = _aggregate_mcp_events()
assert result == {
"compressions": 2,
# (1000-400) + (500-300) = 600 + 200 = 800
"tokens_removed": 800,
"retrievals": 3,
}
def test_returns_zeros_when_no_events() -> None:
with _mock_events([]):
assert _aggregate_mcp_events() == {
"compressions": 0,
"tokens_removed": 0,
"retrievals": 0,
}
def test_unknown_event_types_are_ignored() -> None:
events = [
{"type": "compress", "input_tokens": 100, "output_tokens": 40},
{"type": "unknown_future_kind", "anything": "goes"},
{"type": "retrieve", "hash": "abc"},
{"missing_type": True}, # malformed; should skip
]
with _mock_events(events):
result = _aggregate_mcp_events()
assert result == {"compressions": 1, "tokens_removed": 60, "retrievals": 1}
def test_missing_token_fields_default_to_zero_without_raising() -> None:
events = [
{"type": "compress"}, # no token fields at all
{"type": "compress", "input_tokens": 200, "output_tokens": None}, # None coerced
{"type": "compress", "input_tokens": 100, "output_tokens": 100}, # zero diff
]
with _mock_events(events):
result = _aggregate_mcp_events()
assert result["compressions"] == 3
# First contributes 0 (both missing); second contributes 0 (None→0,
# and output 0 → input - output = 200... wait that's NOT zero); third 0.
# Let me re-derive: max(0, 0 - 0) + max(0, 200 - 0) + max(0, 100 - 100)
# = 0 + 200 + 0 = 200
assert result["tokens_removed"] == 200
def test_read_failure_yields_zeros() -> None:
"""If the shared-stats reader raises, the aggregator must not crash /stats."""
with patch(
"headroom.ccr.mcp_server._read_shared_events",
side_effect=OSError("disk on fire"),
):
assert _aggregate_mcp_events() == {
"compressions": 0,
"tokens_removed": 0,
"retrievals": 0,
}