headroom/examples/strands_mcp_dispatch_test.py

Ignoring revisions in .git-blame-ignore-revs. Click here to bypass and see the normal blame view.

274 lines
11 KiB
Python
Raw Permalink Normal View History

fix(proxy): Strands MCP bundle + backend path fixes + Codex fail-closed protection Three logically-related sets of proxy changes ship in this branch: 1. Strands integration on the Bedrock path (HeadroomBundle + 4 OpenAI handler fixes + LiteLLM cache stats + dep pin) 2. /stats MCP aggregation (cross-process events log → proxy summary) 3. Codex compression-failure fail-closed (WS + HTTP /v1/responses) == 1. Strands integration on the Bedrock path == * HeadroomBundle (headroom/integrations/strands/bundle.py): single-helper MCP wiring for a Strands Agent — Headroom MCP server (headroom_compress / headroom_retrieve / headroom_stats) plus optional Serena MCP and optional in-process compression hook. Constructor builds unstarted MCPClient instances per server; Strands' Agent owns the subprocess lifecycle. Default config: MCP enabled, Serena enabled, hook OFF (proxy is the single source of truth for compression). User-side integration is two lines in any Strands app. * headroom/proxy/handlers/openai.py — backend path now: - calls PrefixCacheTracker.update_from_response (was direct-OpenAI only) - intercepts CCR headroom_retrieve tool_calls server-side, mirroring the Anthropic handler pattern; NO silent fallback, re-raises on CCR errors (per feedback_no_silent_fallbacks) - works for both non-streaming and streaming paths * headroom/proxy/handlers/streaming.py: _stream_openai_via_backend now accepts prefix_tracker + optimized_messages, parses cache stats from the SSE final-usage frame (cache_creation_input_tokens added to the state machine), records CCR retrieve feedback via a new _record_ccr_feedback_from_openai_sse helper. Streaming CCR intercept is intentionally out of scope (mirrors Anthropic streaming behaviour). * headroom/backends/litellm.py: send_openai_message response usage block now carries cache_read_input_tokens / cache_creation_input_tokens (Anthropic/Bedrock dialect) and prompt_tokens_details.cached_tokens (OpenAI dialect). Backwards-compatible — cold-start callers see the same 3-key shape; cache keys appear only when the underlying provider returns them. Pinned by test_no_cache_fields_means_no_cache_keys. * headroom/proxy/auth_mode.py: ("strands-agents/", "strands") added to CLIENT_UA_MAP. Production callers should also set X-Client: strands since the default openai-python UA carries no Strands signal. * pyproject.toml: huggingface-hub>=1.5.0,<2.0 pinned in [ml] so a sibling install (e.g. strands-agents) can't drag the version below the floor transformers 5.x requires (otherwise Kompress silently goes "unavailable"). == 2. /stats MCP aggregation == * headroom/proxy/cost.py: _aggregate_mcp_events() reads the cross-process shared events file the Headroom MCP server already writes to and surfaces summary.mcp with three new keys: - compressions (count of headroom_compress invocations) - tokens_removed (sum of input - output across those) - retrievals (count of headroom_retrieve — the load-bearing over-compression alarm; if it grows linearly with turn count, lossy compressors are dropping info the model actually needs) Defensive on every axis — missing MCP SDK, missing file, malformed events, read errors — never blocks /stats. * examples/strands_bundle_demo.py: stats panel prints the new fields so the demo shows the full proxy-HTTP + MCP-tool story in one view. == 3. Codex compression-failure fail-closed protection == Reported by Camille (2026-05-21): Codex threads were locking with "ran out of room in the model's context window" after Headroom's compression timed out on an oversized response.create frame and forwarded the original ~1.7 MB frame to the upstream, which then rejected it. Codex's auto-compact heuristic gates on the upstream- reported total_usage_tokens (which Headroom had been shrinking on earlier turns), so its compaction never fired and the thread locked. Validated against open Codex issues (CLI + Desktop share codex-rs/core): * #16068 — confirms compaction gates on total_usage_tokens, estimated_token_count is computed but only logged * #19806 — confirms image token estimator unbounded, contributes to the same ContextManager.get_total_token_usage → auto-compaction chain * headroom/proxy/helpers.py: decide_compression_failure_action() with a unit-tested decision matrix: - asyncio.TimeoutError → refuse, always - non-timeout failure + frame > 256 KiB (configurable) → refuse - non-timeout failure + small frame → forward (legacy) Operator escape hatches: - HEADROOM_WS_FAIL_OPEN_ON_COMPRESSION_FAILURE=1 restores legacy - HEADROOM_WS_COMPRESSION_FAIL_THRESHOLD_BYTES tunes the threshold * headroom/proxy/handlers/openai.py (WS /v1/responses): consults the helper after compression failure. On refuse: close client websocket code 1009 with "headroom: compression <reason> — please compact context and retry" reason; set termination_cause for the outer lifecycle finally; return. * headroom/proxy/handlers/openai.py (HTTP /v1/responses): same helper. On refuse: raise HTTPException(413) with a structured error body so FastAPI's HTTPException handler emits a clean 413. The existing `except HTTPException: raise` guard in this handler already ensures the 413 propagates without being swallowed by the 502 catch-all. Anthropic /v1/messages NOT changed in this branch: no equivalent bug report on Anthropic-protocol clients, Claude Code (Anthropic-owned) handles context overflow via its own cache_control/ephemeral primitives, and Cursor/Aider don't maintain the local-Y estimate the Codex bug requires. Deferred until a real report lands; the patch is a one-liner reusing the same helper. == Tests + verification == * tests/test_backends/test_litellm_cache_stats.py — 3 tests pinning cache-stat surfacing across Anthropic/OpenAI dialects + backwards- compat for no-cache responses. * tests/test_proxy/test_openai_backend_path.py — 5 tests (Bedrock cache fields, OpenAI fallback shape, CCR intercept with provider="openai", CCR re-raise on exception, streaming signature contract). * tests/test_proxy/test_mcp_stats_aggregation.py — 5 tests pinning the aggregator across compress+retrieve mixes, empty events, unknown event types, missing token fields, and read failures. * tests/test_proxy/test_compression_failure_action.py — 12 tests pinning the fail-closed decision matrix (timeout always refuses, small transient passes through, oversize refuses, env override variants, custom threshold, invalid threshold falls back, 0/negative ignored). * examples/strands_bedrock_demo.py — model_id bumped from deprecated Claude 3 Haiku to Sonnet 4.5 (the deprecated model now errors on account access). * examples/strands_via_proxy_demo.py — proxy + Bedrock cache + streaming smoke test. * examples/strands_mcp_dispatch_test.py — pure MCP round-trip probe. * examples/strands_bundle_demo.py — full Strands + HeadroomBundle E2E demo (this is the shape a real Strands user copies into their app). Full pytest: 5327 passed, 178 skipped. The previously-failing test_core_operations.py::TestAddBatch::test_add_batch_basic passes now that the huggingface-hub pin in pyproject.toml unblocks transformers imports. E2E verified live against AWS Bedrock (Sonnet 4.5): * cache_write=10,438 on turn A → cache_read=10,438 on turn B * streaming SSE final usage frame carries cache_read_input_tokens * 78.7% reduction on a 50 KB JSON tool_result via SmartCrusher ( dispatched per-content-type by ContentRouter) * Strands Agent + HeadroomBundle: model autonomously called headroom_compress + headroom_retrieve via MCP; CompressionStore round-trip succeeded; final answer correct.
2026-05-21 11:00:14 -07:00
#!/usr/bin/env python3
"""Verify Strands' MCP dispatcher can reach Headroom's MCP server.
This is the focused 15-min test that proves the MCP-everywhere
architecture for Path B works end-to-end:
1. Headroom proxy is running (started separately, port 8787).
2. Build a Strands MCPClient that stdio-spawns `headroom mcp serve`.
3. List MCP tools confirm `headroom_retrieve` is exposed.
4. Round-trip: stash content via the proxy's /v1/compress endpoint
to get a hash, then call `headroom_retrieve(hash)` via MCP,
verify the original comes back.
5. (Bonus) Show the same hash being resolved through both the
proxy's REST /v1/retrieve endpoint AND the MCP path -- proves
the CompressionStore is shared between proxy and MCP server.
If steps 3+4 work, the MCP dispatch path is alive meaning when a
Strands Agent receives a model-emitted `headroom_retrieve` tool_call
(streaming OR non-streaming), it can dispatch it via this same MCP
client and the chain closes.
Run
---
AWS_REGION=us-west-2 python examples/strands_mcp_dispatch_test.py
"""
from __future__ import annotations
import json
import subprocess
import sys
import time
import urllib.error
import urllib.request
from contextlib import suppress
from pathlib import Path
from mcp import StdioServerParameters
from mcp.client.stdio import stdio_client
from strands.tools.mcp import MCPClient
PROXY_PORT = 8787 # matches MCP default
PROXY_URL = f"http://127.0.0.1:{PROXY_PORT}"
def start_proxy() -> subprocess.Popen[bytes]:
"""Spawn the proxy on the MCP default port."""
cmd = [
sys.executable,
"-m",
"headroom.cli",
"proxy",
"--backend",
"bedrock",
"--region",
"us-west-2",
"--port",
str(PROXY_PORT),
]
log = Path("/tmp/headroom_proxy_mcp_test.log").open("wb")
proc = subprocess.Popen(cmd, stdout=log, stderr=subprocess.STDOUT) # noqa: S603
print(f" proxy spawned (pid={proc.pid}); log → /tmp/headroom_proxy_mcp_test.log")
return proc
def wait_for_proxy(timeout_s: float = 30.0) -> None:
deadline = time.time() + timeout_s
while time.time() < deadline:
try:
with urllib.request.urlopen(f"{PROXY_URL}/readyz", timeout=1) as r: # noqa: S310
if r.status == 200:
return
except Exception: # noqa: BLE001
time.sleep(0.5)
raise RuntimeError(f"Proxy did not become ready in {timeout_s}s")
def stop_proxy(proc: subprocess.Popen[bytes]) -> None:
with suppress(ProcessLookupError):
proc.terminate()
try:
proc.wait(timeout=5)
except subprocess.TimeoutExpired:
proc.kill()
def stash_content_via_proxy(content: str) -> str | None:
"""Round-trip through proxy /v1/compress to get a stored hash.
Returns the hash if compression replaced anything, else None.
"""
body = json.dumps({"content": content}).encode()
req = urllib.request.Request( # noqa: S310
f"{PROXY_URL}/v1/compress",
data=body,
headers={"Content-Type": "application/json"},
method="POST",
)
try:
with urllib.request.urlopen(req, timeout=10) as r: # noqa: S310
resp = json.loads(r.read())
except urllib.error.HTTPError as e:
err = e.read().decode("utf-8", errors="replace")[:400]
print(f" /v1/compress HTTP {e.code}: {err}")
return None
print(f" /v1/compress response keys: {list(resp.keys())}")
# The response shape varies — try common shape candidates.
for key in ("hash", "stored_hash", "content_hash"):
if key in resp:
return str(resp[key])
compressed = resp.get("compressed") or resp.get("output") or ""
# Marker form: 'Retrieve original: hash=abc' or 'Retrieve more: hash=abc'
for marker in ("hash=",):
if marker in compressed:
idx = compressed.index(marker) + len(marker)
end = compressed.find(" ", idx)
return compressed[idx : end if end > 0 else idx + 64].strip().rstrip(">")
print(f" no hash found in /v1/compress response — full body: {json.dumps(resp)[:400]}")
return None
def main() -> int:
print("=" * 72)
print(" Strands MCPClient → Headroom MCP server end-to-end probe")
print("=" * 72)
print("\n[1/5] Starting Headroom proxy ...")
proxy = start_proxy()
try:
wait_for_proxy()
print(" proxy ready.\n")
big = json.dumps(
[
{
"id": i,
"name": f"order-{i}",
"status": "completed",
"amount": 100 + i,
"customer": f"cust-{i % 50}",
"description": "A long, repetitive description that takes up space " * 4,
}
for i in range(300)
]
)
print(f"[2/5] Test payload prepared: {len(big)} chars of JSON")
print("\n[3/5] Building Strands MCPClient pointed at `headroom mcp serve` ...")
server_params = StdioServerParameters(
command="headroom",
args=["mcp", "serve", "--proxy-url", PROXY_URL],
)
with MCPClient(lambda: stdio_client(server_params)) as mcp:
print(" MCP connection up.")
print("\n[4/5] Listing MCP tools ...")
try:
tools = mcp.list_tools_sync()
except Exception as e: # noqa: BLE001
print(f" ! list_tools_sync failed: {type(e).__name__}: {e}")
raise
tool_names = [getattr(t, "tool_name", None) or getattr(t, "name", "?") for t in tools]
print(f" tools advertised by Headroom MCP: {tool_names}")
if "headroom_retrieve" not in tool_names:
print(" ! headroom_retrieve NOT in tool list — abort.")
return 4
# Round-trip via MCP: compress to get a hash, then retrieve via MCP.
# This is a pure MCP-only path through Strands' dispatcher — the
# same dispatcher that would resolve a model-emitted
# `headroom_retrieve` tool_call in a real conversation.
print("\n[5a/5] Stashing content via MCP headroom_compress (same dispatcher) ...")
try:
comp = mcp.call_tool_sync(
tool_use_id="probe-compress-1",
name="headroom_compress",
arguments={"content": big},
)
print(f" status={comp.get('status', '?')}")
# Pull the JSON body out of the tool result
comp_payload: dict | None = None
for block in comp.get("content", []) or []:
if isinstance(block, dict):
if "json" in block:
comp_payload = block["json"]
break
if "text" in block:
try:
comp_payload = json.loads(block["text"])
except (json.JSONDecodeError, ValueError):
comp_payload = {"_raw_text": block["text"]}
break
print(f" compress payload keys: {list((comp_payload or {}).keys())}")
# Find the hash. Headroom MCP compress is documented to emit
# markers in the compressed output ("hash=abc..."); also surface
# explicit hash field if present.
stashed_hash: str | None = None
if comp_payload:
for key in ("hash", "stored_hash", "content_hash"):
if key in comp_payload and comp_payload[key]:
stashed_hash = str(comp_payload[key])
break
if not stashed_hash:
compressed_str = (
comp_payload.get("compressed")
or comp_payload.get("output")
or comp_payload.get("_raw_text", "")
)
if "hash=" in compressed_str:
idx = compressed_str.index("hash=") + len("hash=")
tail = compressed_str[idx : idx + 80]
# hash is followed by space, quote, > or end
stashed_hash = ""
for ch in tail:
if ch in " \"'>),\n\r\t":
break
stashed_hash += ch
stashed_hash = stashed_hash or None
if not stashed_hash:
print(
f" ! no hash found in compress response. Full body: "
f"{json.dumps(comp_payload)[:400]}"
)
return 5
print(f" ✓ stashed hash: {stashed_hash[:32]}...")
except Exception as e: # noqa: BLE001
print(f" ! compress call failed: {type(e).__name__}: {e}")
return 5
print("\n[5b/5] Calling headroom_retrieve via Strands MCP dispatcher ...")
try:
result = mcp.call_tool_sync(
tool_use_id="probe-retrieve-1",
name="headroom_retrieve",
arguments={"hash": stashed_hash},
)
print(f" retrieve status: {result.get('status', '?')}")
content = result.get("content", [])
preview = ""
for block in content:
if isinstance(block, dict):
if "text" in block:
preview = block["text"][:300]
break
if "json" in block:
preview = json.dumps(block["json"])[:300]
break
print(f" retrieved content preview: {preview!r}")
if any(needle in preview for needle in ("order-0", "order-1", "order-")):
print(" ✓ retrieved content references the original JSON rows.")
else:
print(
" - preview doesn't obviously match the original. "
"Scan the full preview above to verify."
)
except Exception as e: # noqa: BLE001
print(f" ! retrieve call failed: {type(e).__name__}: {e}")
return 6
print("\n" + "=" * 72)
print(" MCP DISPATCH PROBE COMPLETE")
print(" If steps 3+4+5 succeeded, the MCP-everywhere Path B is live:")
print(" Strands receives a headroom_retrieve tool_call → MCP dispatcher")
print(" → Headroom MCP server → CompressionStore → original content back.")
print("=" * 72)
return 0
finally:
print("\n stopping proxy ...")
stop_proxy(proxy)
if __name__ == "__main__":
sys.exit(main())