mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
Three logically-related sets of proxy changes ship in this branch:
1. Strands integration on the Bedrock path (HeadroomBundle + 4 OpenAI
handler fixes + LiteLLM cache stats + dep pin)
2. /stats MCP aggregation (cross-process events log → proxy summary)
3. Codex compression-failure fail-closed (WS + HTTP /v1/responses)
== 1. Strands integration on the Bedrock path ==
* HeadroomBundle (headroom/integrations/strands/bundle.py): single-helper
MCP wiring for a Strands Agent — Headroom MCP server (headroom_compress
/ headroom_retrieve / headroom_stats) plus optional Serena MCP and
optional in-process compression hook. Constructor builds unstarted
MCPClient instances per server; Strands' Agent owns the subprocess
lifecycle. Default config: MCP enabled, Serena enabled, hook OFF
(proxy is the single source of truth for compression). User-side
integration is two lines in any Strands app.
* headroom/proxy/handlers/openai.py — backend path now:
- calls PrefixCacheTracker.update_from_response (was direct-OpenAI only)
- intercepts CCR headroom_retrieve tool_calls server-side, mirroring
the Anthropic handler pattern; NO silent fallback, re-raises on
CCR errors (per feedback_no_silent_fallbacks)
- works for both non-streaming and streaming paths
* headroom/proxy/handlers/streaming.py: _stream_openai_via_backend now
accepts prefix_tracker + optimized_messages, parses cache stats from
the SSE final-usage frame (cache_creation_input_tokens added to the
state machine), records CCR retrieve feedback via a new
_record_ccr_feedback_from_openai_sse helper. Streaming CCR intercept
is intentionally out of scope (mirrors Anthropic streaming behaviour).
* headroom/backends/litellm.py: send_openai_message response usage block
now carries cache_read_input_tokens / cache_creation_input_tokens
(Anthropic/Bedrock dialect) and prompt_tokens_details.cached_tokens
(OpenAI dialect). Backwards-compatible — cold-start callers see the
same 3-key shape; cache keys appear only when the underlying provider
returns them. Pinned by test_no_cache_fields_means_no_cache_keys.
* headroom/proxy/auth_mode.py: ("strands-agents/", "strands") added to
CLIENT_UA_MAP. Production callers should also set X-Client: strands
since the default openai-python UA carries no Strands signal.
* pyproject.toml: huggingface-hub>=1.5.0,<2.0 pinned in [ml] so a sibling
install (e.g. strands-agents) can't drag the version below the floor
transformers 5.x requires (otherwise Kompress silently goes
"unavailable").
== 2. /stats MCP aggregation ==
* headroom/proxy/cost.py: _aggregate_mcp_events() reads the cross-process
shared events file the Headroom MCP server already writes to and
surfaces summary.mcp with three new keys:
- compressions (count of headroom_compress invocations)
- tokens_removed (sum of input - output across those)
- retrievals (count of headroom_retrieve — the load-bearing
over-compression alarm; if it grows linearly
with turn count, lossy compressors are
dropping info the model actually needs)
Defensive on every axis — missing MCP SDK, missing file, malformed
events, read errors — never blocks /stats.
* examples/strands_bundle_demo.py: stats panel prints the new fields so
the demo shows the full proxy-HTTP + MCP-tool story in one view.
== 3. Codex compression-failure fail-closed protection ==
Reported by Camille (2026-05-21): Codex threads were locking with
"ran out of room in the model's context window" after Headroom's
compression timed out on an oversized response.create frame and
forwarded the original ~1.7 MB frame to the upstream, which then
rejected it. Codex's auto-compact heuristic gates on the upstream-
reported total_usage_tokens (which Headroom had been shrinking on
earlier turns), so its compaction never fired and the thread locked.
Validated against open Codex issues (CLI + Desktop share codex-rs/core):
* #16068 — confirms compaction gates on total_usage_tokens,
estimated_token_count is computed but only logged
* #19806 — confirms image token estimator unbounded, contributes to
the same ContextManager.get_total_token_usage → auto-compaction chain
* headroom/proxy/helpers.py: decide_compression_failure_action() with a
unit-tested decision matrix:
- asyncio.TimeoutError → refuse, always
- non-timeout failure + frame > 256 KiB (configurable) → refuse
- non-timeout failure + small frame → forward (legacy)
Operator escape hatches:
- HEADROOM_WS_FAIL_OPEN_ON_COMPRESSION_FAILURE=1 restores legacy
- HEADROOM_WS_COMPRESSION_FAIL_THRESHOLD_BYTES tunes the threshold
* headroom/proxy/handlers/openai.py (WS /v1/responses): consults the
helper after compression failure. On refuse: close client websocket
code 1009 with "headroom: compression <reason> — please compact
context and retry" reason; set termination_cause for the outer
lifecycle finally; return.
* headroom/proxy/handlers/openai.py (HTTP /v1/responses): same helper.
On refuse: raise HTTPException(413) with a structured error body so
FastAPI's HTTPException handler emits a clean 413. The existing
`except HTTPException: raise` guard in this handler already ensures
the 413 propagates without being swallowed by the 502 catch-all.
Anthropic /v1/messages NOT changed in this branch: no equivalent bug
report on Anthropic-protocol clients, Claude Code (Anthropic-owned)
handles context overflow via its own cache_control/ephemeral
primitives, and Cursor/Aider don't maintain the local-Y estimate the
Codex bug requires. Deferred until a real report lands; the patch is
a one-liner reusing the same helper.
== Tests + verification ==
* tests/test_backends/test_litellm_cache_stats.py — 3 tests pinning
cache-stat surfacing across Anthropic/OpenAI dialects + backwards-
compat for no-cache responses.
* tests/test_proxy/test_openai_backend_path.py — 5 tests (Bedrock cache
fields, OpenAI fallback shape, CCR intercept with provider="openai",
CCR re-raise on exception, streaming signature contract).
* tests/test_proxy/test_mcp_stats_aggregation.py — 5 tests pinning the
aggregator across compress+retrieve mixes, empty events, unknown event
types, missing token fields, and read failures.
* tests/test_proxy/test_compression_failure_action.py — 12 tests pinning
the fail-closed decision matrix (timeout always refuses, small
transient passes through, oversize refuses, env override variants,
custom threshold, invalid threshold falls back, 0/negative ignored).
* examples/strands_bedrock_demo.py — model_id bumped from deprecated
Claude 3 Haiku to Sonnet 4.5 (the deprecated model now errors on
account access).
* examples/strands_via_proxy_demo.py — proxy + Bedrock cache + streaming
smoke test.
* examples/strands_mcp_dispatch_test.py — pure MCP round-trip probe.
* examples/strands_bundle_demo.py — full Strands + HeadroomBundle E2E
demo (this is the shape a real Strands user copies into their app).
Full pytest: 5327 passed, 178 skipped. The previously-failing
test_core_operations.py::TestAddBatch::test_add_batch_basic passes now
that the huggingface-hub pin in pyproject.toml unblocks transformers
imports.
E2E verified live against AWS Bedrock (Sonnet 4.5):
* cache_write=10,438 on turn A → cache_read=10,438 on turn B
* streaming SSE final usage frame carries cache_read_input_tokens
* 78.7% reduction on a 50 KB JSON tool_result via SmartCrusher (
dispatched per-content-type by ContentRouter)
* Strands Agent + HeadroomBundle: model autonomously called
headroom_compress + headroom_retrieve via MCP; CompressionStore
round-trip succeeded; final answer correct.
273 lines
11 KiB
Python
273 lines
11 KiB
Python
#!/usr/bin/env python3
|
|
"""Verify Strands' MCP dispatcher can reach Headroom's MCP server.
|
|
|
|
This is the focused 15-min test that proves the MCP-everywhere
|
|
architecture for Path B works end-to-end:
|
|
|
|
1. Headroom proxy is running (started separately, port 8787).
|
|
2. Build a Strands MCPClient that stdio-spawns `headroom mcp serve`.
|
|
3. List MCP tools — confirm `headroom_retrieve` is exposed.
|
|
4. Round-trip: stash content via the proxy's /v1/compress endpoint
|
|
to get a hash, then call `headroom_retrieve(hash)` via MCP,
|
|
verify the original comes back.
|
|
5. (Bonus) Show the same hash being resolved through both the
|
|
proxy's REST /v1/retrieve endpoint AND the MCP path -- proves
|
|
the CompressionStore is shared between proxy and MCP server.
|
|
|
|
If steps 3+4 work, the MCP dispatch path is alive — meaning when a
|
|
Strands Agent receives a model-emitted `headroom_retrieve` tool_call
|
|
(streaming OR non-streaming), it can dispatch it via this same MCP
|
|
client and the chain closes.
|
|
|
|
Run
|
|
---
|
|
AWS_REGION=us-west-2 python examples/strands_mcp_dispatch_test.py
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
import subprocess
|
|
import sys
|
|
import time
|
|
import urllib.error
|
|
import urllib.request
|
|
from contextlib import suppress
|
|
from pathlib import Path
|
|
|
|
from mcp import StdioServerParameters
|
|
from mcp.client.stdio import stdio_client
|
|
from strands.tools.mcp import MCPClient
|
|
|
|
PROXY_PORT = 8787 # matches MCP default
|
|
PROXY_URL = f"http://127.0.0.1:{PROXY_PORT}"
|
|
|
|
|
|
def start_proxy() -> subprocess.Popen[bytes]:
|
|
"""Spawn the proxy on the MCP default port."""
|
|
cmd = [
|
|
sys.executable,
|
|
"-m",
|
|
"headroom.cli",
|
|
"proxy",
|
|
"--backend",
|
|
"bedrock",
|
|
"--region",
|
|
"us-west-2",
|
|
"--port",
|
|
str(PROXY_PORT),
|
|
]
|
|
log = Path("/tmp/headroom_proxy_mcp_test.log").open("wb")
|
|
proc = subprocess.Popen(cmd, stdout=log, stderr=subprocess.STDOUT) # noqa: S603
|
|
print(f" proxy spawned (pid={proc.pid}); log → /tmp/headroom_proxy_mcp_test.log")
|
|
return proc
|
|
|
|
|
|
def wait_for_proxy(timeout_s: float = 30.0) -> None:
|
|
deadline = time.time() + timeout_s
|
|
while time.time() < deadline:
|
|
try:
|
|
with urllib.request.urlopen(f"{PROXY_URL}/readyz", timeout=1) as r: # noqa: S310
|
|
if r.status == 200:
|
|
return
|
|
except Exception: # noqa: BLE001
|
|
time.sleep(0.5)
|
|
raise RuntimeError(f"Proxy did not become ready in {timeout_s}s")
|
|
|
|
|
|
def stop_proxy(proc: subprocess.Popen[bytes]) -> None:
|
|
with suppress(ProcessLookupError):
|
|
proc.terminate()
|
|
try:
|
|
proc.wait(timeout=5)
|
|
except subprocess.TimeoutExpired:
|
|
proc.kill()
|
|
|
|
|
|
def stash_content_via_proxy(content: str) -> str | None:
|
|
"""Round-trip through proxy /v1/compress to get a stored hash.
|
|
|
|
Returns the hash if compression replaced anything, else None.
|
|
"""
|
|
body = json.dumps({"content": content}).encode()
|
|
req = urllib.request.Request( # noqa: S310
|
|
f"{PROXY_URL}/v1/compress",
|
|
data=body,
|
|
headers={"Content-Type": "application/json"},
|
|
method="POST",
|
|
)
|
|
try:
|
|
with urllib.request.urlopen(req, timeout=10) as r: # noqa: S310
|
|
resp = json.loads(r.read())
|
|
except urllib.error.HTTPError as e:
|
|
err = e.read().decode("utf-8", errors="replace")[:400]
|
|
print(f" /v1/compress HTTP {e.code}: {err}")
|
|
return None
|
|
print(f" /v1/compress response keys: {list(resp.keys())}")
|
|
# The response shape varies — try common shape candidates.
|
|
for key in ("hash", "stored_hash", "content_hash"):
|
|
if key in resp:
|
|
return str(resp[key])
|
|
compressed = resp.get("compressed") or resp.get("output") or ""
|
|
# Marker form: 'Retrieve original: hash=abc' or 'Retrieve more: hash=abc'
|
|
for marker in ("hash=",):
|
|
if marker in compressed:
|
|
idx = compressed.index(marker) + len(marker)
|
|
end = compressed.find(" ", idx)
|
|
return compressed[idx : end if end > 0 else idx + 64].strip().rstrip(">")
|
|
print(f" no hash found in /v1/compress response — full body: {json.dumps(resp)[:400]}")
|
|
return None
|
|
|
|
|
|
def main() -> int:
|
|
print("=" * 72)
|
|
print(" Strands MCPClient → Headroom MCP server end-to-end probe")
|
|
print("=" * 72)
|
|
|
|
print("\n[1/5] Starting Headroom proxy ...")
|
|
proxy = start_proxy()
|
|
try:
|
|
wait_for_proxy()
|
|
print(" proxy ready.\n")
|
|
|
|
big = json.dumps(
|
|
[
|
|
{
|
|
"id": i,
|
|
"name": f"order-{i}",
|
|
"status": "completed",
|
|
"amount": 100 + i,
|
|
"customer": f"cust-{i % 50}",
|
|
"description": "A long, repetitive description that takes up space " * 4,
|
|
}
|
|
for i in range(300)
|
|
]
|
|
)
|
|
print(f"[2/5] Test payload prepared: {len(big)} chars of JSON")
|
|
|
|
print("\n[3/5] Building Strands MCPClient pointed at `headroom mcp serve` ...")
|
|
server_params = StdioServerParameters(
|
|
command="headroom",
|
|
args=["mcp", "serve", "--proxy-url", PROXY_URL],
|
|
)
|
|
with MCPClient(lambda: stdio_client(server_params)) as mcp:
|
|
print(" MCP connection up.")
|
|
|
|
print("\n[4/5] Listing MCP tools ...")
|
|
try:
|
|
tools = mcp.list_tools_sync()
|
|
except Exception as e: # noqa: BLE001
|
|
print(f" ! list_tools_sync failed: {type(e).__name__}: {e}")
|
|
raise
|
|
tool_names = [getattr(t, "tool_name", None) or getattr(t, "name", "?") for t in tools]
|
|
print(f" tools advertised by Headroom MCP: {tool_names}")
|
|
if "headroom_retrieve" not in tool_names:
|
|
print(" ! headroom_retrieve NOT in tool list — abort.")
|
|
return 4
|
|
|
|
# Round-trip via MCP: compress to get a hash, then retrieve via MCP.
|
|
# This is a pure MCP-only path through Strands' dispatcher — the
|
|
# same dispatcher that would resolve a model-emitted
|
|
# `headroom_retrieve` tool_call in a real conversation.
|
|
print("\n[5a/5] Stashing content via MCP headroom_compress (same dispatcher) ...")
|
|
try:
|
|
comp = mcp.call_tool_sync(
|
|
tool_use_id="probe-compress-1",
|
|
name="headroom_compress",
|
|
arguments={"content": big},
|
|
)
|
|
print(f" status={comp.get('status', '?')}")
|
|
# Pull the JSON body out of the tool result
|
|
comp_payload: dict | None = None
|
|
for block in comp.get("content", []) or []:
|
|
if isinstance(block, dict):
|
|
if "json" in block:
|
|
comp_payload = block["json"]
|
|
break
|
|
if "text" in block:
|
|
try:
|
|
comp_payload = json.loads(block["text"])
|
|
except (json.JSONDecodeError, ValueError):
|
|
comp_payload = {"_raw_text": block["text"]}
|
|
break
|
|
print(f" compress payload keys: {list((comp_payload or {}).keys())}")
|
|
# Find the hash. Headroom MCP compress is documented to emit
|
|
# markers in the compressed output ("hash=abc..."); also surface
|
|
# explicit hash field if present.
|
|
stashed_hash: str | None = None
|
|
if comp_payload:
|
|
for key in ("hash", "stored_hash", "content_hash"):
|
|
if key in comp_payload and comp_payload[key]:
|
|
stashed_hash = str(comp_payload[key])
|
|
break
|
|
if not stashed_hash:
|
|
compressed_str = (
|
|
comp_payload.get("compressed")
|
|
or comp_payload.get("output")
|
|
or comp_payload.get("_raw_text", "")
|
|
)
|
|
if "hash=" in compressed_str:
|
|
idx = compressed_str.index("hash=") + len("hash=")
|
|
tail = compressed_str[idx : idx + 80]
|
|
# hash is followed by space, quote, > or end
|
|
stashed_hash = ""
|
|
for ch in tail:
|
|
if ch in " \"'>),\n\r\t":
|
|
break
|
|
stashed_hash += ch
|
|
stashed_hash = stashed_hash or None
|
|
if not stashed_hash:
|
|
print(
|
|
f" ! no hash found in compress response. Full body: "
|
|
f"{json.dumps(comp_payload)[:400]}"
|
|
)
|
|
return 5
|
|
print(f" ✓ stashed hash: {stashed_hash[:32]}...")
|
|
except Exception as e: # noqa: BLE001
|
|
print(f" ! compress call failed: {type(e).__name__}: {e}")
|
|
return 5
|
|
|
|
print("\n[5b/5] Calling headroom_retrieve via Strands MCP dispatcher ...")
|
|
try:
|
|
result = mcp.call_tool_sync(
|
|
tool_use_id="probe-retrieve-1",
|
|
name="headroom_retrieve",
|
|
arguments={"hash": stashed_hash},
|
|
)
|
|
print(f" retrieve status: {result.get('status', '?')}")
|
|
content = result.get("content", [])
|
|
preview = ""
|
|
for block in content:
|
|
if isinstance(block, dict):
|
|
if "text" in block:
|
|
preview = block["text"][:300]
|
|
break
|
|
if "json" in block:
|
|
preview = json.dumps(block["json"])[:300]
|
|
break
|
|
print(f" retrieved content preview: {preview!r}")
|
|
if any(needle in preview for needle in ("order-0", "order-1", "order-")):
|
|
print(" ✓ retrieved content references the original JSON rows.")
|
|
else:
|
|
print(
|
|
" - preview doesn't obviously match the original. "
|
|
"Scan the full preview above to verify."
|
|
)
|
|
except Exception as e: # noqa: BLE001
|
|
print(f" ! retrieve call failed: {type(e).__name__}: {e}")
|
|
return 6
|
|
|
|
print("\n" + "=" * 72)
|
|
print(" MCP DISPATCH PROBE COMPLETE")
|
|
print(" If steps 3+4+5 succeeded, the MCP-everywhere Path B is live:")
|
|
print(" Strands receives a headroom_retrieve tool_call → MCP dispatcher")
|
|
print(" → Headroom MCP server → CompressionStore → original content back.")
|
|
print("=" * 72)
|
|
return 0
|
|
finally:
|
|
print("\n stopping proxy ...")
|
|
stop_proxy(proxy)
|
|
|
|
|
|
if __name__ == "__main__":
|
|
sys.exit(main())
|