headroom/examples
chopratejas 20dc1f28f3 fix(proxy): Strands MCP bundle + backend path fixes + Codex fail-closed protection
Three logically-related sets of proxy changes ship in this branch:

1. Strands integration on the Bedrock path (HeadroomBundle + 4 OpenAI
   handler fixes + LiteLLM cache stats + dep pin)
2. /stats MCP aggregation (cross-process events log → proxy summary)
3. Codex compression-failure fail-closed (WS + HTTP /v1/responses)

== 1. Strands integration on the Bedrock path ==

* HeadroomBundle (headroom/integrations/strands/bundle.py): single-helper
  MCP wiring for a Strands Agent — Headroom MCP server (headroom_compress
  / headroom_retrieve / headroom_stats) plus optional Serena MCP and
  optional in-process compression hook. Constructor builds unstarted
  MCPClient instances per server; Strands' Agent owns the subprocess
  lifecycle. Default config: MCP enabled, Serena enabled, hook OFF
  (proxy is the single source of truth for compression). User-side
  integration is two lines in any Strands app.

* headroom/proxy/handlers/openai.py — backend path now:
  - calls PrefixCacheTracker.update_from_response (was direct-OpenAI only)
  - intercepts CCR headroom_retrieve tool_calls server-side, mirroring
    the Anthropic handler pattern; NO silent fallback, re-raises on
    CCR errors (per feedback_no_silent_fallbacks)
  - works for both non-streaming and streaming paths

* headroom/proxy/handlers/streaming.py: _stream_openai_via_backend now
  accepts prefix_tracker + optimized_messages, parses cache stats from
  the SSE final-usage frame (cache_creation_input_tokens added to the
  state machine), records CCR retrieve feedback via a new
  _record_ccr_feedback_from_openai_sse helper. Streaming CCR intercept
  is intentionally out of scope (mirrors Anthropic streaming behaviour).

* headroom/backends/litellm.py: send_openai_message response usage block
  now carries cache_read_input_tokens / cache_creation_input_tokens
  (Anthropic/Bedrock dialect) and prompt_tokens_details.cached_tokens
  (OpenAI dialect). Backwards-compatible — cold-start callers see the
  same 3-key shape; cache keys appear only when the underlying provider
  returns them. Pinned by test_no_cache_fields_means_no_cache_keys.

* headroom/proxy/auth_mode.py: ("strands-agents/", "strands") added to
  CLIENT_UA_MAP. Production callers should also set X-Client: strands
  since the default openai-python UA carries no Strands signal.

* pyproject.toml: huggingface-hub>=1.5.0,<2.0 pinned in [ml] so a sibling
  install (e.g. strands-agents) can't drag the version below the floor
  transformers 5.x requires (otherwise Kompress silently goes
  "unavailable").

== 2. /stats MCP aggregation ==

* headroom/proxy/cost.py: _aggregate_mcp_events() reads the cross-process
  shared events file the Headroom MCP server already writes to and
  surfaces summary.mcp with three new keys:
    - compressions       (count of headroom_compress invocations)
    - tokens_removed     (sum of input - output across those)
    - retrievals         (count of headroom_retrieve — the load-bearing
                          over-compression alarm; if it grows linearly
                          with turn count, lossy compressors are
                          dropping info the model actually needs)
  Defensive on every axis — missing MCP SDK, missing file, malformed
  events, read errors — never blocks /stats.

* examples/strands_bundle_demo.py: stats panel prints the new fields so
  the demo shows the full proxy-HTTP + MCP-tool story in one view.

== 3. Codex compression-failure fail-closed protection ==

Reported by Camille (2026-05-21): Codex threads were locking with
"ran out of room in the model's context window" after Headroom's
compression timed out on an oversized response.create frame and
forwarded the original ~1.7 MB frame to the upstream, which then
rejected it. Codex's auto-compact heuristic gates on the upstream-
reported total_usage_tokens (which Headroom had been shrinking on
earlier turns), so its compaction never fired and the thread locked.

Validated against open Codex issues (CLI + Desktop share codex-rs/core):
* #16068 — confirms compaction gates on total_usage_tokens,
  estimated_token_count is computed but only logged
* #19806 — confirms image token estimator unbounded, contributes to
  the same ContextManager.get_total_token_usage → auto-compaction chain

* headroom/proxy/helpers.py: decide_compression_failure_action() with a
  unit-tested decision matrix:
    - asyncio.TimeoutError                              → refuse, always
    - non-timeout failure + frame > 256 KiB (configurable) → refuse
    - non-timeout failure + small frame                 → forward (legacy)
  Operator escape hatches:
    - HEADROOM_WS_FAIL_OPEN_ON_COMPRESSION_FAILURE=1 restores legacy
    - HEADROOM_WS_COMPRESSION_FAIL_THRESHOLD_BYTES tunes the threshold

* headroom/proxy/handlers/openai.py (WS /v1/responses): consults the
  helper after compression failure. On refuse: close client websocket
  code 1009 with "headroom: compression <reason> — please compact
  context and retry" reason; set termination_cause for the outer
  lifecycle finally; return.

* headroom/proxy/handlers/openai.py (HTTP /v1/responses): same helper.
  On refuse: raise HTTPException(413) with a structured error body so
  FastAPI's HTTPException handler emits a clean 413. The existing
  `except HTTPException: raise` guard in this handler already ensures
  the 413 propagates without being swallowed by the 502 catch-all.

Anthropic /v1/messages NOT changed in this branch: no equivalent bug
report on Anthropic-protocol clients, Claude Code (Anthropic-owned)
handles context overflow via its own cache_control/ephemeral
primitives, and Cursor/Aider don't maintain the local-Y estimate the
Codex bug requires. Deferred until a real report lands; the patch is
a one-liner reusing the same helper.

== Tests + verification ==

* tests/test_backends/test_litellm_cache_stats.py — 3 tests pinning
  cache-stat surfacing across Anthropic/OpenAI dialects + backwards-
  compat for no-cache responses.
* tests/test_proxy/test_openai_backend_path.py — 5 tests (Bedrock cache
  fields, OpenAI fallback shape, CCR intercept with provider="openai",
  CCR re-raise on exception, streaming signature contract).
* tests/test_proxy/test_mcp_stats_aggregation.py — 5 tests pinning the
  aggregator across compress+retrieve mixes, empty events, unknown event
  types, missing token fields, and read failures.
* tests/test_proxy/test_compression_failure_action.py — 12 tests pinning
  the fail-closed decision matrix (timeout always refuses, small
  transient passes through, oversize refuses, env override variants,
  custom threshold, invalid threshold falls back, 0/negative ignored).

* examples/strands_bedrock_demo.py — model_id bumped from deprecated
  Claude 3 Haiku to Sonnet 4.5 (the deprecated model now errors on
  account access).
* examples/strands_via_proxy_demo.py — proxy + Bedrock cache + streaming
  smoke test.
* examples/strands_mcp_dispatch_test.py — pure MCP round-trip probe.
* examples/strands_bundle_demo.py — full Strands + HeadroomBundle E2E
  demo (this is the shape a real Strands user copies into their app).

Full pytest: 5327 passed, 178 skipped. The previously-failing
test_core_operations.py::TestAddBatch::test_add_batch_basic passes now
that the huggingface-hub pin in pyproject.toml unblocks transformers
imports.

E2E verified live against AWS Bedrock (Sonnet 4.5):
* cache_write=10,438 on turn A → cache_read=10,438 on turn B
* streaming SSE final usage frame carries cache_read_input_tokens
* 78.7% reduction on a 50 KB JSON tool_result via SmartCrusher (
  dispatched per-content-type by ContentRouter)
* Strands Agent + HeadroomBundle: model autonomously called
  headroom_compress + headroom_retrieve via MCP; CompressionStore
  round-trip succeeded; final answer correct.
2026-05-21 11:00:14 -07:00
..
deployment/macos-launchagent fix(deployment): correct port placeholder in LaunchAgent plist template 2026-01-19 13:44:30 -08:00
langchain_demo Fix all ruff lint and format errors for CI 2026-01-10 15:33:44 -08:00
mcp_demo Fix all ruff lint and format errors for CI 2026-01-10 15:33:44 -08:00
vercel-ai-sdk-pr/node-headroom-compression/node_modules Token-level cache hit rate, compression-vs-cache tracking, dashboard SQL, security plan 2026-04-06 18:10:29 -07:00
07-context-compression.ipynb Add notebook for langchain-ai/how_to_fix_your_context PR 2026-03-26 00:25:59 -07:00
context_compression_demo.py Use realistic verbose RAG chunks in demo — triggers Kompress within-item 2026-03-26 00:18:58 -07:00
README.md feat: Add AWS Strands Agents SDK integration 2026-01-31 00:31:37 -08:00
strands_bedrock_demo.py fix(proxy): Strands MCP bundle + backend path fixes + Codex fail-closed protection 2026-05-21 11:00:14 -07:00
strands_bundle_demo.py fix(proxy): Strands MCP bundle + backend path fixes + Codex fail-closed protection 2026-05-21 11:00:14 -07:00
strands_mcp_dispatch_test.py fix(proxy): Strands MCP bundle + backend path fixes + Codex fail-closed protection 2026-05-21 11:00:14 -07:00
strands_via_proxy_demo.py fix(proxy): Strands MCP bundle + backend path fixes + Codex fail-closed protection 2026-05-21 11:00:14 -07:00
test_ccr.py Fix context-blind compression: pass user query to SmartCrusher relevance scorer 2026-03-25 23:15:19 -07:00
test_intelligent_context_toin_ccr.py docs: update documentation for IntelligentContext TOIN + CCR integration 2026-01-27 16:08:36 -08:00

Headroom Examples

This directory contains examples demonstrating Headroom's capabilities.

Quick Start Examples

basic_usage.py

Basic integration with OpenAI client:

export OPENAI_API_KEY='your-key'
python examples/basic_usage.py

anthropic_example.py

Integration with Anthropic Claude:

export ANTHROPIC_API_KEY='your-key'
python examples/anthropic_example.py

streaming_example.py

Streaming responses with optimization:

export OPENAI_API_KEY='your-key'
python examples/streaming_example.py

Evaluation Examples

smart_vs_naive_eval.py

Compare SmartCrusher against naive truncation:

export OPENAI_API_KEY='your-key'
python examples/smart_vs_naive_eval.py

real_world_eval.py

Comprehensive evaluation with Anthropic models:

export ANTHROPIC_API_KEY='your-key'
python examples/real_world_eval.py

real_world_openai_eval.py

Comprehensive evaluation with OpenAI models:

export OPENAI_API_KEY='your-key'
python examples/real_world_openai_eval.py

Demo Directories

langchain_demo/

Full LangChain agent integration demo:

# No API key needed for compression demo
PYTHONPATH=. python -m examples.langchain_demo.show_compression

# Full comparison (requires API key)
export OPENAI_API_KEY='your-key'
PYTHONPATH=. python -m examples.langchain_demo.run_comparison

See langchain_demo/README.md for details.

mcp_demo/

MCP (Model Context Protocol) integration demo:

export OPENAI_API_KEY='your-key'
PYTHONPATH=. python -m examples.mcp_demo.run_agent_eval

strands_bedrock_demo.py

AWS Strands Agents + Bedrock integration demo. Showcases two Headroom integration patterns:

  1. HeadroomHookProvider - Compresses tool outputs in real-time
  2. HeadroomStrandsModel - Optimizes entire conversation context
# Configure AWS credentials
export AWS_ACCESS_KEY_ID='your-access-key'
export AWS_SECRET_ACCESS_KEY='your-secret-key'
export AWS_DEFAULT_REGION='us-west-2'  # Optional, defaults to us-west-2

# Or use AWS profile
export AWS_PROFILE='your-profile-name'

# Run the full demo (both integration patterns)
python examples/strands_bedrock_demo.py

# Run only the hook provider demo
python examples/strands_bedrock_demo.py --hook

# Run only the model wrapper demo
python examples/strands_bedrock_demo.py --model

# Specify a different AWS region
python examples/strands_bedrock_demo.py --region us-east-1

The demo uses Claude 3 Haiku via Bedrock for cost efficiency. It creates agents with 4 tools that return verbose JSON output (search results, logs, database records, metrics) and displays compression statistics with visual comparisons.

Requirements:

  • AWS account with Bedrock enabled
  • Claude 3 Haiku model access in your region
  • pip install strands-agents headroom-ai[strands]

Running Examples

All examples can be run from the repository root:

# Install dependencies
pip install -e ".[dev]"

# Run any example
python examples/<example_name>.py

Expected Results

Example Token Savings Notes
basic_usage 50-70% Simple tool output compression
langchain_demo 70-85% Real agent with multiple tools
mcp_demo 60-80% MCP tool outputs
strands_bedrock_demo 60-85% Strands + Bedrock with verbose tools
real_world_eval 50-90% Varies by scenario

Troubleshooting

ModuleNotFoundError: No module named 'headroom'

Run from the repository root with PYTHONPATH:

PYTHONPATH=. python examples/basic_usage.py

Or install in development mode:

pip install -e .

API Key Errors

Ensure your API keys are set:

export OPENAI_API_KEY='sk-...'
export ANTHROPIC_API_KEY='sk-ant-...'

AWS Credentials Errors (for Strands demo)

Ensure AWS credentials are configured:

# Option 1: Environment variables
export AWS_ACCESS_KEY_ID='your-access-key'
export AWS_SECRET_ACCESS_KEY='your-secret-key'

# Option 2: AWS profile
export AWS_PROFILE='your-profile-name'

# Option 3: AWS credentials file (~/.aws/credentials)

Also ensure Bedrock and the Claude 3 Haiku model are enabled in your AWS account.