## Description
<!-- Briefly explain the change and why it is needed. -->
The metrics docs describe the `headroom_*` Prometheus metric family and
suggest example Grafana panels, but ship no importable dashboard — users
have to build one by hand. This adds a ready-to-import Grafana dashboard
built **only** on documented metric names (`headroom_requests_total`,
`headroom_tokens_saved_total`, `headroom_tokens_input_total`, and the
`headroom_overhead_ms_*` millisecond summary), and links it from the
**Grafana Dashboard** section of `docs/content/docs/metrics.mdx`.
This is a docs/examples-only addition — no source code changes.
Closes #
## Type of Change
- [ ] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [x] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Added `examples/grafana/headroom-dashboard.json` — a ready-to-import
Grafana dashboard (7 panels, uid `headroom-compression`) built entirely
on Headroom's documented `/metrics` names. Panels cover tokens saved,
input tokens, request rate, average processing overhead
(`headroom_overhead_ms_sum` / `headroom_overhead_ms_count` with
min/max), tokens-saved/sec, and request rate by pool. It uses **no
histograms** (the proxy emits none). The `pool`/`source` template
variables use regex matchers (`=~`) so they are optional and match
series without those labels.
- Updated `docs/content/docs/metrics.mdx` — linked the new dashboard
from the **Grafana Dashboard** section with import instructions, keeping
the existing ad-hoc PromQL query table alongside it.
## Testing
<!-- Check what you actually ran, then paste the real command output
below. -->
- [ ] Unit tests pass (`pytest`)
- [ ] Linting passes (`ruff check .`)
- [ ] Type checking passes (`mypy headroom`)
- [ ] New tests added for new functionality
- [x] Manual testing performed
Docs/examples-only change, manually verified: the dashboard JSON is
well-formed and every PromQL query references only the documented
`headroom_*` metric names from `docs/content/docs/metrics.mdx`.
### Test Output
```text
$ python3 -c "import json; d=json.load(open('examples/grafana/headroom-dashboard.json')); print('valid JSON,', len(d['panels']), 'panels, uid', d['uid'])"
valid JSON, 7 panels, uid headroom-compression
```
PromQL queries used by the panels (all against documented `headroom_*`
metrics):
```text
sum(headroom_tokens_saved_total{pool=~"$pool", hook=~"$hook"})
sum(headroom_tokens_input_total{pool=~"$pool", hook=~"$hook"})
sum(rate(headroom_requests_total{pool=~"$pool", hook=~"$hook"}[$__rate_interval]))
sum(rate(headroom_overhead_ms_sum{pool=~"$pool", hook=~"$hook"}[$__rate_interval])) / clamp_min(sum(rate(headroom_overhead_ms_count{pool=~"$pool", hook=~"$hook"}[$__rate_interval])), 1)
sum(rate(headroom_tokens_saved_total{pool=~"$pool", hook=~"$hook"}[$__rate_interval])) by (pool)
max(headroom_overhead_ms_max{pool=~"$pool", hook=~"$hook"})
min(headroom_overhead_ms_min{pool=~"$pool", hook=~"$hook"})
sum(rate(headroom_requests_total{pool=~"$pool", hook=~"$hook"}[$__rate_interval])) by (pool)
```
## Real Behavior Proof
- Environment: local checkout of the PR branch; Python 3 for JSON
validation.
- Exact command / steps: ran the JSON-validation command above (see Test
Output) — parses cleanly, reports 7 panels and uid
`headroom-compression`; then read every panel target and confirmed each
PromQL query references only metric names documented in
`docs/content/docs/metrics.mdx` (`headroom_requests_total`,
`headroom_tokens_saved_total`, `headroom_tokens_input_total`,
`headroom_overhead_ms_{sum,count,min,max}`). No histogram metrics are
referenced.
- Observed result: JSON is valid and importable via Grafana's
**Dashboards → New → Import → Upload**; no datasource UID is hard-coded,
so the importer prompts for a Prometheus datasource. Queries match the
documented metric family.
- Not tested: a full live Grafana import against a running proxy
scraping real `/metrics` was not performed in CI. Verification was
limited to JSON validity and query/metric-name correctness against the
documented metrics.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
Additive docs/examples only — no source code, tests, or runtime behavior
changed.
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [ ] I have added tests that prove my fix is effective or that my
feature works
- [ ] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable
## Screenshots (if applicable)
N/A — dashboard is imported from JSON; see the PromQL and panel list
above.
## Additional Notes
<!-- Mention any N/A checklist items, tradeoffs, follow-ups, or
maintainer context. -->
Test-related checklist items are N/A: this is an additive docs/examples
change with no application code, so `pytest`/`mypy`/`ruff` and new unit
tests do not apply. The dashboard JSON was validated and its queries
checked against the documented metric names instead.
---------
Co-authored-by: Tejas Chopra <chopratejas@gmail.com>
|
||
|---|---|---|
| .. | ||
| deployment/macos-launchagent | ||
| grafana | ||
| langchain_demo | ||
| mcp_demo | ||
| 07-context-compression.ipynb | ||
| context_compression_demo.py | ||
| README.md | ||
| strands_bedrock_demo.py | ||
| strands_bundle_demo.py | ||
| strands_mcp_dispatch_test.py | ||
| strands_via_proxy_demo.py | ||
| tabular_compression_demo.py | ||
| test_ccr.py | ||
Headroom Examples
This directory contains examples demonstrating Headroom's capabilities.
Quick Start Examples
basic_usage.py
Basic integration with OpenAI client:
export OPENAI_API_KEY='your-key'
python examples/basic_usage.py
anthropic_example.py
Integration with Anthropic Claude:
export ANTHROPIC_API_KEY='your-key'
python examples/anthropic_example.py
streaming_example.py
Streaming responses with optimization:
export OPENAI_API_KEY='your-key'
python examples/streaming_example.py
tabular_compression_demo.py
Tabular + spreadsheet compression on generated sample data (no API key needed).
Shows where CSV/markdown tables and .xlsx workbooks compress and where compact,
all-unique data correctly passes through:
python examples/tabular_compression_demo.py # run all scenarios
python examples/tabular_compression_demo.py --write DIR # also save the sample files
Evaluation Examples
smart_vs_naive_eval.py
Compare SmartCrusher against naive truncation:
export OPENAI_API_KEY='your-key'
python examples/smart_vs_naive_eval.py
real_world_eval.py
Comprehensive evaluation with Anthropic models:
export ANTHROPIC_API_KEY='your-key'
python examples/real_world_eval.py
real_world_openai_eval.py
Comprehensive evaluation with OpenAI models:
export OPENAI_API_KEY='your-key'
python examples/real_world_openai_eval.py
Demo Directories
langchain_demo/
Full LangChain agent integration demo:
# No API key needed for compression demo
PYTHONPATH=. python -m examples.langchain_demo.show_compression
# Full comparison (requires API key)
export OPENAI_API_KEY='your-key'
PYTHONPATH=. python -m examples.langchain_demo.run_comparison
See langchain_demo/README.md for details.
mcp_demo/
MCP (Model Context Protocol) integration demo:
export OPENAI_API_KEY='your-key'
PYTHONPATH=. python -m examples.mcp_demo.run_agent_eval
strands_bedrock_demo.py
AWS Strands Agents + Bedrock integration demo. Showcases two Headroom integration patterns:
- HeadroomHookProvider - Compresses tool outputs in real-time
- HeadroomStrandsModel - Optimizes entire conversation context
# Configure AWS credentials
export AWS_ACCESS_KEY_ID='your-access-key'
export AWS_SECRET_ACCESS_KEY='your-secret-key'
export AWS_DEFAULT_REGION='us-west-2' # Optional, defaults to us-west-2
# Or use AWS profile
export AWS_PROFILE='your-profile-name'
# Run the full demo (both integration patterns)
python examples/strands_bedrock_demo.py
# Run only the hook provider demo
python examples/strands_bedrock_demo.py --hook
# Run only the model wrapper demo
python examples/strands_bedrock_demo.py --model
# Specify a different AWS region
python examples/strands_bedrock_demo.py --region us-east-1
The demo uses Claude 3 Haiku via Bedrock for cost efficiency. It creates agents with 4 tools that return verbose JSON output (search results, logs, database records, metrics) and displays compression statistics with visual comparisons.
Requirements:
- AWS account with Bedrock enabled
- Claude 3 Haiku model access in your region
pip install strands-agents headroom-ai[strands]
Running Examples
All examples can be run from the repository root:
# Install dependencies
pip install -e ".[dev]"
# Run any example
python examples/<example_name>.py
Expected Results
| Example | Token Savings | Notes |
|---|---|---|
| basic_usage | 50-70% | Simple tool output compression |
| langchain_demo | 70-85% | Real agent with multiple tools |
| mcp_demo | 60-80% | MCP tool outputs |
| strands_bedrock_demo | 60-85% | Strands + Bedrock with verbose tools |
| real_world_eval | 50-90% | Varies by scenario |
Troubleshooting
ModuleNotFoundError: No module named 'headroom'
Run from the repository root with PYTHONPATH:
PYTHONPATH=. python examples/basic_usage.py
Or install in development mode:
pip install -e .
API Key Errors
Ensure your API keys are set:
export OPENAI_API_KEY='sk-...'
export ANTHROPIC_API_KEY='sk-ant-...'
AWS Credentials Errors (for Strands demo)
Ensure AWS credentials are configured:
# Option 1: Environment variables
export AWS_ACCESS_KEY_ID='your-access-key'
export AWS_SECRET_ACCESS_KEY='your-secret-key'
# Option 2: AWS profile
export AWS_PROFILE='your-profile-name'
# Option 3: AWS credentials file (~/.aws/credentials)
Also ensure Bedrock and the Claude 3 Haiku model are enabled in your AWS account.