mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
## Description The whole-request savings ratios in `/stats` (`proxy_savings_percent`, `savings_percent`) divide by a per-request recount of the full transcript: a session at turn 200 has had its history counted 200 times into the denominator. Long-running cached sessions — 1M-context models especially, since they never compact — therefore read as ~0% savings no matter how well compression performs on content that actually newly enters context. Field example that motivated this: one day of 1M-context Claude Code traffic saved 641K tokens against ~13.4M tokens of genuinely new content (~4.8%), but displayed as 0.14% because the summed full-transcript denominator was 475M. This PR adds a new-content-relative rate alongside the existing fields: - `tokens.new_input_tokens` — provider-billed non-cache-read input (uncached + cache-write tokens, summed from response usage across providers; the cache accumulators already track both). - `tokens.new_input_savings_percent` — `saved / (new_input + saved)`. Tokens Headroom removed never reached the provider, so they're added back to form the baseline: "of the input that would have newly entered context, what fraction did Headroom remove?" Purely additive — no existing field changes, no new accumulators. ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `headroom/proxy/server.py`: compute `new_input_tokens` from `prefix_cache_stats["totals"]` (already built for `/stats`) and emit the two new fields in the `tokens` block. Rate is guarded on `new_input_tokens > 0`: the cache accumulators only see requests with cache activity, so a deployment with no cache metrics (e.g. Bedrock) would otherwise divide savings by themselves and report ~100% — it reports 0 instead. - `tests/test_stats_new_input_savings_rate.py`: endpoint-level tests via `TestClient(create_app(...))` — a long-cached-session request shows 9.09% new-content rate while `proxy_savings_percent` stays diluted at 0.5%; and the no-cache-usage-data case reports 0. - `CHANGELOG.md`: Features entry. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [ ] Manual testing performed ### Test Output ```text $ uv run --frozen --extra dev pytest tests/test_stats_new_input_savings_rate.py -v tests/test_stats_new_input_savings_rate.py::test_stats_reports_new_input_savings_rate PASSED tests/test_stats_new_input_savings_rate.py::test_stats_new_input_rate_is_zero_without_cache_usage_data PASSED ========================= 2 passed, 1 warning in 6.78s ========================= $ uv run --frozen --extra dev pytest tests/test_proxy_savings_history.py tests/test_dashboard_token_savings.py tests/test_proxy_cache_ttl_metrics.py ======================== 57 passed, 1 warning in 10.70s ======================== $ uv run --frozen --extra dev mypy headroom/proxy/server.py Success: no issues found in 1 source file $ ruff check headroom/proxy/server.py tests/test_stats_new_input_savings_rate.py All checks passed! $ ruff format --check headroom/proxy/server.py tests/test_stats_new_input_savings_rate.py 2 files already formatted ``` ## Real Behavior Proof - Environment: macOS 15 (darwin 24.6.0), Python 3.10 via `uv run --frozen --extra dev`. - Exact command / steps: `TestClient(create_app(config))`, record a request shaped like a late turn of a long cached session (`input_tokens=1_000_000, tokens_saved=5_000, cache_read=900_000, cache_write=45_000, uncached=5_000`), then `GET /stats`. - Observed result: `tokens.new_input_tokens == 50_000`, `tokens.new_input_savings_percent == 9.09`, while `proxy_savings_percent` stays `0.5` — the dilution the new field exists to correct, reproduced side by side. - Not tested: not run against a live proxy with real provider traffic; `ruff`/`mypy` run scoped to the changed files rather than the whole repo. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Screenshots (if applicable) N/A — JSON API addition; dashboard adoption can follow separately. ## Additional Notes - No linked issue; companion to the nested tool_result image token-counting fix (same investigation — that PR fixes the inflated numerator/denominator counts, this one fixes the metric that divides by transcript recounts). - Caveat worth a reviewer's eye: the numerator (`tokens_saved_total`, local tokenizer) and denominator (provider-reported usage) come from different counters. They're on the same scale, but the rate is honest-approximate rather than exact — comment in code says so. - Deliberately did not change the dashboard headline or any existing field semantics; consumers can opt into the new rate. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: JerrettDavis <mxjerrett@gmail.com>
81 lines
3 KiB
Python
81 lines
3 KiB
Python
"""New-content-relative savings rate in /stats (tokens.new_input_savings_percent).
|
|
|
|
The whole-request ratios recount the full transcript on every turn, so long
|
|
cached sessions dilute toward 0% regardless of how well compression performs
|
|
on content that newly enters context. The new rate divides by provider-billed
|
|
non-cache-read input (uncached + cache-write) plus the tokens compression
|
|
removed before they could be billed.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import asyncio
|
|
|
|
from fastapi.testclient import TestClient
|
|
|
|
from headroom.proxy.server import ProxyConfig, create_app
|
|
|
|
|
|
def _make_client(tmp_path, monkeypatch) -> TestClient:
|
|
monkeypatch.setenv("HEADROOM_SAVINGS_PATH", str(tmp_path / "proxy_savings.json"))
|
|
config = ProxyConfig(
|
|
cache_enabled=False,
|
|
rate_limit_enabled=False,
|
|
log_requests=False,
|
|
)
|
|
return TestClient(create_app(config))
|
|
|
|
|
|
def test_stats_reports_new_input_savings_rate(tmp_path, monkeypatch):
|
|
with _make_client(tmp_path, monkeypatch) as client:
|
|
proxy = client.app.state.proxy
|
|
# A late turn of a long cached session: the local transcript recount
|
|
# (input_tokens) dwarfs what the provider newly billed (uncached +
|
|
# cache_write = 50k), so the whole-request ratio dilutes to ~0.5%
|
|
# while the new-content rate reports the undiluted 9.09%.
|
|
asyncio.run(
|
|
proxy.metrics.record_request(
|
|
provider="anthropic",
|
|
model="claude-opus-4-6",
|
|
input_tokens=1_000_000,
|
|
output_tokens=200,
|
|
tokens_saved=5_000,
|
|
latency_ms=10.0,
|
|
cache_read_tokens=900_000,
|
|
cache_write_tokens=45_000,
|
|
uncached_input_tokens=5_000,
|
|
)
|
|
)
|
|
|
|
stats = client.get("/stats")
|
|
assert stats.status_code == 200
|
|
tokens = stats.json()["tokens"]
|
|
|
|
assert tokens["new_input_tokens"] == 50_000
|
|
# 5_000 saved / (50_000 billed-new + 5_000 saved) = 9.09%
|
|
assert tokens["new_input_savings_percent"] == 9.09
|
|
# The transcript-diluted ratio stays as-is — the new rate sits alongside,
|
|
# it does not replace existing fields.
|
|
assert tokens["proxy_savings_percent"] == 0.5
|
|
|
|
|
|
def test_stats_new_input_rate_is_zero_without_cache_usage_data(tmp_path, monkeypatch):
|
|
with _make_client(tmp_path, monkeypatch) as client:
|
|
proxy = client.app.state.proxy
|
|
# Savings recorded but no cache usage observed (provider without
|
|
# cache metrics): the rate must report 0, not savings/savings=100%.
|
|
asyncio.run(
|
|
proxy.metrics.record_request(
|
|
provider="bedrock",
|
|
model="claude-opus-4-6",
|
|
input_tokens=10_000,
|
|
output_tokens=200,
|
|
tokens_saved=2_000,
|
|
latency_ms=10.0,
|
|
)
|
|
)
|
|
|
|
tokens = client.get("/stats").json()["tokens"]
|
|
|
|
assert tokens["new_input_tokens"] == 0
|
|
assert tokens["new_input_savings_percent"] == 0
|