fix(proxy/cost): price cache savings by most-used model, not first-seen (#2023)

## Description

`build_prefix_cache_stats` (`headroom/proxy/cost.py`) values each
provider's cache-read savings
using a single "base input price per token". It derives that price by
scanning
`cost_tracker._tokens_sent_by_model` and **breaking on the first**
provider-matching model that
has a price — even though the comment says "most-used model":

```python
# Get the base input price per token for the most-used model on this provider
input_price_per_token = None
if cost_tracker:
    for model_name in cost_tracker._tokens_sent_by_model:   # insertion order, NOT usage order
        ...
        if is_match:
            price_per_1m = cost_tracker._get_list_price(model_name)
            if price_per_1m:
                input_price_per_token = price_per_1m / 1_000_000
                break                                        # first match wins
```

`_tokens_sent_by_model` is insertion-ordered, so the price used depends
on which model was
*recorded first*, not on usage volume. A Claude Code session sends both
Sonnet (main loop) and
Haiku (titles/subagents). If Haiku ($0.80/M) was seen before Sonnet
($3/M), **all** of the
provider's cache-read savings are priced at Haiku's rate — understating
the dashboard's cache
savings by ~3.75×. Reverse the order and it overstates.

Closes: no issue filed — found while auditing the cache-savings pricing.

## Fix

Pick the provider-matching, priced model with the **highest token
volume** instead of breaking
on the first match:

```python
best_tokens = -1
for model_name, tokens_sent in cost_tracker._tokens_sent_by_model.items():
    if is_match and tokens_sent > best_tokens:
        price_per_1m = cost_tracker._get_list_price(model_name)
        if price_per_1m:
            input_price_per_token = price_per_1m / 1_000_000
            best_tokens = tokens_sent
```

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)

## Changes Made

- `headroom/proxy/cost.py`: select the highest-volume provider-matching
model (with a known price) rather than the first-recorded one.
- `tests/test_proxy_cache_ttl_metrics.py`: add
`test_prefix_cache_stats_prices_by_most_used_model` using real distinct
per-model prices. (The existing cache-stats tests monkeypatch
`_get_list_price` to a constant `100.0`, which masked the
model-selection logic — hence the bug slipped through.)

## Testing

- [x] New regression test added
(`tests/test_proxy_cache_ttl_metrics.py`)
- [x] Linting/formatting clean — run with the CI-pinned `ruff==0.15.17`
- [ ] Full `pytest` deferred to CI (local-OOM reason below).

```text
$ uvx ruff@0.15.17 check headroom/proxy/cost.py tests/test_proxy_cache_ttl_metrics.py
All checks passed!
```

## Real Behavior Proof

- Environment: Windows 11, Python 3.10, headroom from this branch.
Importing `headroom` pulls in the torch/transformers stack and a full
`pytest` gets OOM-killed on this box, so I verified the selection logic
with a dependency-free script and left the full pytest to CI.
- Exact command / steps: ran a `{haiku: 500, sonnet: 50000}` token map
(Haiku recorded first, Sonnet the higher volume) through both the old
first-match and new highest-volume selection with real prices.
- Observed result: the old logic picks Haiku's $0.80/M (first-inserted);
the new logic picks Sonnet's $3/M (highest volume) and is
insertion-order independent:

```text
OLD picks Haiku price: 0.80/M  (first-inserted)
NEW picks Sonnet price: 3.00/M  (highest volume)
  -> old understates the input price by 3.75x (3.75x)
NEW is insertion-order independent
COST MOST-USED-MODEL FIX VERIFIED
```

- Not tested: rendering the live dashboard (needs the running app). The
fix is confined to the price-selection loop and the new test drives
`build_prefix_cache_stats` directly. Full local `pytest` deferred to CI
(OOM, per above).

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [ ] New and existing unit tests pass locally with my changes — ran
lint + a standalone logic check; full pytest deferred to CI (local OOM,
disclosed above)
- [x] I have updated the CHANGELOG.md if applicable

## Additional Notes

- No new dependencies; a single-loop change plus a test with realistic
prices.
- @JerrettDavis tagging you — this skews the dashboard's per-provider
cache-savings dollar figure by the ratio between a provider's models
(≈3.75× for Sonnet/Haiku), so it seemed worth surfacing. Thanks!

---------

Co-authored-by: JerrettDavis <mxjerrett@gmail.com>
Co-authored-by: Tejas Chopra <chopratejas@gmail.com>
This commit is contained in:
Abhay Singh 2026-07-13 19:07:58 +05:30 committed by GitHub
parent 415e03c168
commit b4f807f21a
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
3 changed files with 54 additions and 4 deletions

View file

@ -108,6 +108,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
* **proxy/vertex:** route Vertex `publisher=google` (Gemini) requests to the region matching the request path. `vertex_generate_content`, `vertex_stream_generate_content`, and `vertex_count_tokens` discarded the path's `location` and forwarded to the single fixed host from `_api_target(proxy, "vertex")` (default `us-central1`), instead of the region-aware `_vertex_target_for_location` the sibling Anthropic `rawPredict` route already uses. So a request to `.../locations/europe-west1/publishers/google/...` was sent to a `us-central1` host, which Vertex rejects on the region/host mismatch. The three google routes now derive the host from the request's `location` (operator-pinned upstreams are still honored). * **proxy/vertex:** route Vertex `publisher=google` (Gemini) requests to the region matching the request path. `vertex_generate_content`, `vertex_stream_generate_content`, and `vertex_count_tokens` discarded the path's `location` and forwarded to the single fixed host from `_api_target(proxy, "vertex")` (default `us-central1`), instead of the region-aware `_vertex_target_for_location` the sibling Anthropic `rawPredict` route already uses. So a request to `.../locations/europe-west1/publishers/google/...` was sent to a `us-central1` host, which Vertex rejects on the region/host mismatch. The three google routes now derive the host from the request's `location` (operator-pinned upstreams are still honored).
* **proxy/anthropic:** give each Anthropic conversation its own session id. `SessionTrackerStore.compute_session_id` derived its fallback id from `model` + system text harvested only from `role:"system"` entries inside `messages` — but Anthropic carries the system prompt as a top-level `body["system"]` field, so genuine Anthropic requests (which never carry `x-headroom-session-id`) collapsed to `md5(model:[])` and every conversation on the same model shared one `PrefixCacheTracker`. That let session-sticky state cross-contaminate: conversation A's sticky `headroom_retrieve`/memory tools and `anthropic-beta` headers were injected into conversation B, and frozen-prefix/compression-cache state mixed across conversations. The Anthropic handler now folds the top-level `system` into the session-id inputs (prepending a synthetic `role:"system"` message used only to derive the id), giving distinct conversations distinct ids. * **proxy/anthropic:** give each Anthropic conversation its own session id. `SessionTrackerStore.compute_session_id` derived its fallback id from `model` + system text harvested only from `role:"system"` entries inside `messages` — but Anthropic carries the system prompt as a top-level `body["system"]` field, so genuine Anthropic requests (which never carry `x-headroom-session-id`) collapsed to `md5(model:[])` and every conversation on the same model shared one `PrefixCacheTracker`. That let session-sticky state cross-contaminate: conversation A's sticky `headroom_retrieve`/memory tools and `anthropic-beta` headers were injected into conversation B, and frozen-prefix/compression-cache state mixed across conversations. The Anthropic handler now folds the top-level `system` into the session-id inputs (prepending a synthetic `role:"system"` message used only to derive the id), giving distinct conversations distinct ids.
* **cache/semantic:** key entries by the full-context hash, not the trailing query text. `SemanticCache.put` stored each response under `sha256(query)[:16]` where `query` is only the last user message, and the exact-match branch of `get` returned the slot without checking the stored entry's `messages_hash`. Two requests that share a trailing message ("continue", "yes", "run the tests") but differ in earlier context therefore collided on one slot — the second overwrote the first, and the first's hash then resolved to the second's cached response (wrong data served). Entries are now keyed by `messages_hash` when present, and `get` verifies `entry.messages_hash` before returning. * **cache/semantic:** key entries by the full-context hash, not the trailing query text. `SemanticCache.put` stored each response under `sha256(query)[:16]` where `query` is only the last user message, and the exact-match branch of `get` returned the slot without checking the stored entry's `messages_hash`. Two requests that share a trailing message ("continue", "yes", "run the tests") but differ in earlier context therefore collided on one slot — the second overwrote the first, and the first's hash then resolved to the second's cached response (wrong data served). Entries are now keyed by `messages_hash` when present, and `get` verifies `entry.messages_hash` before returning.
* **proxy/cost:** value prefix-cache savings with the most-used model's price, not the first-recorded one. `build_prefix_cache_stats` scanned `cost_tracker._tokens_sent_by_model` and broke on the *first* provider-matching model with a price — despite the "most-used model" comment — so a Claude Code session (Sonnet for the main loop, Haiku for titles/subagents) priced all of a provider's cache-read savings at whichever model happened to be recorded first. If Haiku ($0.80/M) came before Sonnet ($3/M), the dashboard understated cache savings ~3.75x (and vice-versa). It now picks the provider-matching, priced model with the highest token volume.
* **proxy/openai:** stop overriding an explicit client `stream_options.include_usage` on the streaming chat path. To count tokens from the trailing usage chunk, the handler set `include_usage: True` unconditionally — including flipping an explicit client `false` to `true`. The upstream then appended a usage-only chunk (`choices: []`) the client never requested, and the common `chunk.choices[0].delta` loop raised `IndexError`. The option is now only filled in when the client left the choice open (no `stream_options`, or a dict without `include_usage`); an explicit `true`/`false` is respected. * **proxy/openai:** stop overriding an explicit client `stream_options.include_usage` on the streaming chat path. To count tokens from the trailing usage chunk, the handler set `include_usage: True` unconditionally — including flipping an explicit client `false` to `true`. The upstream then appended a usage-only chunk (`choices: []`) the client never requested, and the common `chunk.choices[0].delta` loop raised `IndexError`. The option is now only filled in when the client left the choice open (no `stream_options`, or a dict without `include_usage`); an explicit `true`/`false` is respected.
* **proxy/openai:** stop PRE_SEND from reintroducing `tools: []` after the direct #728 fix. The OpenAI request handler now mirrors the existing `tools or _original_tools is not None` body-write guard during PRE_SEND write-back, so providers that reject empty tool arrays no longer see a tools field when the client omitted it, while explicit client `tools: []` remains preserved ([#1983](https://github.com/headroomlabs-ai/headroom/issues/1983)). * **proxy/openai:** stop PRE_SEND from reintroducing `tools: []` after the direct #728 fix. The OpenAI request handler now mirrors the existing `tools or _original_tools is not None` body-write guard during PRE_SEND write-back, so providers that reject empty tool arrays no longer see a tools field when the client omitted it, while explicit client `tools: []` remains preserved ([#1983](https://github.com/headroomlabs-ai/headroom/issues/1983)).
* **proxy/openai:** keep the exact Responses function name `terminal` resident during OpenAI tool-search deferral so cache-mode optimization stops forwarding `terminal.terminal` and triggering the reserved-namespace 400 on Codex Responses ([#1946](https://github.com/headroomlabs-ai/headroom/issues/1946)). * **proxy/openai:** keep the exact Responses function name `terminal` resident during OpenAI tool-search deferral so cache-mode optimization stops forwarding `terminal.terminal` and triggering the reserved-namespace 400 on Codex Responses ([#1946](https://github.com/headroomlabs-ai/headroom/issues/1946)).

View file

@ -139,10 +139,17 @@ def build_prefix_cache_stats(
read_mult: float = econ["read_multiplier"] # type: ignore[assignment] read_mult: float = econ["read_multiplier"] # type: ignore[assignment]
write_mult: float = econ["write_multiplier"] # type: ignore[assignment] write_mult: float = econ["write_multiplier"] # type: ignore[assignment]
# Get the base input price per token for the most-used model on this provider # Get the base input price per token for the most-used model on this
# provider. Pick the provider-matching, priced model with the highest
# token volume — not the first one recorded. A Claude Code session sends
# both Sonnet (main loop) and Haiku (titles/subagents); breaking on the
# first-inserted model would price all cache savings at whichever happened
# to be seen first (e.g. Haiku's $0.80/M vs Sonnet's $3/M), skewing the
# dashboard's savings figure ~3.75x.
input_price_per_token = None input_price_per_token = None
if cost_tracker: if cost_tracker:
for model_name in cost_tracker._tokens_sent_by_model: best_tokens = -1
for model_name, tokens_sent in cost_tracker._tokens_sent_by_model.items():
# Match model to provider # Match model to provider
_openai_prefixes = ("gpt", "o1", "o3", "o4") _openai_prefixes = ("gpt", "o1", "o3", "o4")
is_match = ( is_match = (
@ -151,11 +158,11 @@ def build_prefix_cache_stats(
or (provider == "gemini" and "gemini" in model_name) or (provider == "gemini" and "gemini" in model_name)
or (provider == "bedrock" and "claude" in model_name) or (provider == "bedrock" and "claude" in model_name)
) )
if is_match: if is_match and tokens_sent > best_tokens:
price_per_1m = cost_tracker._get_list_price(model_name) price_per_1m = cost_tracker._get_list_price(model_name)
if price_per_1m: if price_per_1m:
input_price_per_token = price_per_1m / 1_000_000 input_price_per_token = price_per_1m / 1_000_000
break best_tokens = tokens_sent
# Calculate savings: # Calculate savings:
# Cache reads save (1.0 - read_mult) per token vs uncached input price. # Cache reads save (1.0 - read_mult) per token vs uncached input price.

View file

@ -110,6 +110,48 @@ def test_prefix_cache_stats_subtracts_write_premium_from_provider_net_savings(
assert anthropic["net_savings_usd"] == 0.0021 assert anthropic["net_savings_usd"] == 0.0021
def test_prefix_cache_stats_prices_by_most_used_model(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""Cache-read savings must be valued with the highest-volume model's price,
not whichever model was recorded first. A Claude Code session sends Haiku
(titles) and Sonnet (main loop); pricing the savings at Haiku's rate because
it was inserted first understates the dashboard figure ~3.75x."""
prices = {"claude-haiku-4-5": 0.80, "claude-sonnet-4-5": 3.00}
monkeypatch.setattr(CostTracker, "_get_list_price", lambda _self, m: prices.get(m))
def _savings(tokens_by_model: dict[str, int]) -> float:
metrics = PrometheusMetrics()
# Use a large read count so the reported savings_usd (rounded to 4 dp)
# stays exact and the price ratio is not lost to rounding.
metrics.cache_by_provider["anthropic"].update(
{
"requests": 1,
"hit_requests": 1,
"cache_read_tokens": 1_000_000,
"cache_write_tokens": 0,
"cache_write_5m_tokens": 0,
"cache_write_1h_tokens": 0,
"cache_write_5m_requests": 0,
"cache_write_1h_requests": 0,
}
)
tracker = CostTracker()
tracker._tokens_sent_by_model.update(tokens_by_model)
stats = build_prefix_cache_stats(metrics, tracker)
return stats["by_provider"]["anthropic"]["savings_usd"]
# Haiku recorded first, but Sonnet carries the higher token volume.
haiku_first = _savings({"claude-haiku-4-5": 500, "claude-sonnet-4-5": 50_000})
sonnet_only = _savings({"claude-sonnet-4-5": 50_000})
haiku_only = _savings({"claude-haiku-4-5": 500})
# Priced by Sonnet regardless of insertion order, not by first-seen Haiku.
assert haiku_first == sonnet_only
assert haiku_first > haiku_only
assert haiku_only == pytest.approx(sonnet_only * 0.80 / 3.00)
def test_prefix_cache_stats_subtracts_write_premium_from_total_net_savings( def test_prefix_cache_stats_subtracts_write_premium_from_total_net_savings(
monkeypatch: pytest.MonkeyPatch, monkeypatch: pytest.MonkeyPatch,
) -> None: ) -> None: