mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
Two coupled defects on Anthropic's 1M-context tier: Headroom **under-budgeted** those sessions and **under-priced** them by ~2x. The second gets worse once the first is fixed, so they ship together. --- # Part 1 — `[1m]` was lost before the context budget was sized `sanitize_anthropic_model_id()` strips a trailing `[1m]`, which is correct for the wire — upstream Anthropic rejects the suffix, and #2027 added the strip for exactly that reason. But `[1m]` is not only an ANSI artifact. Claude Code appends it to a model id to request the **1M context tier**, and only sends the `context-1m` beta header when it is present (#1158 — what `headroom wrap claude --1m` sets up). `get_context_limit()` sanitized *before* resolving, so the tier was gone by lookup time: ```python provider.get_context_limit("claude-sonnet-4-5[1m]") # 200_000 ← real window is 1M ``` The request still reached Anthropic correctly and still got a 1M window — the beta header goes through untouched. What broke is our **budget**: Headroom sized a 1M session at 200K and began compacting at a fifth of the available room. Models whose base is already 1M (`claude-opus-5`, `claude-sonnet-5`) resolved to 1M either way, which is why this went unnoticed. It bites the Sonnet 4 / 4.5 family — the models `[1m]` exists for. **Fix:** read the tier off the id *before* sanitizing; raise the resolved limit to at least 1M. `max()` rather than assignment, so a base wider than 1M keeps its own window. Detection is deliberately narrower than the sanitizer — only a literal `[1m]`; `[0m]`, `[1;32m]` and real `ESC[` sequences still strip without promoting. | model id | wire id (unchanged) | limit before | limit after | |---|---|---|---| | `claude-sonnet-4-5` | `claude-sonnet-4-5` | 200K | 200K | | `claude-sonnet-4-5[1m]` | `claude-sonnet-4-5` | **200K** | **1M** | | `claude-opus-5[1m]` | `claude-opus-5` | 1M | 1M | | `claude-sonnet-4-5[0m]` | `claude-sonnet-4-5` | 200K | 200K | | `ESC[1m claude-sonnet-4-5 ESC[0m` | `claude-sonnet-4-5` | 200K | 200K | The wire id is unchanged in every case, so #2027 holds — guarded by a regression test. --- # Part 2 — the pricing that reports those sessions was wrong ### 2a. The LiteLLM cost path was dead in every provider `litellm.completion_cost()` no longer accepts `prompt_tokens` / `completion_tokens`. Every call raised `TypeError`: ``` TypeError: completion_cost() got an unexpected keyword argument 'prompt_tokens' ``` All five providers — `anthropic`, `openai`, `google`, `cohere`, `litellm` — caught it with a bare `except` and silently fell through to their hand-maintained tables. The "up-to-date pricing from LiteLLM" the docstrings promise **has not run at all**. Anthropic additionally passed `input_tokens - cached_tokens`, the wrong convention (LiteLLM expects the cache-inclusive total), which would also have suppressed the long-context threshold even had the call worked. Replaced with `litellm.cost_per_token()` behind one shared helper, `pricing.litellm_pricing.estimate_cost_from_tokens()`, which reuses the existing gateway-alias candidate chain and returns `None` (not an exception) when LiteLLM can't price a model. ### 2b. Neither path applied Anthropic's long-context premium On the Sonnet 4 / 4.5 family a prompt over 200K re-prices the **whole** request — input 2×, output 1.5×, cache 2× — not just the tokens past the threshold. Rates confirmed from LiteLLM's `*_above_200k_tokens` fields. | request (`claude-sonnet-4-5`) | reported before | true | error | |---|---|---|---| | 100K in / 5K out | $0.3750 | $0.3750 | — | | 300K in / 5K out | $0.9750 | **$1.9125** | −49% | | 300K in (150K cached) / 5K out | $0.5700 | **$1.1025** | −48% | LiteLLM applies this itself once the call works. The manual fallback needed `_apply_long_context_premium()` — the LiteLLM dependency is gated `python_version < '3.14'`, so on 3.14 the fallback is the *only* path. **Both paths now agree to four decimal places on every case under test.** --- ## What I checked and did *not* change The fork report that prompted this claimed the Anthropic tables were materially stale ("Opus 4.x priced wrong"). **That does not hold.** I audited every entry against LiteLLM's vendored table: - **Anthropic** — every model LiteLLM knows matches exactly, Opus 4.x included. - **OpenAI** — all 17 entries match; the two that don't resolve are retired models. The defect was the mechanism, not the numbers, so the rate cards are untouched. One thing the repaired path fixes for free: OpenAI's cached-input discount is **50% on gpt-4o, 75% on gpt-4.1, 90% on gpt-5**, but the manual path applies a flat 50% estimate. With LiteLLM live, real per-model rates are used. The flat estimate remains only as the offline fallback. ## Scope **No Rust change needed.** `crates/headroom-proxy/src/compression/model_limits.rs` resolves context windows but has **no in-tree callers**; the Rust `[1m]` handling is wire-body sanitization only, correct as-is, and its integration tests assert behavior this PR does not touch. **Judgment call worth a reviewer's eye:** the `[1m]` marker is honored for *any* model, including ones with no 1M tier (`claude-haiku-4-5-20251001[1m]` → 1M). Gating on an allowlist would be more precise but reintroduces a hand-maintained table that rots — the failure mode `model_limits.rs` already documents against. Since `[1m]` is set by our own wrapper and Claude Code's opt-in, honoring it seemed the better default. Happy to tighten. ## Tests - `TestContext1MSuffix` — detection, the 200K→1M promotion, the `max()` floor, ANSI non-promotion, and the wire-id guard for #2027. - `TestLongContextPricing` — the premium on both paths (parametrized), threshold boundary (200,000 vs 200,001), an untiered model charged no premium, and the two halves meeting: a `[1m]` request gets both the 1M window and the premium rate. - `TestLiteLLMCostHelper` — unknown model returns `None`, a known model prices correctly, and `input_tokens` is cache-inclusive. Two existing tests were updated, both pinned to the broken behavior: - `test_estimate_cost_basic` probed a "per 1M" rate by sending exactly 1M tokens, which now crosses the 200K threshold. Re-probed at 100K. (Worth knowing: `claude-3-5-sonnet-20241022` is retired and no longer in LiteLLM, so the alias chain resolves it to `claude-sonnet-4-20250514` and it inherits that model's tier. Harmless — a 200K-window model can't exceed 200K in reality — but it explains the number.) - `test_litellm_provider_info_and_cost_fallbacks` monkeypatched `litellm.completion_cost`; repointed at the new helper seam. ``` ruff check / ruff format / mypy — clean across all six changed source files ``` 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Tejas Chopra <tejas@Tejass-MacBook-Pro.local> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
291 lines
13 KiB
Python
291 lines
13 KiB
Python
"""Tests for Anthropic provider."""
|
|
|
|
import pytest
|
|
|
|
|
|
class TestAnthropicModelSanitization:
|
|
def test_sanitize_model_id_removes_ansi_escape_sequences(self):
|
|
from headroom.providers.anthropic import sanitize_anthropic_model_id
|
|
|
|
assert sanitize_anthropic_model_id("claude-opus-4-8\x1b[1m") == "claude-opus-4-8"
|
|
|
|
def test_sanitize_model_id_removes_displayed_style_suffix(self):
|
|
from headroom.providers.anthropic import sanitize_anthropic_model_id
|
|
|
|
assert sanitize_anthropic_model_id("claude-opus-4-8[1m]") == "claude-opus-4-8"
|
|
assert sanitize_anthropic_model_id("glm-5.2[1m]") == "glm-5.2"
|
|
|
|
def test_sanitize_model_metadata_cleans_nested_model_ids(self):
|
|
from headroom.providers.anthropic import sanitize_anthropic_model_metadata
|
|
|
|
payload = {
|
|
"data": [
|
|
{"id": "claude-opus-4-8\x1b[1m", "display_name": "Claude Opus 4.8"},
|
|
{"id": "claude-sonnet-4-5[1m]"},
|
|
],
|
|
"model": "claude-opus-4-8[1m]",
|
|
}
|
|
|
|
assert sanitize_anthropic_model_metadata(payload) == {
|
|
"data": [
|
|
{"id": "claude-opus-4-8", "display_name": "Claude Opus 4.8"},
|
|
{"id": "claude-sonnet-4-5"},
|
|
],
|
|
"model": "claude-opus-4-8",
|
|
}
|
|
|
|
|
|
class TestContext1MSuffix:
|
|
"""`[1m]` is a 1M-context tier request, not just an ANSI artifact (#1158).
|
|
|
|
Claude Code appends `[1m]` to a model id and only then sends the
|
|
`context-1m` beta header, so the real upstream window is 1M even when the
|
|
base model defaults to 200K. The suffix must still be stripped off the wire
|
|
(upstream rejects it, #2027) but must not be lost before we size the budget.
|
|
"""
|
|
|
|
@pytest.fixture
|
|
def provider(self):
|
|
from headroom.providers.anthropic import AnthropicProvider
|
|
|
|
return AnthropicProvider()
|
|
|
|
def test_1m_suffix_is_detected(self):
|
|
from headroom.providers.anthropic import has_context_1m_suffix
|
|
|
|
assert has_context_1m_suffix("claude-sonnet-4-5[1m]")
|
|
assert has_context_1m_suffix("claude-sonnet-4-5[1m][1m]")
|
|
assert not has_context_1m_suffix("claude-sonnet-4-5")
|
|
|
|
def test_ansi_artifacts_are_not_mistaken_for_a_tier_request(self):
|
|
from headroom.providers.anthropic import has_context_1m_suffix
|
|
|
|
# A dangling reset, a compound style, and a real escape sequence are
|
|
# terminal noise -- none of them means "give me 1M".
|
|
assert not has_context_1m_suffix("claude-sonnet-4-5[0m]")
|
|
assert not has_context_1m_suffix("claude-sonnet-4-5[1;32m]")
|
|
assert not has_context_1m_suffix("\x1b[1mclaude-sonnet-4-5\x1b[0m")
|
|
|
|
def test_1m_suffix_raises_a_200k_model_to_1m(self, provider):
|
|
# The regression: sanitizing before the lookup resolved this to the
|
|
# base model's 200K window, so a 1M request was budgeted at 1/5 size.
|
|
assert provider.get_context_limit("claude-sonnet-4-5") == 200_000
|
|
assert provider.get_context_limit("claude-sonnet-4-5[1m]") == 1_000_000
|
|
|
|
def test_1m_suffix_never_lowers_an_already_larger_window(self, provider):
|
|
# max(), not a flat assignment: a base model wider than 1M keeps its own.
|
|
assert provider.get_context_limit("claude-opus-5[1m]") >= 1_000_000
|
|
|
|
def test_ansi_artifact_does_not_inflate_the_window(self, provider):
|
|
assert provider.get_context_limit("claude-sonnet-4-5[0m]") == 200_000
|
|
assert provider.get_context_limit("\x1b[1mclaude-sonnet-4-5\x1b[0m") == 200_000
|
|
|
|
def test_wire_model_id_still_drops_the_suffix(self):
|
|
# Upstream rejects `[1m]`; the tier fix must not regress #2027.
|
|
from headroom.providers.anthropic import sanitize_anthropic_model_id
|
|
|
|
assert sanitize_anthropic_model_id("claude-sonnet-4-5[1m]") == "claude-sonnet-4-5"
|
|
|
|
|
|
class TestLongContextPricing:
|
|
"""Anthropic's long-context premium above a 200K prompt.
|
|
|
|
On the Sonnet 4 / 4.5 family a prompt over 200K re-prices the *whole*
|
|
request -- input, output and cache alike -- at input 2x, output 1.5x,
|
|
cache 2x. Both the LiteLLM path and the manual fallback must apply it, or
|
|
Headroom under-reports the cost of exactly the sessions `[1m]` unlocks.
|
|
"""
|
|
|
|
@pytest.fixture
|
|
def provider(self):
|
|
from headroom.providers.anthropic import AnthropicProvider
|
|
|
|
return AnthropicProvider()
|
|
|
|
@pytest.fixture
|
|
def manual_provider(self, monkeypatch):
|
|
"""Provider with the LiteLLM path disabled, exercising the fallback."""
|
|
import headroom.providers.anthropic as anthropic_module
|
|
|
|
monkeypatch.setattr(anthropic_module, "estimate_cost_from_tokens", lambda *a, **k: None)
|
|
return anthropic_module.AnthropicProvider()
|
|
|
|
# 100K in / 5K out -> 100K*$3 + 5K*$15 = $0.375
|
|
# 300K in / 5K out -> 300K*$6 + 5K*$22.5 = $1.9125 (premium)
|
|
# 300K in of which 150K cached, 5K out
|
|
# -> 150K*$6 + 150K*$0.60 + 5K*$22.5 = $1.1025
|
|
_CASES = [
|
|
(100_000, 5_000, 0, 0.3750),
|
|
(300_000, 5_000, 0, 1.9125),
|
|
(300_000, 5_000, 150_000, 1.1025),
|
|
]
|
|
|
|
@pytest.mark.parametrize(("input_tokens", "output_tokens", "cached_tokens", "expected"), _CASES)
|
|
def test_litellm_path(self, provider, input_tokens, output_tokens, cached_tokens, expected):
|
|
cost = provider.estimate_cost(
|
|
input_tokens, output_tokens, "claude-sonnet-4-5", cached_tokens
|
|
)
|
|
assert cost == pytest.approx(expected, rel=1e-4)
|
|
|
|
@pytest.mark.parametrize(("input_tokens", "output_tokens", "cached_tokens", "expected"), _CASES)
|
|
def test_manual_fallback_matches_litellm(
|
|
self, manual_provider, input_tokens, output_tokens, cached_tokens, expected
|
|
):
|
|
cost = manual_provider.estimate_cost(
|
|
input_tokens, output_tokens, "claude-sonnet-4-5", cached_tokens
|
|
)
|
|
assert cost == pytest.approx(expected, rel=1e-4)
|
|
|
|
def test_untiered_model_is_not_charged_a_premium(self, manual_provider):
|
|
# Opus is flat-rated across its whole window: 300K*$5 + 5K*$25 = $1.625.
|
|
cost = manual_provider.estimate_cost(300_000, 5_000, "claude-opus-4-5-20251101", 0)
|
|
assert cost == pytest.approx(1.625, rel=1e-4)
|
|
|
|
def test_premium_applies_only_above_the_threshold(self, manual_provider):
|
|
at = manual_provider.estimate_cost(200_000, 0, "claude-sonnet-4-5", 0)
|
|
just_over = manual_provider.estimate_cost(200_001, 0, "claude-sonnet-4-5", 0)
|
|
assert at == pytest.approx(0.60, rel=1e-4) # 200K * $3
|
|
assert just_over == pytest.approx(1.2000, rel=1e-3) # re-priced at $6
|
|
|
|
def test_1m_suffix_request_is_priced_at_the_premium(self, manual_provider):
|
|
# The two halves of this PR meeting: `[1m]` unlocks the window, and a
|
|
# session that fills it is billed at the long-context rate.
|
|
assert manual_provider.get_context_limit("claude-sonnet-4-5[1m]") == 1_000_000
|
|
cost = manual_provider.estimate_cost(300_000, 5_000, "claude-sonnet-4-5[1m]", 0)
|
|
assert cost == pytest.approx(1.9125, rel=1e-4)
|
|
|
|
|
|
class TestLiteLLMCostHelper:
|
|
"""The shared helper each provider now uses for LiteLLM-backed pricing.
|
|
|
|
It replaces a `litellm.completion_cost(prompt_tokens=...)` call that had
|
|
stopped accepting those kwargs and raised TypeError on every invocation.
|
|
"""
|
|
|
|
def test_returns_none_for_unknown_model(self):
|
|
from headroom.pricing.litellm_pricing import estimate_cost_from_tokens
|
|
|
|
assert estimate_cost_from_tokens("no-such-model-xyz", 1000, 1000) is None
|
|
|
|
def test_prices_a_known_model(self):
|
|
from headroom.pricing.litellm_pricing import estimate_cost_from_tokens
|
|
|
|
# gpt-4o: $2.50/1M in, $10/1M out -> 100K in + 5K out = $0.30
|
|
assert estimate_cost_from_tokens("gpt-4o", 100_000, 5_000) == pytest.approx(0.30, rel=1e-4)
|
|
|
|
def test_input_tokens_are_cache_inclusive(self):
|
|
from headroom.pricing.litellm_pricing import estimate_cost_from_tokens
|
|
|
|
# The cached portion is a subset of input_tokens, not additional to it,
|
|
# so a fully-cached prompt costs strictly less than an uncached one.
|
|
uncached = estimate_cost_from_tokens("gpt-4o", 100_000, 5_000)
|
|
cached = estimate_cost_from_tokens("gpt-4o", 100_000, 5_000, cached_tokens=50_000)
|
|
assert cached < uncached
|
|
|
|
|
|
class TestAnthropicTokenCounting:
|
|
@pytest.fixture
|
|
def anthropic_provider(self):
|
|
from headroom.providers.anthropic import AnthropicProvider
|
|
|
|
return AnthropicProvider()
|
|
|
|
def test_count_text_fallback(self, anthropic_provider):
|
|
# Without API client, should use tiktoken fallback
|
|
counter = anthropic_provider.get_token_counter("claude-3-5-sonnet-20241022")
|
|
count = counter.count_text("Hello world")
|
|
assert count > 0
|
|
|
|
def test_count_messages_basic(self, anthropic_provider):
|
|
counter = anthropic_provider.get_token_counter("claude-3-5-sonnet-20241022")
|
|
messages = [{"role": "user", "content": "Hello"}]
|
|
count = counter.count_messages(messages)
|
|
assert count > 0
|
|
|
|
def test_count_messages_tolerates_null_tool_calls(self, anthropic_provider):
|
|
# OpenAI-format assistant messages routinely carry `tool_calls: null`
|
|
# (and occasionally `function: null`) on a no-tool turn. The estimated
|
|
# counter iterated the value after only a key-presence check, so it
|
|
# raised `TypeError: 'NoneType' object is not iterable`.
|
|
counter = anthropic_provider.get_token_counter("claude-3-5-sonnet-20241022")
|
|
messages = [
|
|
{"role": "assistant", "content": "hi", "tool_calls": None},
|
|
{"role": "assistant", "content": "x", "tool_calls": [{"id": "a", "function": None}]},
|
|
]
|
|
assert counter.count_messages(messages) > 0
|
|
|
|
def test_count_text_allows_literal_special_tokens(self, anthropic_provider):
|
|
counter = anthropic_provider.get_token_counter("claude-3-5-sonnet-20241022")
|
|
count = counter.count_text("prefix <|fim_suffix|> suffix")
|
|
assert count > 0
|
|
|
|
|
|
class TestAnthropicModelLimits:
|
|
@pytest.fixture
|
|
def anthropic_provider(self):
|
|
from headroom.providers.anthropic import AnthropicProvider
|
|
|
|
return AnthropicProvider()
|
|
|
|
def test_get_context_limit_claude_sonnet(self, anthropic_provider):
|
|
limit = anthropic_provider.get_context_limit("claude-3-5-sonnet-20241022")
|
|
assert limit == 200000
|
|
|
|
def test_get_context_limit_claude_opus(self, anthropic_provider):
|
|
limit = anthropic_provider.get_context_limit("claude-3-opus-20240229")
|
|
assert limit == 200000
|
|
|
|
def test_get_context_limit_strips_ansi_model_suffix(self, anthropic_provider):
|
|
assert anthropic_provider.get_context_limit("claude-opus-4-7[1m]") == 1000000
|
|
|
|
def test_get_context_limit_claude_5_family(self, anthropic_provider):
|
|
assert anthropic_provider.get_context_limit("claude-fable-5") == 1000000
|
|
assert anthropic_provider.get_context_limit("claude-opus-4-8") == 1000000
|
|
assert anthropic_provider.get_context_limit("claude-sonnet-5") == 1000000
|
|
|
|
def test_supports_model_known(self, anthropic_provider):
|
|
assert anthropic_provider.supports_model("claude-3-5-sonnet-20241022")
|
|
|
|
def test_supports_model_prefix(self, anthropic_provider):
|
|
assert anthropic_provider.supports_model("claude-3-5-sonnet-latest")
|
|
|
|
def test_token_counter_cache_uses_sanitized_model_id(self, anthropic_provider):
|
|
plain = anthropic_provider.get_token_counter("claude-opus-4-7")
|
|
styled = anthropic_provider.get_token_counter("claude-opus-4-7\x1b[1m")
|
|
|
|
assert styled is plain
|
|
|
|
|
|
class TestAnthropicCostEstimation:
|
|
@pytest.fixture
|
|
def anthropic_provider(self):
|
|
from headroom.providers.anthropic import AnthropicProvider
|
|
|
|
return AnthropicProvider()
|
|
|
|
def test_estimate_cost_basic(self, anthropic_provider):
|
|
# Probed at 100K, below the 200K long-context threshold: a 1M-token
|
|
# probe would cross it and bill at the premium rate, which is a
|
|
# separate property (covered by TestLongContextPricing).
|
|
cost = anthropic_provider.estimate_cost(
|
|
input_tokens=100_000,
|
|
output_tokens=0,
|
|
model="claude-3-5-sonnet-20241022",
|
|
)
|
|
# $3.00 per 1M input
|
|
assert cost == pytest.approx(0.30, rel=0.1)
|
|
|
|
def test_pricing_lookup_strips_ansi_model_suffix(self, anthropic_provider):
|
|
assert anthropic_provider._get_pricing("claude-opus-4-7[1m]") == (
|
|
anthropic_provider._get_pricing("claude-opus-4-7")
|
|
)
|
|
|
|
def test_pricing_claude_5_family(self, anthropic_provider):
|
|
fable = anthropic_provider._get_pricing("claude-fable-5")
|
|
assert fable == {"input": 10.00, "output": 50.00, "cached_input": 1.00}
|
|
|
|
opus = anthropic_provider._get_pricing("claude-opus-4-8")
|
|
assert opus == {"input": 5.00, "output": 25.00, "cached_input": 0.50}
|
|
|
|
sonnet = anthropic_provider._get_pricing("claude-sonnet-5")
|
|
assert sonnet == {"input": 3.00, "output": 15.00, "cached_input": 0.30}
|