headroom/tests/test_pricing_from_litellm.py
Tejas Chopra 0e1d6bfa79
refactor(pricing): make LiteLLM the source of truth, not the hardcoded table (#2779)
## Description

> **Stacked on #2777.** That PR corrects the built-in table's *values*;
this one stops the table being *authoritative*. Both are wanted — the
fallback should be right **and** not in charge. Merge #2777 first.

You asked whether the token/savings code could be simpler and whether
the hardcoding could go. This is the hardcoding half, and the
encouraging finding is that **almost none of it needed writing** — the
infrastructure already existed and was simply unused.

`headroom/pricing/litellm_pricing.py` (300 lines, LiteLLM-backed, with
an `ImportError` fallback and gateway-prefix handling) has been in the
tree the whole time. `ModelInfo`'s own docstring says:

> *"Pricing is fetched dynamically from LiteLLM's database. Use
`ModelRegistry.estimate_cost()` to get current pricing."*

Yet **zero of the four providers called it** (`grep -c litellm_pricing`
→ openai 0, anthropic 0, google 0, cohere 0). Each kept a parallel
hardcoded table. `_get_pricing` had no LiteLLM lookup at all, unlike
`get_context_limit` — which is precisely how it went ~18 months stale
and priced `gpt-4.1-nano` **300× over**.

## Changes Made

**1. Resolution order now mirrors `get_context_limit`**, so limits and
prices can't disagree:

```
explicit user config  ->  LiteLLM  ->  built-in table  ->  family  ->  unknown default
```

Config beats LiteLLM because a configured price is a decision, not a
guess. The table stays because it must: the `litellm` dependency is
gated `python_version < '3.14'`, and LiteLLM doesn't know every model.
It just isn't in charge, so its drift only reaches installs with no
LiteLLM.

**2. Gateway-routed names now resolve at all.** `litellm.model_cost`
keys the *unwrapped* form, so `bedrock/anthropic.claude-...` missed
every candidate and silently took the $2.50/$10.00 unknown default:

| model | before | after |
|---|---|---|
| `bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0` | $2.50 / $10.00 |
**$3.00 / $15.00** |
| `bedrock/us.anthropic.claude-3-5-sonnet-...-v2:0` | $2.50 / $10.00 |
**$3.00 / $15.00** |
| `vertex_ai/claude-sonnet-4-5` | $2.50 / $10.00 | **$3.00 / $15.00** |
| `groq/llama-3.3-70b-versatile` | $2.50 / $10.00 | **$0.59 / $0.79** |
| `gemini-2.5-flash` | $2.50 / $10.00 | **$0.30 / $2.50** |
| `deepseek-chat` | $2.50 / $10.00 | **$0.28 / $0.42** |

`pricing_lookup_candidates` only ever *prepended* provider prefixes. It
now also tries progressively unwrapped forms, derived by splitting on
`/` — deliberately **not** another hardcoded gateway-prefix list. A
wrong guess costs nothing: each candidate is an exact dict lookup, so it
just misses.

**3. The staleness warning became meaningful.** It fires only when the
fallback table is actually used. Before, it was unconditional — and with
`_PRICING_LAST_UPDATED = 2025-01-14` against a 60-day window it had been
firing for ~18 months, which trains people to ignore it.

**4. `pricing_per_1m` rounds to 6dp.** LiteLLM stores cost *per token*,
so `× 1e6` leaves float noise ($0.4/1M arrives as
`0.39999999999999997`).

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [x] Code refactoring

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check` + `ruff format`)
- [x] Type checking passes (`mypy`)
- [x] New tests added for new functionality

### Test Output

```text
$ pytest tests/test_pricing_from_litellm.py -q
10 passed
```

Full pricing / cost / provider / models / savings / reporting /
tokenizer set:

```text
$ pytest tests/test_*{pricing,cost,provider,models,savings,utils,reporting,token}*.py -q
3 failed, 1026 passed, 38 skipped

pre-existing on main (all three in my recorded baseline):
  test_bundled_tools_savings.py::test_compressed_payload_preserves_answer_anthropic
  test_compress_route_tokenizer_by_model.py::...[deepseek/deepseek-v4]   (needs transformers)
  test_compress_route_tokenizer_by_model.py::...[deepseek/deepseek-v4]   (needs transformers)
```

Deferring the full sharded run to CI — no maturin/Rust core locally.

### One test of mine changed, and why

Three cases in #2777's `test_openai_pricing_resolution.py` failed on
exact equality once prices started coming from LiteLLM — `gpt-4.1-mini`
arrived as `0.39999999999999997` rather than `0.4`. The values were
right; binary floating point isn't exact. Switched those to
`pytest.approx(..., abs=0.001)` — money compared to the cent — which
passes whether the number comes from LiteLLM or the literal table.

## Deliberately NOT in this PR

- **Encodings stay hardcoded.** LiteLLM carries no tiktoken encoding
data, and `_lookup_encoding_name`'s `None` return is load-bearing (it's
the "not an OpenAI model" signal from #2761). Encodings also track
tokenizer generations, not monthly price changes — they aren't the drift
problem.
- **Anthropic / Cohere / Google providers.** Same shape, same fix, but
Anthropic's pricing is a `{input, output, cached_input}` dict rather
than a tuple, and its matcher is worse (`if model in known_model or
known_model in model` — bidirectional substring). Worth its own PR
rather than tripling this diff.
- **Context limits.** Already LiteLLM-first; the layering there was
correct all along.
- `accounts/fireworks/models/kimi-k2` still falls back — LiteLLM
genuinely has no entry. The file already shows the pattern for filling
such gaps (`_register_minimax_pricing`, `_inject_deepseek_pricing`) if
we want it.
2026-08-04 11:32:26 -07:00

87 lines
3.2 KiB
Python

"""Pricing comes from LiteLLM; the built-in table is only a fallback.
`_PRICING` used to be authoritative, with no LiteLLM lookup in front of it. That
is how it went ~18 months stale and priced `gpt-4.1-nano` 300x over. The table is
still needed -- the `litellm` dependency is gated `python_version < '3.14'`, and
LiteLLM does not know every model -- but it must not outrank the live source.
Resolution order (mirroring `get_context_limit`, so limits and prices agree):
1. explicit user config (`HEADROOM_MODEL_LIMITS` / `models.json`)
2. LiteLLM
3. built-in table -> family pattern -> unknown default
"""
from __future__ import annotations
import pytest
from headroom.pricing.litellm_model_resolution import unwrapped_model_forms
from headroom.providers.openai import OpenAIProvider
litellm = pytest.importorskip("litellm")
def test_unwrapped_model_forms_drops_leading_segments() -> None:
"""Pure function: no gateway-prefix list to maintain."""
assert unwrapped_model_forms("bedrock/anthropic.claude-x") == ("anthropic.claude-x",)
assert unwrapped_model_forms("accounts/fireworks/models/kimi-k2") == (
"fireworks/models/kimi-k2",
"models/kimi-k2",
"kimi-k2",
)
assert unwrapped_model_forms("gpt-4o") == ()
@pytest.mark.parametrize(
("model", "want_in", "want_out"),
[
# Gateway-routed names. litellm.model_cost keys the UNWRAPPED form, so
# these all returned None (-> $2.50/$10.00 unknown default) before.
("bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0", 3.00, 15.00),
("bedrock/us.anthropic.claude-3-5-sonnet-20241022-v2:0", 3.00, 15.00),
("vertex_ai/claude-sonnet-4-5", 3.00, 15.00),
("groq/llama-3.3-70b-versatile", 0.59, 0.79),
# Non-OpenAI models reachable through the OpenAI-compatible passthrough.
("gemini-2.5-flash", 0.30, 2.50),
("deepseek-chat", 0.28, 0.42),
],
)
def test_provider_prices_models_its_table_never_covered(
model: str, want_in: float, want_out: float
) -> None:
got_in, got_out = OpenAIProvider()._get_pricing(model)
assert (round(got_in, 2), round(got_out, 2)) == (want_in, want_out)
def test_explicit_config_outranks_litellm() -> None:
"""A configured price is a decision, not a guess."""
provider = OpenAIProvider()
provider._pricing_overrides["gpt-4o"] = (99.0, 111.0)
assert provider._get_pricing("gpt-4o") == (99.0, 111.0)
def test_falls_back_to_the_builtin_table_without_litellm(monkeypatch: pytest.MonkeyPatch) -> None:
"""Offline / Python >= 3.14 installs must still get sane numbers.
Pinned because the fallback is exactly where the table's correctness still
matters -- it is the only thing those installs see.
"""
import headroom.pricing.litellm_pricing as lp
monkeypatch.setattr(lp, "LITELLM_AVAILABLE", False)
provider = OpenAIProvider()
assert provider._get_pricing("gpt-4.1-nano") == (0.10, 0.40)
assert provider._get_pricing("gpt-4") == (30.00, 60.00)
assert provider._get_pricing("o3") == (2.00, 8.00)
def test_unknown_model_still_returns_a_usable_default() -> None:
"""Never raise, never return None, for a model nobody knows."""
got = OpenAIProvider()._get_pricing("totally-made-up-model-xyz")
assert got is not None
assert got[0] > 0