headroom/tests/test_openai_pricing_resolution.py
Tejas Chopra 0e1d6bfa79
refactor(pricing): make LiteLLM the source of truth, not the hardcoded table (#2779)
## Description

> **Stacked on #2777.** That PR corrects the built-in table's *values*;
this one stops the table being *authoritative*. Both are wanted — the
fallback should be right **and** not in charge. Merge #2777 first.

You asked whether the token/savings code could be simpler and whether
the hardcoding could go. This is the hardcoding half, and the
encouraging finding is that **almost none of it needed writing** — the
infrastructure already existed and was simply unused.

`headroom/pricing/litellm_pricing.py` (300 lines, LiteLLM-backed, with
an `ImportError` fallback and gateway-prefix handling) has been in the
tree the whole time. `ModelInfo`'s own docstring says:

> *"Pricing is fetched dynamically from LiteLLM's database. Use
`ModelRegistry.estimate_cost()` to get current pricing."*

Yet **zero of the four providers called it** (`grep -c litellm_pricing`
→ openai 0, anthropic 0, google 0, cohere 0). Each kept a parallel
hardcoded table. `_get_pricing` had no LiteLLM lookup at all, unlike
`get_context_limit` — which is precisely how it went ~18 months stale
and priced `gpt-4.1-nano` **300× over**.

## Changes Made

**1. Resolution order now mirrors `get_context_limit`**, so limits and
prices can't disagree:

```
explicit user config  ->  LiteLLM  ->  built-in table  ->  family  ->  unknown default
```

Config beats LiteLLM because a configured price is a decision, not a
guess. The table stays because it must: the `litellm` dependency is
gated `python_version < '3.14'`, and LiteLLM doesn't know every model.
It just isn't in charge, so its drift only reaches installs with no
LiteLLM.

**2. Gateway-routed names now resolve at all.** `litellm.model_cost`
keys the *unwrapped* form, so `bedrock/anthropic.claude-...` missed
every candidate and silently took the $2.50/$10.00 unknown default:

| model | before | after |
|---|---|---|
| `bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0` | $2.50 / $10.00 |
**$3.00 / $15.00** |
| `bedrock/us.anthropic.claude-3-5-sonnet-...-v2:0` | $2.50 / $10.00 |
**$3.00 / $15.00** |
| `vertex_ai/claude-sonnet-4-5` | $2.50 / $10.00 | **$3.00 / $15.00** |
| `groq/llama-3.3-70b-versatile` | $2.50 / $10.00 | **$0.59 / $0.79** |
| `gemini-2.5-flash` | $2.50 / $10.00 | **$0.30 / $2.50** |
| `deepseek-chat` | $2.50 / $10.00 | **$0.28 / $0.42** |

`pricing_lookup_candidates` only ever *prepended* provider prefixes. It
now also tries progressively unwrapped forms, derived by splitting on
`/` — deliberately **not** another hardcoded gateway-prefix list. A
wrong guess costs nothing: each candidate is an exact dict lookup, so it
just misses.

**3. The staleness warning became meaningful.** It fires only when the
fallback table is actually used. Before, it was unconditional — and with
`_PRICING_LAST_UPDATED = 2025-01-14` against a 60-day window it had been
firing for ~18 months, which trains people to ignore it.

**4. `pricing_per_1m` rounds to 6dp.** LiteLLM stores cost *per token*,
so `× 1e6` leaves float noise ($0.4/1M arrives as
`0.39999999999999997`).

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [x] Code refactoring

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check` + `ruff format`)
- [x] Type checking passes (`mypy`)
- [x] New tests added for new functionality

### Test Output

```text
$ pytest tests/test_pricing_from_litellm.py -q
10 passed
```

Full pricing / cost / provider / models / savings / reporting /
tokenizer set:

```text
$ pytest tests/test_*{pricing,cost,provider,models,savings,utils,reporting,token}*.py -q
3 failed, 1026 passed, 38 skipped

pre-existing on main (all three in my recorded baseline):
  test_bundled_tools_savings.py::test_compressed_payload_preserves_answer_anthropic
  test_compress_route_tokenizer_by_model.py::...[deepseek/deepseek-v4]   (needs transformers)
  test_compress_route_tokenizer_by_model.py::...[deepseek/deepseek-v4]   (needs transformers)
```

Deferring the full sharded run to CI — no maturin/Rust core locally.

### One test of mine changed, and why

Three cases in #2777's `test_openai_pricing_resolution.py` failed on
exact equality once prices started coming from LiteLLM — `gpt-4.1-mini`
arrived as `0.39999999999999997` rather than `0.4`. The values were
right; binary floating point isn't exact. Switched those to
`pytest.approx(..., abs=0.001)` — money compared to the cent — which
passes whether the number comes from LiteLLM or the literal table.

## Deliberately NOT in this PR

- **Encodings stay hardcoded.** LiteLLM carries no tiktoken encoding
data, and `_lookup_encoding_name`'s `None` return is load-bearing (it's
the "not an OpenAI model" signal from #2761). Encodings also track
tokenizer generations, not monthly price changes — they aren't the drift
problem.
- **Anthropic / Cohere / Google providers.** Same shape, same fix, but
Anthropic's pricing is a `{input, output, cached_input}` dict rather
than a tuple, and its matcher is worse (`if model in known_model or
known_model in model` — bidirectional substring). Worth its own PR
rather than tripling this diff.
- **Context limits.** Already LiteLLM-first; the layering there was
correct all along.
- `accounts/fireworks/models/kimi-k2` still falls back — LiteLLM
genuinely has no entry. The file already shows the pattern for filling
such gaps (`_register_minimax_pricing`, `_inject_deepseek_pricing`) if
we want it.
2026-08-04 11:32:26 -07:00

84 lines
3 KiB
Python

"""OpenAI pricing must not be shadowed by a shorter model family.
``_get_pricing`` matches by prefix in plain dict order, so the first *inserted*
prefix won rather than the most specific one. ``gpt-4.1`` fell into the ``gpt-4``
entry and was priced at the legacy $30/$60:
gpt-4.1 $30.00 in vs $2.00 actual 15x
gpt-4.1-mini $30.00 in vs $0.40 actual 75x
gpt-4.1-nano $30.00 in vs $0.10 actual 300x
Unlike ``get_context_limit``, ``_get_pricing`` has no litellm lookup in front of
it, so this table is the only source for ``client.py``'s cost_before/cost_after,
``reporting/generator.py`` and ``evals/cost_tracker.py``. (The *proxy* cost path
is unaffected -- it calls ``litellm.cost_per_token`` directly, as does
``savings_ledger``.)
Values here were verified against litellm's ``model_cost``.
"""
from __future__ import annotations
import pytest
from headroom.providers.openai import OpenAIProvider
# (model, input $/1M, output $/1M)
EXPECTED = [
# The shadowed cases.
("gpt-4.1", 2.00, 8.00),
("gpt-4.1-mini", 0.40, 1.60),
("gpt-4.1-nano", 0.10, 0.40),
("gpt-4.1-2025-04-14", 2.00, 8.00),
# Fell through to the unknown-model default (GPT-4o tier).
("gpt-5", 1.25, 10.00),
("gpt-5-mini", 0.25, 2.00),
("gpt-5-nano", 0.05, 0.40),
("o4-mini", 1.10, 4.40),
# Stale entry: o3 was cut to $2/$8 in June 2025.
("o3", 2.00, 8.00),
# Must not regress.
("gpt-4o", 2.50, 10.00),
("gpt-4o-mini", 0.15, 0.60),
("gpt-4", 30.00, 60.00),
("gpt-4-turbo", 10.00, 30.00),
("gpt-3.5-turbo", 0.50, 1.50),
("o3-mini", 1.10, 4.40),
("o1", 15.00, 60.00),
]
@pytest.mark.parametrize(("model", "want_in", "want_out"), EXPECTED)
def test_pricing_prefers_the_most_specific_prefix(
model: str, want_in: float, want_out: float
) -> None:
got_in, got_out = OpenAIProvider()._get_pricing(model)
# Tolerance, not equality: these rates may come from LiteLLM's per-token
# figures, and the x1e6 conversion is not exact in binary floating point
# ($0.4/1M arrives as 0.39999999999999997). Money compared to the cent.
assert got_in == pytest.approx(want_in, abs=0.001)
assert got_out == pytest.approx(want_out, abs=0.001)
def test_nano_is_not_priced_as_legacy_gpt4() -> None:
"""The 300x case, stated plainly: nano must be the cheapest gpt-4.1 tier."""
provider = OpenAIProvider()
nano_in, _ = provider._get_pricing("gpt-4.1-nano")
legacy_in, _ = provider._get_pricing("gpt-4")
assert nano_in < legacy_in / 100
def test_pricing_metadata_is_not_stale() -> None:
"""The staleness warning is a real feature; keep the stamp honest.
If someone edits _PRICING without re-verifying, this starts failing rather
than silently shipping a stale table behind a fresh-looking date.
"""
from headroom.providers.openai import _PRICING_LAST_UPDATED, _PRICING_STALE_DAYS
assert _PRICING_STALE_DAYS > 0
# Sanity: the stamp should postdate the gpt-4.1/gpt-5 entries it covers.
assert _PRICING_LAST_UPDATED.year >= 2026