fix: surface output reduction without a restart, and explain $0.00 savings on Python 3.14 (#1296)

## Description

I was running headroom through pipx on Python 3.14 and hit two issues.

The Proxy $ Saved tile was stuck at $0.00 even though tokens were
tracking fine. Pricing comes from litellm, and litellm does not install
on Python 3.14 because of a version lock, so there was just nothing to
price against. Rather than hardcode a price table that goes stale, I
added a `litellm_available` flag to `/stats` and the tile now tells you
to reinstall on 3.13 when pricing isn't there, like the output-shaper
tile already does.

The other one was Output Tokens Saved showing "—" after I turned on the
shaper. The recorder reads the learned baseline once at startup, so if
you run `learn --verbosity --apply` while the proxy is already up it
never gets picked up, and a later flush writes the empty baseline over
the one learn just saved. Now it re-reads the baseline before estimating
and before each flush, so it works without a restart.

Closes # N/A (no tracking issue)

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [x] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- `output_savings.py`: re-read the baseline from disk before estimating
and before each flush, so a baseline learned while the proxy is running
takes effect (and a re-learn with the same sample count too).
- `server.py`: expose a `litellm_available` flag on `/stats`.
- `dashboard.html`: when savings are zero and litellm is missing, point
to reinstalling on 3.13 instead of showing $0.00.
- tests and docs (`test_output_savings.py`, README, metrics, CHANGELOG).

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ pytest tests/test_output_savings.py -q
34 passed, 1 warning in 0.11s
```

## Real Behavior Proof

- Environment: macOS, Python 3.13 (litellm present) and 3.14 (litellm
absent), running this branch.
- Exact command / steps: record shaped traffic, write a baseline to the
same file while the recorder is live (no restart), then estimate and
flush.
- Observed result: the recorder goes from `available: False` to
`available: True` once the baseline is written mid-run, and keeps it
after a flush. Before this it stayed `False` and the flush reset the
baseline. Raw output:
  ```text
  shaper traffic recorded, baseline not learned yet -> available: False
  learn --apply wrote baseline while proxy up; restart NOT performed
after baseline write -> available: True | method: estimated | pct: 50.3
  baseline kept after flush -> disk samples: 4
  ```
- Not tested: I did not render the tile hint in a browser, I checked the
flag on `/stats` and read the template instead.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable

## Additional Notes

Just a final note, ruff and mypy are clean on what I changed. The
repo-wide `ruff check .` and `mypy headroom` do report a few problems,
but they're in files I didn't touch and already exist on the base
commit, so I left them alone to keep this small. Happy to do a separate
cleanup PR.

---------

Co-authored-by: JD Davis <mxjerrett@gmail.com>
This commit is contained in:
Lucas Santos 2026-06-26 19:15:42 +02:00 committed by GitHub
parent a0cb7982e3
commit c30ec4cda8
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
7 changed files with 164 additions and 1 deletions

View file

@ -44,6 +44,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
* **proxy:** retry upstream 429 rate limits honoring `Retry-After` instead of passing them straight to the client. Both the non-streaming (`_retry_request`) and streaming (`_stream_response`) forwarders returned an upstream 429 verbatim, so a parallel agent fan-out that exceeded the per-minute limit aborted every run; 429s are now retried with backoff (honoring the upstream `Retry-After`, capped at `retry_max_delay_ms`), surfacing only the exhausted 429 to the client ([#1221](https://github.com/chopratejas/headroom/issues/1221)).
* **proxy:** force Responses API `store=true` when Headroom injects memory tools so `previous_response_id` continuations work after memory tool calls from clients that requested `store=false` ([#1103](https://github.com/chopratejas/headroom/pull/1103)).
* **proxy:** build SSL contexts for custom CA bundles so enterprise/private PKI roots work with Python/OpenSSL strict verification.
* **dashboard:** the Proxy $ Saved tile no longer shows a bare `$0.00` when cost pricing is unavailable. Pricing depends on litellm, which pyproject gates off on Python 3.14+, so `/stats` now exposes a top-level `litellm_available` flag and the tile points you to reinstall on Python 3.13 when it is false ([#1296](https://github.com/chopratejas/headroom/pull/1296)).
* **proxy:** the output-savings recorder now reloads the learned baseline before estimating and before each flush, so a baseline written by `headroom learn --verbosity --apply` while the proxy is running takes effect without a restart and the periodic flush no longer overwrites it. Fixes Output Tokens Saved staying at "—" after enabling the shaper ([#1296](https://github.com/chopratejas/headroom/pull/1296)).
* **tokenizers:** bound token-counting of oversized tool-content blobs instead of running `count_text` over the whole serialized string. `count_messages` runs on the proxy request path; serializing is cheap, but `count_text` over a multi-megabyte `tool_result` / `tool_use` string took seconds and could freeze `/health` and in-flight requests. For payloads over ~50KB serialized, `count_text` now runs on an even-spread sample of the string and scales by length; it stays model-accurate, bounded for any blob shape, and biased to under-count. Smaller payloads stay exact.
* **codex:** stop persisting a project-specific `--db` path in the global `headroom_memory` MCP config, so `headroom wrap codex --memory` falls back to the active cwd's `.headroom/memory.db` at runtime while keeping the current project's local bootstrap work scoped correctly ([#1147](https://github.com/chopratejas/headroom/issues/1147)).
* **ccr:** stop emitting Anthropic request-side retrieval markers on frozen-prefix turns when `headroom_retrieve` injection is deferred, so cache-preserving requests forward original content instead of irrecoverable marker-only payloads ([#1006](https://github.com/chopratejas/headroom/issues/1006)).

View file

@ -315,6 +315,8 @@ Using `pipx`? Choose a supported interpreter explicitly:
pipx install --python python3.13 "headroom-ai[all]"
```
> **Pick 3.13 if you want dollar savings.** The dashboard's *Proxy $ Saved* tile prices compression with [LiteLLM](https://github.com/BerriAI/litellm), and LiteLLM can't be installed on Python 3.14+. On 3.14 token savings still track, but the dollar figure stays `$0.00`. If you already installed on 3.14, switch with `pipx reinstall headroom-ai --python python3.13` and restart the proxy.
→ [Installation guide](https://headroom-docs.vercel.app/docs/installation) — Docker tags, persistent service, PowerShell, devcontainers.
### Updating

View file

@ -192,8 +192,15 @@
<div class="mt-2 text-xs text-gray-500">
<span x-show="(stats.persistent_savings?.lifetime?.compression_savings_usd || 0) > 0"
x-text="formatNumber(stats.tokens?.proxy_compression_saved || 0) + ' proxy tokens only; ' + cliFilteringLabel + ' excluded from $'"></span>
<span x-show="!((stats.persistent_savings?.lifetime?.compression_savings_usd || 0) > 0)"
<span x-show="!((stats.persistent_savings?.lifetime?.compression_savings_usd || 0) > 0) && stats.litellm_available !== false"
x-text="formatNumber(stats.requests?.total || 0) + ' requests processed'"></span>
<!-- Pricing is LiteLLM-only. LiteLLM currently can't install on
newer Python (e.g. 3.14+); state the fact, not the version,
so the hint stays correct if that changes. -->
<span x-show="!((stats.persistent_savings?.lifetime?.compression_savings_usd || 0) > 0) && stats.litellm_available === false">
Cost pricing needs LiteLLM, which isn't installed in this environment. If you're on a Python version LiteLLM doesn't support yet, reinstall Headroom on Python 3.13:
<code class="text-gray-400">pipx reinstall headroom-ai --python python3.13</code>
</span>
</div>
</div>

View file

@ -451,8 +451,34 @@ class SavingsRecorder:
return True
return False
def _reload_baseline_locked(self) -> None:
"""Adopt the on-disk baseline written by ``learn --verbosity --apply``.
``learn`` rewrites the baseline in place in the same file a running proxy
holds open, while the recorder only ever appends treatment/control
samples and never touches the baseline. Without re-reading it, two things
break: (1) a baseline learned while the proxy is up never takes effect
until a restart, so treatment lookups all miss (``m == 0``) and the
output-reduction tile stays at ""; and (2) our periodic flush would
write our in-memory (empty) baseline straight over the one ``learn`` just
persisted.
Adopt the disk baseline whenever it carries samples and differs from
ours. Comparing content (not just sample count) means a re-learn with the
same number of samples still takes effect, and the empty-disk guard keeps
a truncated file from wiping a baseline we already hold."""
try:
disk = SavingsLedger.load(self._path)
except OSError:
return
if disk.baseline.total_samples == 0:
return
if disk.baseline.to_dict() != self._ledger.baseline.to_dict():
self._ledger.baseline = disk.baseline
def _flush_locked(self) -> None:
try:
self._reload_baseline_locked()
self._ledger.save(self._path)
self._since_flush = 0
except OSError:
@ -464,6 +490,7 @@ class SavingsRecorder:
def estimate(self) -> SavingsEstimate:
with self._lock:
self._reload_baseline_locked()
return self._ledger.best_estimate()

View file

@ -153,6 +153,7 @@ from headroom.proxy.project_context import (
from headroom.proxy.prometheus_metrics import PrometheusMetrics # noqa: F401
from headroom.proxy.rate_limiter import TokenBucketRateLimiter # noqa: F401
from headroom.proxy.request_logger import RequestLogger # noqa: F401
from headroom.proxy.savings_tracker import LITELLM_AVAILABLE
from headroom.proxy.semantic_cache import SemanticCache # noqa: F401
from headroom.proxy.ssl_context import build_httpx_verify
from headroom.proxy.warmup import WarmupRegistry
@ -3015,6 +3016,12 @@ def create_app(config: ProxyConfig | None = None) -> FastAPI:
},
"savings_history": m.savings_history[-100:], # Last 100 data points
"display_session": display_session,
# Whether LiteLLM is importable. Pricing (the "$ Saved" tile) is
# derived entirely from LiteLLM's cost tables, and LiteLLM is gated
# off on Python >=3.14 in pyproject — so when this is False the
# dashboard tells the user to reinstall on 3.13 instead of just
# showing $0.00 forever.
"litellm_available": LITELLM_AVAILABLE,
"persistent_savings": persistent_savings,
"prefix_cache": prefix_cache_stats,
"cost": _merge_cost_stats(

View file

@ -5,12 +5,14 @@ from __future__ import annotations
from headroom.proxy.output_savings import (
BaselineModel,
SavingsLedger,
SavingsRecorder,
assign_arm,
conversation_key_from_body,
echo_ratio,
input_bucket,
model_family,
stratum_key,
stratum_label,
)
# ---------------------------------------------------------------------------
@ -285,3 +287,110 @@ class TestEchoRatio:
def test_short_output_returns_zero(self):
assert echo_ratio("a b", "a b c d e f g h", n=8) == 0.0
# ---------------------------------------------------------------------------
# recorder baseline reload (learn-while-running)
# ---------------------------------------------------------------------------
class TestRecorderBaselineReload:
"""The recorder must pick up a baseline that ``learn --verbosity --apply``
writes while the proxy is already running, and a flush must never overwrite
that learned baseline with the recorder's own empty in-memory copy."""
@staticmethod
def _key() -> str:
return stratum_key(
turn_kind="code",
input_tokens=8000,
model="claude-opus-4-8",
has_tools=True,
)
def test_adopts_baseline_learned_after_start(self, tmp_path):
path = str(tmp_path / "output_savings.json")
key = self._key()
recorder = SavingsRecorder(path, flush_every=1)
for output_tokens in (200, 210, 190):
recorder.record_from_labels([stratum_label("treatment", key)], output_tokens)
# No baseline to compare against yet, so there is nothing to estimate.
assert recorder.estimate().n_requests == 0
# Simulate `learn --verbosity --apply` writing a baseline to the same
# file while the recorder is live (no restart).
learned = SavingsLedger.load(path)
for output_tokens in (400, 420, 380, 410):
learned.baseline.observe(key, output_tokens)
learned.save(path)
estimate = recorder.estimate()
assert estimate.n_requests > 0
assert estimate.kind == "estimated"
assert estimate.tokens_saved > 0
def test_flush_does_not_clobber_learned_baseline(self, tmp_path):
path = str(tmp_path / "output_savings.json")
key = self._key()
# Recorder starts before any baseline exists, so its in-memory baseline
# is empty.
recorder = SavingsRecorder(path, flush_every=1)
learned = SavingsLedger.load(path)
for output_tokens in (400, 420, 380, 410):
learned.baseline.observe(key, output_tokens)
learned.save(path)
assert SavingsLedger.load(path).baseline.total_samples == 4
recorder.record_from_labels([stratum_label("treatment", key)], 200)
recorder.flush()
# The flush must keep the learned baseline rather than writing the empty
# in-memory one over it.
assert SavingsLedger.load(path).baseline.total_samples == 4
def test_does_not_downgrade_to_empty_disk_baseline(self, tmp_path):
path = str(tmp_path / "output_savings.json")
key = self._key()
# Recorder already holds a learned baseline in memory.
recorder = SavingsRecorder(path, flush_every=1)
recorder._ledger.baseline.observe(key, 400)
recorder._ledger.baseline.observe(key, 420)
assert recorder._ledger.baseline.total_samples == 2
# A stale/empty file on disk must not erase a baseline we already hold.
SavingsLedger().save(path)
recorder.flush()
assert recorder._ledger.baseline.total_samples == 2
def test_relearn_with_same_sample_count_is_adopted(self, tmp_path):
path = str(tmp_path / "output_savings.json")
key = self._key()
recorder = SavingsRecorder(path, flush_every=1)
for output_tokens in (200, 210, 190):
recorder.record_from_labels([stratum_label("treatment", key)], output_tokens)
# First learn writes a baseline; the recorder adopts it.
first = SavingsLedger.load(path)
for output_tokens in (400, 400, 400, 400):
first.baseline.observe(key, output_tokens)
first.save(path)
baseline_tokens_v1 = recorder.estimate().baseline_tokens
assert baseline_tokens_v1 > 0
# Re-running learn replaces the baseline in place with the SAME number of
# samples but different values. A sample-count guard would miss this; the
# recorder must still pick the new baseline up.
relearned = SavingsLedger.load(path)
relearned.baseline = BaselineModel()
for output_tokens in (800, 800, 800, 800):
relearned.baseline.observe(key, output_tokens)
relearned.save(path)
assert recorder.estimate().baseline_tokens > baseline_tokens_v1

View file

@ -57,6 +57,15 @@ Use `HEADROOM_SAVINGS_PATH` to override the file location directly, or
set `HEADROOM_WORKSPACE_DIR` to relocate the entire state root. See the
[Filesystem Contract](filesystem-contract.md) for details.
> **`compression_savings_usd` needs LiteLLM (Python 3.13).** Dollar figures are
> priced entirely from LiteLLM's cost tables, and LiteLLM can't be installed on
> Python 3.14+. On 3.14 the token counts are unaffected but every USD field
> (and the dashboard's *Proxy $ Saved* tile) reads `0`. `/stats` exposes a
> top-level `"litellm_available"` boolean so clients can tell "genuinely $0"
> apart from "pricing unavailable"; the dashboard uses it to prompt a reinstall
> on 3.13 (`pipx reinstall headroom-ai --python python3.13`) rather than showing
> a misleading `$0.00`.
For Anthropic-style providers that return cache-write TTL buckets, `/stats`
also surfaces observed cache TTL usage under `prefix_cache`: