feat: add Anthropic Claude Code subscription window tracking
- New headroom/subscription/ package:
- models.py: RateLimitWindow, ExtraUsage (cents->USD), SubscriptionSnapshot
with per-model 7d windows (opus/sonnet), HeadroomContribution,
WindowDiscrepancy, SubscriptionState
- client.py: async httpx client for GET /api/oauth/usage with OAuth token
resolution (env var -> credentials file -> live proxy header)
- tracker.py: background polling singleton (asyncio Task), notify_active(),
update_contribution(), anomaly detection (surge pricing, cache miss),
atomic persistence, OTEL callback
- session_tracking.py: JSONL transcript reader, Sonnet-normalised model
weights, compute_window_tokens() for 5h/7d window boundaries
- __init__.py: clean public re-exports
- Integration:
- proxy/models.py: subscription_tracking_enabled, poll_interval_s,
active_window_s config fields
- proxy/server.py: tracker startup/shutdown, /subscription-window endpoint,
subscription_window key in /stats
- proxy/handlers/anthropic.py: OAuth Bearer detection, notify_active() +
update_contribution() per request
- observability/metrics.py: 5 observable gauges (5h/7d utilisation %,
5h/7d seconds-to-reset, overage USD) using Observation type
- dashboard/templates/dashboard.html: subscription window panel with 5h/7d
progress bars, countdown, overage, Headroom contribution bar + anomaly
alerts
- Tests: 33 new unit tests in tests/test_subscription_tracker.py (all pass)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-04-11 21:53:44 -05:00
|
|
|
from __future__ import annotations
|
2026-04-24 15:33:30 +02:00
|
|
|
|
fix: remove rtk and lean-ctx CLI context tools (#2677)
## Description
Removes both third-party CLI context tools — **rtk** and **lean-ctx** —
and with them the context-tool selector itself. Headroom no longer
downloads, installs or configures either one, and there is no
replacement.
The previous pass (#2344) gated only three entry points inside
`headroom/cli/wrap.py`. That left the feature reachable in practice:
| Gap | Effect |
|---|---|
| `scripts/install.sh:1544`, `install.ps1:1681` | Ran `rtk init --global
--auto-patch` from bash/PowerShell, **bypassing the Python gate
entirely** — `curl \| sh` still wrote a Claude Code `PreToolUse` hook
regardless of `HEADROOM_RTK` |
| `wrap.py` `_setup_context_tool_for_agent` | **`wrap openhands` was
broken by default**: `rtk_required=True` met a gate returning `None` →
`SystemExit(1)`. Invisible because all 8 openhands tests patched
`_ensure_rtk_binary` to a fake path |
| `proxy/helpers.py`, `subscription/tracker.py` | Proxy shelled out to
`rtk gain` from `/stats`, the dashboard and `headroom perf`; the tracker
polled it per contribution (`_RTK_WIRING_DEFAULT = "enabled"`) |
| No cleanup path | Nothing removed artifacts an earlier default had
installed, so a machine that once ran the old default kept rtk in the
loop forever (#1669, #1955) |
Also worth noting: the rtk binary download had **no SHA or signature
verification** — only `rtk --version` as a smoke test.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [x] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [x] Code refactoring (no functional changes)
## Changes Made
**Removed** — `headroom/rtk/` and `headroom/lean_ctx/` packages,
`headroom/cli/wrap_rtk_metrics.py`, `_selected_context_tool` /
`_setup_context_tool_for_agent` / `_VALID_CONTEXT_TOOLS`, the `--rtk` /
`--no-rtk` / `--no-project-rtk` / `--keep-rtk` flags across all 18 wrap
subcommands, `HEADROOM_RTK*`, the proxy-side `rtk gain` polling, the
dashboard CLI-filtering panel (rows + all 8 `cliFiltering*` Alpine
getters), `paths.rtk_path()` / `lean_ctx_path()`, the SDK path helpers,
`benchmarks/rtk_loop_learn_eval.py`, and the `headroom/rtk/**` CI path
filters.
**Fails loudly, not silently** — `--context-tool` / `--no-context-tool`
/ `HEADROOM_CONTEXT_TOOL` are kept solely to error out. They live in
shell profiles, aliases and CI jobs, and accepting them as a no-op would
read as Headroom having quietly stopped working. The installers reject
them too, which matters more than it looks: their arg parsers forward
the first unknown flag **and everything after it** to the wrapped tool,
so a leftover `--no-rtk` would have silently swallowed a following
`--port` and then been ignored downstream.
**New `headroom/context_tool_cleanup.py`** — deleting the code cannot
help a machine that already ran the old default, since the hooks,
binaries and injected guidance are durable on disk.
`purge_context_tool_artifacts()` runs once per `wrap`/`unwrap` and
removes the registered hook entries, the generated hook scripts, the
Headroom-managed `~/.local/bin` symlinks, the vendored
`~/.headroom/bin/{rtk,lean-ctx}` binaries, the `lean-ctx` MCP server
entry and the marker-fenced instruction blocks. Deliberately
conservative: idempotent, **skips** a malformed config rather than
overwriting it, and only unlinks a symlink resolving inside Headroom's
own bin dir so a user's own build is untouched. It reports on
**stderr**, because `wrap/unwrap openclaw --prepare-only` emit
machine-readable JSON on stdout as their entire contract. Skipped for
`wrap selfheal` (runs from a SessionStart hook; must not race Claude
Code's writer for `~/.claude.json`) and for `--help`, which must stay
read-only.
**Client-config hardening** (discovered while investigating a "corrupted
Serena settings file" report) — `wrap.py` reset a settings file to `{}`
when an existing file would not parse, then wrote that back. One
hand-edited typo or a transient `EACCES`/`EINTR` on a valid file
destroyed the user's `permissions`, `env` and `hooks`, on **every
`headroom wrap claude`**. It now refuses to write. Separately,
`fsutil.write_text` is now atomic (temp file + `fsync` + `os.replace`),
fixing all 14 non-atomic client-config writes at once; it follows
symlinks rather than replacing them (dotfile managers) and preserves an
existing file's mode.
**Deliberately kept** — `rtk` stays in the wrapper-peel list in
`transforms/content_router.py`. It sits beside `sudo`/`env`/`timeout` as
shell-command grammar, so `rtk cat f` is still classified as a file read
for anyone running their own rtk install, which the purge intentionally
leaves alone.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ ruff check headroom/ tests/ e2e/ --exclude headroom/dashboard/templates
All checks passed!
$ ruff format --check headroom/ tests/ e2e/ --exclude headroom/dashboard/templates
1255 files already formatted
$ mypy headroom/
Success: no issues found in 508 source files
$ pytest tests/test_context_tool_cleanup.py -q
11 passed
$ pytest tests/test_fsutil.py -q
12 passed
$ pytest tests/test_cli/test_wrap_codex.py -q # 89 tests
89 passed in 431.68s
$ pytest tests/test_cli/test_wrap_opencode.py -q
39 passed in 257.46s
$ pytest tests/test_cli/test_wrap_helpers.py -q
45 passed
$ pytest tests/test_paths.py -q
75 passed
$ pytest tests/test_cli/test_unwrap_claude.py -q
14 passed
$ pytest tests/test_proxy_savings_history.py -q
39 passed
$ pytest tests/test_cli/test_wrap_copilot.py -q
27 passed
$ pytest tests/test_cli/test_wrap_zcode.py -q
20 passed
$ pytest tests/test_subscription_tracker.py -q
9 passed
$ pytest tests/test_proxy_dashboard_stats_cache.py -q
5 passed, 1 skipped
```
Repo-wide grep for 14 removed symbols (`headroom.rtk`,
`headroom.lean_ctx`, `_ensure_rtk_binary`, `_selected_context_tool`,
`_get_context_tool_stats`, `rtk_path`, `lean_ctx_path`,
`wrap_rtk_metrics`, `HEADROOM_RTK`, `cli_tokens_avoided`,
`tokens_saved_rtk`, …) across `*.py`, `*.ts`, `*.sh`, `*.ps1`, `*.yml`,
`*.html`: **zero hits**.
Notable test changes: `test_wrap_openhands.py` no longer patches
`_ensure_rtk_binary` and asserts `wrap openhands --prepare-only` exits 0
unpatched — the regression that was previously masked.
`test_wrap_continue.py` and `test_wrap_hintfile_agents.py` were removed
(every test drove RTK instruction injection). A new
`test_subscription_tracker.py::test_load_state_written_before_cli_context_tools_were_removed`
proves a pre-removal `subscription_state.json` still loads.
## Real Behavior Proof
- **Environment:** macOS 15.4 (darwin 25.4.0), Python 3.12.6, Headroom @
this branch, real `~/.headroom` and `~/.claude` on the dev machine.
- **Exact command / steps and observed result:**
```text
# 1. Retired flag fails loudly instead of silently no-op'ing
$ headroom wrap codex --prepare-only --context-tool rtk
Error: CLI context tools (rtk, lean-ctx) have been removed from Headroom: they
rewrote shell commands through a third-party binary Headroom no longer manages.
Drop --context-tool / --no-context-tool and unset HEADROOM_CONTEXT_TOOL;
`headroom wrap` uninstalls what they left behind on first run.
$ HEADROOM_CONTEXT_TOOL=lean-ctx headroom wrap codex --prepare-only
Error: CLI context tools (rtk, lean-ctx) have been removed from Headroom: ...
# 2. install.sh rejects the retired flags (extracted parse_wrap_args harness)
['--no-rtk', '--port', '9999'] rc=1 ERROR: CLI context tools ... Drop --no-rtk
['--context-tool=rtk'] rc=1 ERROR: CLI context tools ... Drop --context-tool
$ bash -n scripts/install.sh # syntax OK
# 3. Purge ran against the real machine, which had all the orphaned artifacts
$ python -c "from headroom.context_tool_cleanup import purge_context_tool_artifacts; ..."
removed ~/.headroom/bin/lean-ctx (51 MB)
removed ~/.headroom/bin/rtk (7.7 MB)
removed ~/.local/bin/rtk (symlink into ~/.headroom/bin)
removed ~/.claude/hooks/rtk-rewrite.sh
removed 8 lean-ctx-* hook scripts
# ~/.claude.json afterwards: 90 top-level keys, 19 projects, mcpServers unchanged
# → ~59 MB reclaimed, no unrelated key touched
# 4. stdout stays machine-readable while the purge reports (planted a fake artifact)
$ headroom wrap openclaw --prepare-only --gateway-provider-id codex >out 2>err
$ cat out
{"enabled":true,"config":{"proxyPort":8787,...}} # parses as JSON
$ cat err
Retired CLI context tool cleanup: removed /Users/tcms/.headroom/bin/rtk
# 5. --help is inert (planted artifact survives), a real run purges
$ headroom wrap codex --help → artifact survived: CORRECT
$ headroom wrap openclaw --prepare-only → purged: CORRECT
# 6. MCP purge dry-run against a copy of the real 82 KB ~/.claude.json
top-level keys 90 -> 90; projects 19 -> 19; LOST keys: none
all content outside mcpServers byte-identical: True
```
Dashboard rendered via the Playwright test after the panel removal:
"Token Savings" shows only `Proxy 0 (0.0%)` / `Of total wire: 36.86%`,
and "Token Usage" reads Before Compression → Proxy Removed → After
Compression with no "Filtered (this session)" row. Nothing below the
removed panel broke.
- **Not tested:** Windows and Linux (macOS only) — `install.ps1` is
verified by brace-balance and inspection, not executed, since no `pwsh`
is available locally. The wrap e2e suite (`e2e/wrap/run.py`) was updated
but not run; it needs the Docker e2e image. `serena project index`
interaction is exercised in the stacked base PR.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md` — it is generated by
release-please from my Conventional Commit PR title (a CI guard enforces
this)
## Additional Notes
**Stacked on #2676** (`tejas/serena-config-bootstrap`) — please merge
that first; this PR's base should then be retargeted to `main`, or it
will read as containing that fix too.
**Breaking-change migration for users:**
- Drop `--rtk`, `--no-rtk`, `--no-project-rtk`, `--keep-rtk`,
`--context-tool`, `--no-context-tool` from any alias, script or CI job,
and unset `HEADROOM_RTK*` / `HEADROOM_CONTEXT_TOOL`. They now error
rather than being ignored, so the failure is immediate and
self-explaining.
- Previously-installed artifacts are purged automatically on the next
`wrap`/`unwrap`; no manual cleanup needed.
- `headroom perf --json` no longer carries a `cli_filtering` key, and
`/stats` no longer returns a `context_tool` section.
**Docs:** `docs/rtk-architecture.md` deleted; RTK/lean-ctx removed from
`README.md`,
`docs/content/docs/{configuration,opencode,grok-build,docker-install,filesystem-contract}.mdx`,
`docs/observability.md` and the matching `wiki/` pages.
`REALIGNMENT/09-phase-G-rtk-observability.md` is marked SUPERSEDED
rather than deleted, to keep the planning record.
**Follow-ups not in scope:** `_emit_wrap_interrupted` was deleted as
dead code — its only caller was the `except KeyboardInterrupt` guarding
the binary download, so with no download there is nothing slow left to
interrupt.
2026-07-30 22:59:41 -07:00
|
|
|
import json
|
2026-04-23 07:39:52 -05:00
|
|
|
import sys
|
fix(subscription): run transcript token scan off the event loop (#1263)
## Description
The subscription tracker's poll loop scans Claude Code transcripts to
compute window-token usage. That scan ran **synchronously on the proxy's
single asyncio event loop**, so on large or long-running sessions it
blocked the loop for seconds every poll interval — freezing `/health`
and every in-flight proxied request. This moves the scan off the loop
with `asyncio.to_thread`.
Closes # <!-- no existing issue; root cause found via faulthandler.
Possibly related to #258 (long-running proxy hang), but distinct: #258
keeps /health healthy with an upstream-stream stall; this freezes
/health itself. -->
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [x] Performance improvement
## Changes Made
- `headroom/subscription/tracker.py` — `_maybe_poll()` now calls `await
asyncio.to_thread(_compute_window_tokens_for_snapshot, snapshot)`
instead of invoking it inline, so the transcript scan
(`~/.claude/projects/**/*.jsonl` read + `json.loads` per line) no longer
runs on the event-loop thread. The computed result is wired through
unchanged.
- `tests/test_subscription_tracker.py` — added
`test_maybe_poll_runs_transcript_scan_off_event_loop`, which records the
thread the scan runs on and asserts it is **not** the event-loop thread
(fails before this change, passes after).
- `CHANGELOG.md` — Unreleased → Bug Fixes entry.
## Root Cause
Captured with `faulthandler` (`SIGUSR1`) during a live wedge — the event
loop frozen mid-`json.loads`:
```
Current thread (most recent call first):
File ".../python3.14/json/decoder.py", line 361 in raw_decode
File ".../python3.14/json/__init__.py", line 352 in loads
File ".../headroom/subscription/session_tracking.py", line 127 in compute_window_tokens
File ".../headroom/subscription/tracker.py", line 872 in _compute_window_tokens_for_snapshot
File ".../headroom/subscription/tracker.py", line 731 in _maybe_poll
File ".../headroom/subscription/tracker.py", line 693 in _poll_loop
File ".../python3.14/asyncio/events.py", line 94 in _run
```
`_poll_loop` fires every `poll_interval_s` (default **300s**);
`compute_window_tokens` reads **every** `~/.claude/projects/**/*.jsonl`
transcript and `json.loads` each line. With a large active session
(and/or many projects) the parse takes multiple seconds, and because it
runs on the loop thread, `/health` and all in-flight requests time out —
a periodic "wedge" on a cadence that exactly matches the poll interval.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ uv run ruff check headroom/subscription/tracker.py tests/test_subscription_tracker.py
All checks passed!
$ uv run ruff format --check headroom/subscription/tracker.py tests/test_subscription_tracker.py
2 files already formatted
$ uv run mypy headroom/subscription/tracker.py
Success: no issues found in 1 source file
$ uv run pytest tests/test_subscription_tracker.py -q
...... [100%]
6 passed in 0.42s
# Regression test fails before the fix, passes after:
$ git stash push -- headroom/subscription/tracker.py # remove the fix
$ uv run pytest tests/test_subscription_tracker.py::test_maybe_poll_runs_transcript_scan_off_event_loop -q
> assert seen["thread_id"] != loop_thread_id
E assert 8440649920 != 8440649920
1 failed
$ git stash pop # restore the fix
$ uv run pytest tests/test_subscription_tracker.py::test_maybe_poll_runs_transcript_scan_off_event_loop -q
1 passed
```
## Real Behavior Proof
- **Environment:** macOS (Darwin 25), Python 3.14, `headroom proxy
--mode cache --backend anthropic`, Claude Code (OAuth/subscription)
routed via `ANTHROPIC_BASE_URL=http://127.0.0.1:8787`, a large,
long-running ~1M-token session.
- **Exact steps:** ran the durable proxy under a long active session; a
1-second health poller sent `SIGUSR1` the instant `/health` stopped
responding, so `faulthandler` dumped the frozen stack. Confirmed the
captured frame above. Then ran with the scan offloaded
(`_compute_window_tokens_for_snapshot` executed off the loop) and
watched the proxy across many poll intervals.
- **Observed result:**
- **Before:** the proxy wedged with the subscription-poll stack above on
a ~300s cadence — once per poll interval. `/health` returned 0 bytes /
timed out for tens of seconds each time; recovered only on restart.
- **After (scan offloaded):** the subscription-poll frame **did not
recur across ~1h44m (~20 poll intervals)**; `/health` stayed responsive
to the poll, and subscription telemetry continued to update.
- **Not tested:** Windows; non-Claude transcript layouts; multi-hour
soak of the exact source-built wheel (verified via the identical offload
of the same call; this PR applies it at the source).
- **Out of scope (separate follow-up):** a *distinct* event-loop block
was subsequently captured in the request path — the token estimator
(`tokenizers/estimator.py` → `tokenizers/base.py`
`count_messages`/`_count_content_parts` → `json.dumps`) runs
synchronously in `handle_anthropic_messages`. Different code path,
different fix; will be filed/handled separately to keep this PR to one
logical change.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review <!-- draft -->
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation (CHANGELOG)
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md
## Additional Notes
Single logical change. The fix preserves the telemetry result
(`_state.window_tokens`) unchanged; it only changes *where* the blocking
scan runs. No new dependencies. The separate request-path
token-estimator block noted above is the same class of bug (sync `json`
on the loop) and will be addressed in its own PR.
Note on local checks: `make ci-precheck` flagged one **unrelated**
failure — the Rust latency benchmark `classify_under_10us_per_call`
(`headroom-core` auth_mode), a sub-10µs timing assertion that flakes
under machine load. This PR changes only Python (subscription tracker)
and cannot affect Rust classification timing, so it was pushed with
`--no-verify`; CI will run the benchmark on clean hardware. Python
checks (`pytest`/`ruff`/`mypy`) all pass (output above).
2026-06-24 22:43:06 +08:00
|
|
|
import threading
|
fix(subscription): only reset 5h contribution on real rollover, not API jitter (#1255)
## Description
The 5-hour-window rollover detector in
`SubscriptionTracker._maybe_reset_contribution` zeroes the
`HeadroomContribution` counters on **every poll** instead of once per
window, so the dashboard's per-window savings figure stays pinned near
0%.
Root cause: the rollover check compared `five_hour.resets_at` between
consecutive polls with a bare `!=`. Anthropic's usage API reports that
timestamp with **second-level jitter** — on my account it flaps between
`01:59:59Z` and `02:00:00Z` on consecutive polls *within the same
window* — so the `!=` is true on essentially every poll and fires a
spurious `5h window rolled over; resetting headroom contribution
counters`.
The fix treats only a **forward jump larger than `_ROLLOVER_MIN_ADVANCE`
(1 minute)** as a genuine rollover. Jitter is sub-second; a real
rollover advances `resets_at` by ~5 hours, so the threshold cleanly
separates the two.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `headroom/subscription/tracker.py`: replaced the `curr_resets_at !=
prev_resets_at` rollover test with `curr_resets_at - prev_resets_at >
_ROLLOVER_MIN_ADVANCE`, and added the `_ROLLOVER_MIN_ADVANCE =
timedelta(minutes=1)` constant with a comment explaining the API jitter.
- `tests/test_subscription_tracker.py`: extended `_make_snapshot` to
accept an explicit `resets_at`; added
`test_second_level_reset_jitter_does_not_reset_contribution` (1-second
flap must NOT reset) and
`test_genuine_five_hour_rollover_resets_contribution` (5-hour jump still
resets).
- `CHANGELOG.md`: added a Bug Fixes entry under Unreleased.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ uv run pytest tests/test_subscription_tracker.py -v
test_tracker_notify_active_update_and_basic_state PASSED [ 14%]
test_tracker_start_stop_and_rollover_reset PASSED [ 28%]
test_second_level_reset_jitter_does_not_reset_contribution PASSED [ 42%]
test_genuine_five_hour_rollover_resets_contribution PASSED [ 57%]
test_maybe_poll_handles_inactive_and_none_snapshot PASSED [ 71%]
test_maybe_poll_success_updates_state_and_metrics PASSED [ 85%]
test_persist_and_load_state_round_trip PASSED [100%]
======================= 7 passed in 0.10s =======================
# Fails-before proof: stash the source fix, keep the new tests, re-run the jitter test
$ git stash push -- headroom/subscription/tracker.py
$ uv run pytest tests/test_subscription_tracker.py::test_second_level_reset_jitter_does_not_reset_contribution -q
E AssertionError: assert 0 == 99
E + where 0 = HeadroomContribution(tokens_submitted=0, ...).tokens_submitted
FAILED tests/test_subscription_tracker.py::test_second_level_reset_jitter_does_not_reset_contribution
1 failed in 0.09s
$ uv run ruff check headroom/subscription/tracker.py tests/test_subscription_tracker.py
All checks passed!
$ uv run mypy headroom/subscription/tracker.py
Success: no issues found in 1 source file
```
## Real Behavior Proof
- Environment: Linux (kernel 7.0), Python 3.14, `uv` 0.11.23, headroom
proxy on `127.0.0.1:8787`, Anthropic OAuth subscription account (Claude
Max), model `claude-opus-4-8`, Claude Code `claude-cli/2.1.185`.
- Exact command / steps: Inspected the live proxy on `main` before
patching. `~/.headroom/logs/proxy.log` contained 65 `5h window rolled
over; resetting headroom contribution counters` lines over a ~6h
session; I computed the gaps between consecutive events, and dumped
`five_hour.resets_at` from `~/.headroom/subscription_state.json`
history.
- Observed result: Median gap between resets was **exactly 300.0s** (=
the default `poll_interval_s`), not ~5h — i.e. it reset every poll. The
persisted history showed `five_hour.resets_at` flapping across only 4
distinct values, all within ~2s of `02:00:00Z` (e.g.
`2026-06-22T01:59:59Z` ↔ `2026-06-22T02:00:00Z`), and `contribution` was
all-zeros. After the patch, the unit tests reproduce this exact flap
(`base` vs `base + 1s`) and the counters are preserved; a genuine +5h
jump still resets.
- Not tested: I did not run the patched proxy live for a full 5-hour
window to observe a real rollover end-to-end (would need a multi-hour
session); the genuine-rollover path is covered by unit test only. No
change to the dashboard rendering code. Verified on Linux/Python 3.14
only.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Additional Notes
Docs checklist item is N/A — this is an internal accounting fix with no
user-facing API/config change. The threshold constant
(`_ROLLOVER_MIN_ADVANCE = 1 min`) is deliberately generous over the
observed sub-second jitter while remaining far below a real ~5h advance;
happy to tune or switch to an "old deadline has elapsed" guard
(`prev_resets_at <= now`) if maintainers prefer that framing.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-06-26 15:13:44 -04:00
|
|
|
from datetime import datetime, timedelta
|
feat: add Anthropic Claude Code subscription window tracking
- New headroom/subscription/ package:
- models.py: RateLimitWindow, ExtraUsage (cents->USD), SubscriptionSnapshot
with per-model 7d windows (opus/sonnet), HeadroomContribution,
WindowDiscrepancy, SubscriptionState
- client.py: async httpx client for GET /api/oauth/usage with OAuth token
resolution (env var -> credentials file -> live proxy header)
- tracker.py: background polling singleton (asyncio Task), notify_active(),
update_contribution(), anomaly detection (surge pricing, cache miss),
atomic persistence, OTEL callback
- session_tracking.py: JSONL transcript reader, Sonnet-normalised model
weights, compute_window_tokens() for 5h/7d window boundaries
- __init__.py: clean public re-exports
- Integration:
- proxy/models.py: subscription_tracking_enabled, poll_interval_s,
active_window_s config fields
- proxy/server.py: tracker startup/shutdown, /subscription-window endpoint,
subscription_window key in /stats
- proxy/handlers/anthropic.py: OAuth Bearer detection, notify_active() +
update_contribution() per request
- observability/metrics.py: 5 observable gauges (5h/7d utilisation %,
5h/7d seconds-to-reset, overage USD) using Observation type
- dashboard/templates/dashboard.html: subscription window panel with 5h/7d
progress bars, countdown, overage, Headroom contribution bar + anomaly
alerts
- Tests: 33 new unit tests in tests/test_subscription_tracker.py (all pass)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-04-11 21:53:44 -05:00
|
|
|
from pathlib import Path
|
2026-04-23 07:39:52 -05:00
|
|
|
from types import SimpleNamespace
|
2026-04-24 15:33:30 +02:00
|
|
|
|
feat: add Anthropic Claude Code subscription window tracking
- New headroom/subscription/ package:
- models.py: RateLimitWindow, ExtraUsage (cents->USD), SubscriptionSnapshot
with per-model 7d windows (opus/sonnet), HeadroomContribution,
WindowDiscrepancy, SubscriptionState
- client.py: async httpx client for GET /api/oauth/usage with OAuth token
resolution (env var -> credentials file -> live proxy header)
- tracker.py: background polling singleton (asyncio Task), notify_active(),
update_contribution(), anomaly detection (surge pricing, cache miss),
atomic persistence, OTEL callback
- session_tracking.py: JSONL transcript reader, Sonnet-normalised model
weights, compute_window_tokens() for 5h/7d window boundaries
- __init__.py: clean public re-exports
- Integration:
- proxy/models.py: subscription_tracking_enabled, poll_interval_s,
active_window_s config fields
- proxy/server.py: tracker startup/shutdown, /subscription-window endpoint,
subscription_window key in /stats
- proxy/handlers/anthropic.py: OAuth Bearer detection, notify_active() +
update_contribution() per request
- observability/metrics.py: 5 observable gauges (5h/7d utilisation %,
5h/7d seconds-to-reset, overage USD) using Observation type
- dashboard/templates/dashboard.html: subscription window panel with 5h/7d
progress bars, countdown, overage, Headroom contribution bar + anomaly
alerts
- Tests: 33 new unit tests in tests/test_subscription_tracker.py (all pass)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-04-11 21:53:44 -05:00
|
|
|
import pytest
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
import headroom.subscription.tracker as tracker_module
|
feat: add Anthropic Claude Code subscription window tracking
- New headroom/subscription/ package:
- models.py: RateLimitWindow, ExtraUsage (cents->USD), SubscriptionSnapshot
with per-model 7d windows (opus/sonnet), HeadroomContribution,
WindowDiscrepancy, SubscriptionState
- client.py: async httpx client for GET /api/oauth/usage with OAuth token
resolution (env var -> credentials file -> live proxy header)
- tracker.py: background polling singleton (asyncio Task), notify_active(),
update_contribution(), anomaly detection (surge pricing, cache miss),
atomic persistence, OTEL callback
- session_tracking.py: JSONL transcript reader, Sonnet-normalised model
weights, compute_window_tokens() for 5h/7d window boundaries
- __init__.py: clean public re-exports
- Integration:
- proxy/models.py: subscription_tracking_enabled, poll_interval_s,
active_window_s config fields
- proxy/server.py: tracker startup/shutdown, /subscription-window endpoint,
subscription_window key in /stats
- proxy/handlers/anthropic.py: OAuth Bearer detection, notify_active() +
update_contribution() per request
- observability/metrics.py: 5 observable gauges (5h/7d utilisation %,
5h/7d seconds-to-reset, overage USD) using Observation type
- dashboard/templates/dashboard.html: subscription window panel with 5h/7d
progress bars, countdown, overage, Headroom contribution bar + anomaly
alerts
- Tests: 33 new unit tests in tests/test_subscription_tracker.py (all pass)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-04-11 21:53:44 -05:00
|
|
|
from headroom.subscription.models import (
|
|
|
|
|
HeadroomContribution,
|
|
|
|
|
RateLimitWindow,
|
|
|
|
|
SubscriptionSnapshot,
|
|
|
|
|
WindowDiscrepancy,
|
|
|
|
|
WindowTokens,
|
2026-04-23 07:39:52 -05:00
|
|
|
_utc_now,
|
feat: add Anthropic Claude Code subscription window tracking
- New headroom/subscription/ package:
- models.py: RateLimitWindow, ExtraUsage (cents->USD), SubscriptionSnapshot
with per-model 7d windows (opus/sonnet), HeadroomContribution,
WindowDiscrepancy, SubscriptionState
- client.py: async httpx client for GET /api/oauth/usage with OAuth token
resolution (env var -> credentials file -> live proxy header)
- tracker.py: background polling singleton (asyncio Task), notify_active(),
update_contribution(), anomaly detection (surge pricing, cache miss),
atomic persistence, OTEL callback
- session_tracking.py: JSONL transcript reader, Sonnet-normalised model
weights, compute_window_tokens() for 5h/7d window boundaries
- __init__.py: clean public re-exports
- Integration:
- proxy/models.py: subscription_tracking_enabled, poll_interval_s,
active_window_s config fields
- proxy/server.py: tracker startup/shutdown, /subscription-window endpoint,
subscription_window key in /stats
- proxy/handlers/anthropic.py: OAuth Bearer detection, notify_active() +
update_contribution() per request
- observability/metrics.py: 5 observable gauges (5h/7d utilisation %,
5h/7d seconds-to-reset, overage USD) using Observation type
- dashboard/templates/dashboard.html: subscription window panel with 5h/7d
progress bars, countdown, overage, Headroom contribution bar + anomaly
alerts
- Tests: 33 new unit tests in tests/test_subscription_tracker.py (all pass)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-04-11 21:53:44 -05:00
|
|
|
)
|
2026-04-23 07:39:52 -05:00
|
|
|
from headroom.subscription.tracker import SubscriptionTracker
|
2026-04-24 15:33:30 +02:00
|
|
|
|
|
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
def _make_snapshot(
|
fix(subscription): only reset 5h contribution on real rollover, not API jitter (#1255)
## Description
The 5-hour-window rollover detector in
`SubscriptionTracker._maybe_reset_contribution` zeroes the
`HeadroomContribution` counters on **every poll** instead of once per
window, so the dashboard's per-window savings figure stays pinned near
0%.
Root cause: the rollover check compared `five_hour.resets_at` between
consecutive polls with a bare `!=`. Anthropic's usage API reports that
timestamp with **second-level jitter** — on my account it flaps between
`01:59:59Z` and `02:00:00Z` on consecutive polls *within the same
window* — so the `!=` is true on essentially every poll and fires a
spurious `5h window rolled over; resetting headroom contribution
counters`.
The fix treats only a **forward jump larger than `_ROLLOVER_MIN_ADVANCE`
(1 minute)** as a genuine rollover. Jitter is sub-second; a real
rollover advances `resets_at` by ~5 hours, so the threshold cleanly
separates the two.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `headroom/subscription/tracker.py`: replaced the `curr_resets_at !=
prev_resets_at` rollover test with `curr_resets_at - prev_resets_at >
_ROLLOVER_MIN_ADVANCE`, and added the `_ROLLOVER_MIN_ADVANCE =
timedelta(minutes=1)` constant with a comment explaining the API jitter.
- `tests/test_subscription_tracker.py`: extended `_make_snapshot` to
accept an explicit `resets_at`; added
`test_second_level_reset_jitter_does_not_reset_contribution` (1-second
flap must NOT reset) and
`test_genuine_five_hour_rollover_resets_contribution` (5-hour jump still
resets).
- `CHANGELOG.md`: added a Bug Fixes entry under Unreleased.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ uv run pytest tests/test_subscription_tracker.py -v
test_tracker_notify_active_update_and_basic_state PASSED [ 14%]
test_tracker_start_stop_and_rollover_reset PASSED [ 28%]
test_second_level_reset_jitter_does_not_reset_contribution PASSED [ 42%]
test_genuine_five_hour_rollover_resets_contribution PASSED [ 57%]
test_maybe_poll_handles_inactive_and_none_snapshot PASSED [ 71%]
test_maybe_poll_success_updates_state_and_metrics PASSED [ 85%]
test_persist_and_load_state_round_trip PASSED [100%]
======================= 7 passed in 0.10s =======================
# Fails-before proof: stash the source fix, keep the new tests, re-run the jitter test
$ git stash push -- headroom/subscription/tracker.py
$ uv run pytest tests/test_subscription_tracker.py::test_second_level_reset_jitter_does_not_reset_contribution -q
E AssertionError: assert 0 == 99
E + where 0 = HeadroomContribution(tokens_submitted=0, ...).tokens_submitted
FAILED tests/test_subscription_tracker.py::test_second_level_reset_jitter_does_not_reset_contribution
1 failed in 0.09s
$ uv run ruff check headroom/subscription/tracker.py tests/test_subscription_tracker.py
All checks passed!
$ uv run mypy headroom/subscription/tracker.py
Success: no issues found in 1 source file
```
## Real Behavior Proof
- Environment: Linux (kernel 7.0), Python 3.14, `uv` 0.11.23, headroom
proxy on `127.0.0.1:8787`, Anthropic OAuth subscription account (Claude
Max), model `claude-opus-4-8`, Claude Code `claude-cli/2.1.185`.
- Exact command / steps: Inspected the live proxy on `main` before
patching. `~/.headroom/logs/proxy.log` contained 65 `5h window rolled
over; resetting headroom contribution counters` lines over a ~6h
session; I computed the gaps between consecutive events, and dumped
`five_hour.resets_at` from `~/.headroom/subscription_state.json`
history.
- Observed result: Median gap between resets was **exactly 300.0s** (=
the default `poll_interval_s`), not ~5h — i.e. it reset every poll. The
persisted history showed `five_hour.resets_at` flapping across only 4
distinct values, all within ~2s of `02:00:00Z` (e.g.
`2026-06-22T01:59:59Z` ↔ `2026-06-22T02:00:00Z`), and `contribution` was
all-zeros. After the patch, the unit tests reproduce this exact flap
(`base` vs `base + 1s`) and the counters are preserved; a genuine +5h
jump still resets.
- Not tested: I did not run the patched proxy live for a full 5-hour
window to observe a real rollover end-to-end (would need a multi-hour
session); the genuine-rollover path is covered by unit test only. No
change to the dashboard rendering code. Verified on Linux/Python 3.14
only.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Additional Notes
Docs checklist item is N/A — this is an internal accounting fix with no
user-facing API/config change. The threshold constant
(`_ROLLOVER_MIN_ADVANCE = 1 min`) is deliberately generous over the
observed sub-second jitter while remaining far below a real ~5h advance;
happy to tune or switch to an "old deadline has elapsed" guard
(`prev_resets_at <= now`) if maintainers prefer that framing.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-06-26 15:13:44 -04:00
|
|
|
*,
|
|
|
|
|
token_prefix: str = "token123",
|
|
|
|
|
reset_offset_hours: int = 5,
|
|
|
|
|
resets_at: datetime | None = None,
|
2026-04-23 07:39:52 -05:00
|
|
|
) -> SubscriptionSnapshot:
|
|
|
|
|
return SubscriptionSnapshot(
|
|
|
|
|
five_hour=RateLimitWindow(
|
|
|
|
|
used=10,
|
|
|
|
|
limit=100,
|
|
|
|
|
utilization_pct=10.0,
|
fix(subscription): only reset 5h contribution on real rollover, not API jitter (#1255)
## Description
The 5-hour-window rollover detector in
`SubscriptionTracker._maybe_reset_contribution` zeroes the
`HeadroomContribution` counters on **every poll** instead of once per
window, so the dashboard's per-window savings figure stays pinned near
0%.
Root cause: the rollover check compared `five_hour.resets_at` between
consecutive polls with a bare `!=`. Anthropic's usage API reports that
timestamp with **second-level jitter** — on my account it flaps between
`01:59:59Z` and `02:00:00Z` on consecutive polls *within the same
window* — so the `!=` is true on essentially every poll and fires a
spurious `5h window rolled over; resetting headroom contribution
counters`.
The fix treats only a **forward jump larger than `_ROLLOVER_MIN_ADVANCE`
(1 minute)** as a genuine rollover. Jitter is sub-second; a real
rollover advances `resets_at` by ~5 hours, so the threshold cleanly
separates the two.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `headroom/subscription/tracker.py`: replaced the `curr_resets_at !=
prev_resets_at` rollover test with `curr_resets_at - prev_resets_at >
_ROLLOVER_MIN_ADVANCE`, and added the `_ROLLOVER_MIN_ADVANCE =
timedelta(minutes=1)` constant with a comment explaining the API jitter.
- `tests/test_subscription_tracker.py`: extended `_make_snapshot` to
accept an explicit `resets_at`; added
`test_second_level_reset_jitter_does_not_reset_contribution` (1-second
flap must NOT reset) and
`test_genuine_five_hour_rollover_resets_contribution` (5-hour jump still
resets).
- `CHANGELOG.md`: added a Bug Fixes entry under Unreleased.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ uv run pytest tests/test_subscription_tracker.py -v
test_tracker_notify_active_update_and_basic_state PASSED [ 14%]
test_tracker_start_stop_and_rollover_reset PASSED [ 28%]
test_second_level_reset_jitter_does_not_reset_contribution PASSED [ 42%]
test_genuine_five_hour_rollover_resets_contribution PASSED [ 57%]
test_maybe_poll_handles_inactive_and_none_snapshot PASSED [ 71%]
test_maybe_poll_success_updates_state_and_metrics PASSED [ 85%]
test_persist_and_load_state_round_trip PASSED [100%]
======================= 7 passed in 0.10s =======================
# Fails-before proof: stash the source fix, keep the new tests, re-run the jitter test
$ git stash push -- headroom/subscription/tracker.py
$ uv run pytest tests/test_subscription_tracker.py::test_second_level_reset_jitter_does_not_reset_contribution -q
E AssertionError: assert 0 == 99
E + where 0 = HeadroomContribution(tokens_submitted=0, ...).tokens_submitted
FAILED tests/test_subscription_tracker.py::test_second_level_reset_jitter_does_not_reset_contribution
1 failed in 0.09s
$ uv run ruff check headroom/subscription/tracker.py tests/test_subscription_tracker.py
All checks passed!
$ uv run mypy headroom/subscription/tracker.py
Success: no issues found in 1 source file
```
## Real Behavior Proof
- Environment: Linux (kernel 7.0), Python 3.14, `uv` 0.11.23, headroom
proxy on `127.0.0.1:8787`, Anthropic OAuth subscription account (Claude
Max), model `claude-opus-4-8`, Claude Code `claude-cli/2.1.185`.
- Exact command / steps: Inspected the live proxy on `main` before
patching. `~/.headroom/logs/proxy.log` contained 65 `5h window rolled
over; resetting headroom contribution counters` lines over a ~6h
session; I computed the gaps between consecutive events, and dumped
`five_hour.resets_at` from `~/.headroom/subscription_state.json`
history.
- Observed result: Median gap between resets was **exactly 300.0s** (=
the default `poll_interval_s`), not ~5h — i.e. it reset every poll. The
persisted history showed `five_hour.resets_at` flapping across only 4
distinct values, all within ~2s of `02:00:00Z` (e.g.
`2026-06-22T01:59:59Z` ↔ `2026-06-22T02:00:00Z`), and `contribution` was
all-zeros. After the patch, the unit tests reproduce this exact flap
(`base` vs `base + 1s`) and the counters are preserved; a genuine +5h
jump still resets.
- Not tested: I did not run the patched proxy live for a full 5-hour
window to observe a real rollover end-to-end (would need a multi-hour
session); the genuine-rollover path is covered by unit test only. No
change to the dashboard rendering code. Verified on Linux/Python 3.14
only.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Additional Notes
Docs checklist item is N/A — this is an internal accounting fix with no
user-facing API/config change. The threshold constant
(`_ROLLOVER_MIN_ADVANCE = 1 min`) is deliberately generous over the
observed sub-second jitter while remaining far below a real ~5h advance;
happy to tune or switch to an "old deadline has elapsed" guard
(`prev_resets_at <= now`) if maintainers prefer that framing.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-06-26 15:13:44 -04:00
|
|
|
resets_at=(
|
|
|
|
|
resets_at
|
|
|
|
|
if resets_at is not None
|
|
|
|
|
else _utc_now() + timedelta(hours=reset_offset_hours)
|
|
|
|
|
),
|
2026-04-23 07:39:52 -05:00
|
|
|
),
|
|
|
|
|
seven_day=RateLimitWindow(used=20, limit=200, utilization_pct=10.0),
|
|
|
|
|
token_prefix=token_prefix,
|
|
|
|
|
)
|
2026-04-24 15:33:30 +02:00
|
|
|
|
|
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
def test_tracker_notify_active_update_and_basic_state(monkeypatch: pytest.MonkeyPatch) -> None:
|
|
|
|
|
monkeypatch.setattr(SubscriptionTracker, "_load_persisted_state", lambda self: None)
|
|
|
|
|
tracker = SubscriptionTracker(enabled=False)
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
assert tracker.is_available() is False
|
|
|
|
|
assert tracker.latest_snapshot is None
|
|
|
|
|
assert tracker.is_active() is False
|
|
|
|
|
assert isinstance(tracker.get_stats(), dict)
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
tracker.notify_active("")
|
|
|
|
|
tracker.notify_active("Basic token")
|
|
|
|
|
tracker.notify_active("Bearer sk-ant-api-key")
|
|
|
|
|
assert tracker._current_token is None
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
tracker.notify_active("Bearer oauth-token-123")
|
|
|
|
|
assert tracker._current_token == "oauth-token-123"
|
|
|
|
|
assert tracker._full_tokens["oauth-to"] == 1
|
|
|
|
|
assert tracker.is_active() is True
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
tracker.update_contribution(
|
|
|
|
|
tokens_submitted=10,
|
|
|
|
|
tokens_saved_compression=5,
|
|
|
|
|
tokens_saved_cache_reads=3,
|
|
|
|
|
compression_savings_usd=1.25,
|
|
|
|
|
cache_savings_usd=-2.0,
|
|
|
|
|
)
|
|
|
|
|
contribution = tracker._state.contribution
|
|
|
|
|
assert contribution.tokens_submitted == 10
|
|
|
|
|
assert contribution.tokens_saved_compression == 5
|
|
|
|
|
assert contribution.tokens_saved_cache_reads == 3
|
2026-05-08 10:59:17 -07:00
|
|
|
assert contribution.to_dict()["tokens_saved"]["compression"] == 5
|
|
|
|
|
assert contribution.to_dict()["tokens_saved"]["proxy_compression"] == 5
|
2026-04-23 07:39:52 -05:00
|
|
|
assert contribution.compression_savings_usd == 1.25
|
|
|
|
|
assert contribution.cache_savings_usd == 0.0
|
2026-04-24 15:33:30 +02:00
|
|
|
|
|
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
@pytest.mark.asyncio
|
|
|
|
|
async def test_tracker_start_stop_and_rollover_reset(
|
|
|
|
|
monkeypatch: pytest.MonkeyPatch, tmp_path: Path
|
|
|
|
|
) -> None:
|
|
|
|
|
monkeypatch.setattr(SubscriptionTracker, "_load_persisted_state", lambda self: None)
|
|
|
|
|
tracker = SubscriptionTracker(persist_path=tmp_path / "state.json")
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
async def fake_poll_loop() -> None:
|
|
|
|
|
assert tracker._stop_event is not None
|
|
|
|
|
await tracker._stop_event.wait()
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
tracker._poll_loop = fake_poll_loop # type: ignore[method-assign]
|
|
|
|
|
await tracker.start()
|
|
|
|
|
first_task = tracker._poll_task
|
|
|
|
|
assert first_task is not None
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
await tracker.start()
|
|
|
|
|
assert tracker._poll_task is first_task
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
await tracker.stop()
|
|
|
|
|
assert tracker._stop_event is not None and tracker._stop_event.is_set()
|
|
|
|
|
assert tracker._persist_path.exists()
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
tracker._state.history = [
|
|
|
|
|
_make_snapshot(reset_offset_hours=5),
|
|
|
|
|
_make_snapshot(reset_offset_hours=6),
|
|
|
|
|
]
|
|
|
|
|
tracker._state.contribution = HeadroomContribution(tokens_submitted=99)
|
|
|
|
|
tracker._maybe_reset_contribution(tracker._state.history[-1])
|
|
|
|
|
assert tracker._state.contribution.tokens_submitted == 0
|
2026-04-24 15:33:30 +02:00
|
|
|
|
|
|
|
|
|
fix(subscription): only reset 5h contribution on real rollover, not API jitter (#1255)
## Description
The 5-hour-window rollover detector in
`SubscriptionTracker._maybe_reset_contribution` zeroes the
`HeadroomContribution` counters on **every poll** instead of once per
window, so the dashboard's per-window savings figure stays pinned near
0%.
Root cause: the rollover check compared `five_hour.resets_at` between
consecutive polls with a bare `!=`. Anthropic's usage API reports that
timestamp with **second-level jitter** — on my account it flaps between
`01:59:59Z` and `02:00:00Z` on consecutive polls *within the same
window* — so the `!=` is true on essentially every poll and fires a
spurious `5h window rolled over; resetting headroom contribution
counters`.
The fix treats only a **forward jump larger than `_ROLLOVER_MIN_ADVANCE`
(1 minute)** as a genuine rollover. Jitter is sub-second; a real
rollover advances `resets_at` by ~5 hours, so the threshold cleanly
separates the two.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `headroom/subscription/tracker.py`: replaced the `curr_resets_at !=
prev_resets_at` rollover test with `curr_resets_at - prev_resets_at >
_ROLLOVER_MIN_ADVANCE`, and added the `_ROLLOVER_MIN_ADVANCE =
timedelta(minutes=1)` constant with a comment explaining the API jitter.
- `tests/test_subscription_tracker.py`: extended `_make_snapshot` to
accept an explicit `resets_at`; added
`test_second_level_reset_jitter_does_not_reset_contribution` (1-second
flap must NOT reset) and
`test_genuine_five_hour_rollover_resets_contribution` (5-hour jump still
resets).
- `CHANGELOG.md`: added a Bug Fixes entry under Unreleased.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ uv run pytest tests/test_subscription_tracker.py -v
test_tracker_notify_active_update_and_basic_state PASSED [ 14%]
test_tracker_start_stop_and_rollover_reset PASSED [ 28%]
test_second_level_reset_jitter_does_not_reset_contribution PASSED [ 42%]
test_genuine_five_hour_rollover_resets_contribution PASSED [ 57%]
test_maybe_poll_handles_inactive_and_none_snapshot PASSED [ 71%]
test_maybe_poll_success_updates_state_and_metrics PASSED [ 85%]
test_persist_and_load_state_round_trip PASSED [100%]
======================= 7 passed in 0.10s =======================
# Fails-before proof: stash the source fix, keep the new tests, re-run the jitter test
$ git stash push -- headroom/subscription/tracker.py
$ uv run pytest tests/test_subscription_tracker.py::test_second_level_reset_jitter_does_not_reset_contribution -q
E AssertionError: assert 0 == 99
E + where 0 = HeadroomContribution(tokens_submitted=0, ...).tokens_submitted
FAILED tests/test_subscription_tracker.py::test_second_level_reset_jitter_does_not_reset_contribution
1 failed in 0.09s
$ uv run ruff check headroom/subscription/tracker.py tests/test_subscription_tracker.py
All checks passed!
$ uv run mypy headroom/subscription/tracker.py
Success: no issues found in 1 source file
```
## Real Behavior Proof
- Environment: Linux (kernel 7.0), Python 3.14, `uv` 0.11.23, headroom
proxy on `127.0.0.1:8787`, Anthropic OAuth subscription account (Claude
Max), model `claude-opus-4-8`, Claude Code `claude-cli/2.1.185`.
- Exact command / steps: Inspected the live proxy on `main` before
patching. `~/.headroom/logs/proxy.log` contained 65 `5h window rolled
over; resetting headroom contribution counters` lines over a ~6h
session; I computed the gaps between consecutive events, and dumped
`five_hour.resets_at` from `~/.headroom/subscription_state.json`
history.
- Observed result: Median gap between resets was **exactly 300.0s** (=
the default `poll_interval_s`), not ~5h — i.e. it reset every poll. The
persisted history showed `five_hour.resets_at` flapping across only 4
distinct values, all within ~2s of `02:00:00Z` (e.g.
`2026-06-22T01:59:59Z` ↔ `2026-06-22T02:00:00Z`), and `contribution` was
all-zeros. After the patch, the unit tests reproduce this exact flap
(`base` vs `base + 1s`) and the counters are preserved; a genuine +5h
jump still resets.
- Not tested: I did not run the patched proxy live for a full 5-hour
window to observe a real rollover end-to-end (would need a multi-hour
session); the genuine-rollover path is covered by unit test only. No
change to the dashboard rendering code. Verified on Linux/Python 3.14
only.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Additional Notes
Docs checklist item is N/A — this is an internal accounting fix with no
user-facing API/config change. The threshold constant
(`_ROLLOVER_MIN_ADVANCE = 1 min`) is deliberately generous over the
observed sub-second jitter while remaining far below a real ~5h advance;
happy to tune or switch to an "old deadline has elapsed" guard
(`prev_resets_at <= now`) if maintainers prefer that framing.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-06-26 15:13:44 -04:00
|
|
|
def test_second_level_reset_jitter_does_not_reset_contribution(
|
|
|
|
|
monkeypatch: pytest.MonkeyPatch, tmp_path: Path
|
|
|
|
|
) -> None:
|
|
|
|
|
"""Regression: the usage API reports ``resets_at`` with second-level jitter
|
|
|
|
|
within a single window (observed flapping between ``01:59:59Z`` and
|
|
|
|
|
``02:00:00Z`` on consecutive polls). That must NOT be treated as a rollover,
|
|
|
|
|
or the contribution counters get zeroed every poll and the dashboard sticks
|
|
|
|
|
at ~0% savings.
|
|
|
|
|
"""
|
|
|
|
|
monkeypatch.setattr(SubscriptionTracker, "_load_persisted_state", lambda self: None)
|
|
|
|
|
tracker = SubscriptionTracker(persist_path=tmp_path / "state.json")
|
|
|
|
|
|
|
|
|
|
base = _utc_now() + timedelta(hours=3)
|
|
|
|
|
tracker._state.history = [
|
|
|
|
|
_make_snapshot(resets_at=base),
|
|
|
|
|
_make_snapshot(resets_at=base + timedelta(seconds=1)),
|
|
|
|
|
]
|
|
|
|
|
tracker._state.contribution = HeadroomContribution(tokens_submitted=99)
|
|
|
|
|
tracker._maybe_reset_contribution(tracker._state.history[-1])
|
|
|
|
|
assert tracker._state.contribution.tokens_submitted == 99
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_genuine_five_hour_rollover_resets_contribution(
|
|
|
|
|
monkeypatch: pytest.MonkeyPatch, tmp_path: Path
|
|
|
|
|
) -> None:
|
|
|
|
|
"""A real rollover advances ``resets_at`` by ~5 hours and still resets."""
|
|
|
|
|
monkeypatch.setattr(SubscriptionTracker, "_load_persisted_state", lambda self: None)
|
|
|
|
|
tracker = SubscriptionTracker(persist_path=tmp_path / "state.json")
|
|
|
|
|
|
|
|
|
|
base = _utc_now()
|
|
|
|
|
tracker._state.history = [
|
|
|
|
|
_make_snapshot(resets_at=base),
|
|
|
|
|
_make_snapshot(resets_at=base + timedelta(hours=5)),
|
|
|
|
|
]
|
|
|
|
|
tracker._state.contribution = HeadroomContribution(tokens_submitted=99)
|
|
|
|
|
tracker._maybe_reset_contribution(tracker._state.history[-1])
|
|
|
|
|
assert tracker._state.contribution.tokens_submitted == 0
|
|
|
|
|
|
|
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
@pytest.mark.asyncio
|
|
|
|
|
async def test_maybe_poll_handles_inactive_and_none_snapshot(
|
|
|
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
|
|
|
) -> None:
|
|
|
|
|
monkeypatch.setattr(SubscriptionTracker, "_load_persisted_state", lambda self: None)
|
|
|
|
|
tracker = SubscriptionTracker()
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
monkeypatch.setattr("headroom.subscription.client.read_cached_oauth_token", lambda: None)
|
|
|
|
|
await tracker._maybe_poll()
|
|
|
|
|
assert tracker._state.poll_count == 0
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
monkeypatch.setattr(
|
|
|
|
|
"headroom.subscription.client.read_cached_oauth_token", lambda: "cached-token"
|
|
|
|
|
)
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
async def fetch_none(token: str | None):
|
|
|
|
|
return None
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
tracker._client = SimpleNamespace(fetch=fetch_none)
|
|
|
|
|
await tracker._maybe_poll()
|
|
|
|
|
assert tracker._state.last_error == "fetch returned None"
|
|
|
|
|
assert tracker._state.poll_errors == 1
|
2026-04-24 15:33:30 +02:00
|
|
|
|
|
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
@pytest.mark.asyncio
|
|
|
|
|
async def test_maybe_poll_success_updates_state_and_metrics(
|
|
|
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
|
|
|
) -> None:
|
|
|
|
|
monkeypatch.setattr(SubscriptionTracker, "_load_persisted_state", lambda self: None)
|
|
|
|
|
tracker = SubscriptionTracker()
|
|
|
|
|
tracker.notify_active("Bearer live-oauth-token")
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
snapshot = _make_snapshot()
|
|
|
|
|
discrepancies = [WindowDiscrepancy(kind="cache_miss", description="miss", severity="warning")]
|
|
|
|
|
metrics_calls: list[dict] = []
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
async def fetch_snapshot(token: str | None):
|
|
|
|
|
assert token == "live-oauth-token"
|
|
|
|
|
return snapshot
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
tracker._client = SimpleNamespace(fetch=fetch_snapshot)
|
|
|
|
|
monkeypatch.setattr(
|
|
|
|
|
tracker_module, "_compute_window_tokens_for_snapshot", lambda snap: WindowTokens(input=7)
|
|
|
|
|
)
|
|
|
|
|
monkeypatch.setattr(tracker_module, "_detect_discrepancies", lambda snap, tokens: discrepancies)
|
|
|
|
|
monkeypatch.setattr(
|
|
|
|
|
tracker, "_persist_state", lambda: metrics_calls.append({"persisted": True})
|
|
|
|
|
)
|
|
|
|
|
monkeypatch.setitem(
|
|
|
|
|
sys.modules,
|
|
|
|
|
"headroom.observability.metrics",
|
|
|
|
|
SimpleNamespace(
|
|
|
|
|
get_otel_metrics=lambda: SimpleNamespace(
|
|
|
|
|
record_subscription_window=lambda state: metrics_calls.append(state)
|
2026-04-18 01:18:23 +07:00
|
|
|
)
|
2026-04-23 07:39:52 -05:00
|
|
|
),
|
|
|
|
|
)
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
await tracker._maybe_poll()
|
|
|
|
|
assert tracker.latest_snapshot is snapshot
|
|
|
|
|
assert tracker._state.window_tokens.input == 7
|
|
|
|
|
assert tracker._state.discrepancies[-1].kind == "cache_miss"
|
|
|
|
|
assert tracker._state.last_error is None
|
|
|
|
|
assert tracker._state.poll_count == 1
|
|
|
|
|
assert metrics_calls[0] == {"persisted": True}
|
|
|
|
|
assert isinstance(metrics_calls[1], dict)
|
2026-04-24 15:33:30 +02:00
|
|
|
|
|
|
|
|
|
fix(subscription): run transcript token scan off the event loop (#1263)
## Description
The subscription tracker's poll loop scans Claude Code transcripts to
compute window-token usage. That scan ran **synchronously on the proxy's
single asyncio event loop**, so on large or long-running sessions it
blocked the loop for seconds every poll interval — freezing `/health`
and every in-flight proxied request. This moves the scan off the loop
with `asyncio.to_thread`.
Closes # <!-- no existing issue; root cause found via faulthandler.
Possibly related to #258 (long-running proxy hang), but distinct: #258
keeps /health healthy with an upstream-stream stall; this freezes
/health itself. -->
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [x] Performance improvement
## Changes Made
- `headroom/subscription/tracker.py` — `_maybe_poll()` now calls `await
asyncio.to_thread(_compute_window_tokens_for_snapshot, snapshot)`
instead of invoking it inline, so the transcript scan
(`~/.claude/projects/**/*.jsonl` read + `json.loads` per line) no longer
runs on the event-loop thread. The computed result is wired through
unchanged.
- `tests/test_subscription_tracker.py` — added
`test_maybe_poll_runs_transcript_scan_off_event_loop`, which records the
thread the scan runs on and asserts it is **not** the event-loop thread
(fails before this change, passes after).
- `CHANGELOG.md` — Unreleased → Bug Fixes entry.
## Root Cause
Captured with `faulthandler` (`SIGUSR1`) during a live wedge — the event
loop frozen mid-`json.loads`:
```
Current thread (most recent call first):
File ".../python3.14/json/decoder.py", line 361 in raw_decode
File ".../python3.14/json/__init__.py", line 352 in loads
File ".../headroom/subscription/session_tracking.py", line 127 in compute_window_tokens
File ".../headroom/subscription/tracker.py", line 872 in _compute_window_tokens_for_snapshot
File ".../headroom/subscription/tracker.py", line 731 in _maybe_poll
File ".../headroom/subscription/tracker.py", line 693 in _poll_loop
File ".../python3.14/asyncio/events.py", line 94 in _run
```
`_poll_loop` fires every `poll_interval_s` (default **300s**);
`compute_window_tokens` reads **every** `~/.claude/projects/**/*.jsonl`
transcript and `json.loads` each line. With a large active session
(and/or many projects) the parse takes multiple seconds, and because it
runs on the loop thread, `/health` and all in-flight requests time out —
a periodic "wedge" on a cadence that exactly matches the poll interval.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ uv run ruff check headroom/subscription/tracker.py tests/test_subscription_tracker.py
All checks passed!
$ uv run ruff format --check headroom/subscription/tracker.py tests/test_subscription_tracker.py
2 files already formatted
$ uv run mypy headroom/subscription/tracker.py
Success: no issues found in 1 source file
$ uv run pytest tests/test_subscription_tracker.py -q
...... [100%]
6 passed in 0.42s
# Regression test fails before the fix, passes after:
$ git stash push -- headroom/subscription/tracker.py # remove the fix
$ uv run pytest tests/test_subscription_tracker.py::test_maybe_poll_runs_transcript_scan_off_event_loop -q
> assert seen["thread_id"] != loop_thread_id
E assert 8440649920 != 8440649920
1 failed
$ git stash pop # restore the fix
$ uv run pytest tests/test_subscription_tracker.py::test_maybe_poll_runs_transcript_scan_off_event_loop -q
1 passed
```
## Real Behavior Proof
- **Environment:** macOS (Darwin 25), Python 3.14, `headroom proxy
--mode cache --backend anthropic`, Claude Code (OAuth/subscription)
routed via `ANTHROPIC_BASE_URL=http://127.0.0.1:8787`, a large,
long-running ~1M-token session.
- **Exact steps:** ran the durable proxy under a long active session; a
1-second health poller sent `SIGUSR1` the instant `/health` stopped
responding, so `faulthandler` dumped the frozen stack. Confirmed the
captured frame above. Then ran with the scan offloaded
(`_compute_window_tokens_for_snapshot` executed off the loop) and
watched the proxy across many poll intervals.
- **Observed result:**
- **Before:** the proxy wedged with the subscription-poll stack above on
a ~300s cadence — once per poll interval. `/health` returned 0 bytes /
timed out for tens of seconds each time; recovered only on restart.
- **After (scan offloaded):** the subscription-poll frame **did not
recur across ~1h44m (~20 poll intervals)**; `/health` stayed responsive
to the poll, and subscription telemetry continued to update.
- **Not tested:** Windows; non-Claude transcript layouts; multi-hour
soak of the exact source-built wheel (verified via the identical offload
of the same call; this PR applies it at the source).
- **Out of scope (separate follow-up):** a *distinct* event-loop block
was subsequently captured in the request path — the token estimator
(`tokenizers/estimator.py` → `tokenizers/base.py`
`count_messages`/`_count_content_parts` → `json.dumps`) runs
synchronously in `handle_anthropic_messages`. Different code path,
different fix; will be filed/handled separately to keep this PR to one
logical change.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review <!-- draft -->
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation (CHANGELOG)
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md
## Additional Notes
Single logical change. The fix preserves the telemetry result
(`_state.window_tokens`) unchanged; it only changes *where* the blocking
scan runs. No new dependencies. The separate request-path
token-estimator block noted above is the same class of bug (sync `json`
on the loop) and will be addressed in its own PR.
Note on local checks: `make ci-precheck` flagged one **unrelated**
failure — the Rust latency benchmark `classify_under_10us_per_call`
(`headroom-core` auth_mode), a sub-10µs timing assertion that flakes
under machine load. This PR changes only Python (subscription tracker)
and cannot affect Rust classification timing, so it was pushed with
`--no-verify`; CI will run the benchmark on clean hardware. Python
checks (`pytest`/`ruff`/`mypy`) all pass (output above).
2026-06-24 22:43:06 +08:00
|
|
|
@pytest.mark.asyncio
|
|
|
|
|
async def test_maybe_poll_runs_transcript_scan_off_event_loop(
|
|
|
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
|
|
|
) -> None:
|
|
|
|
|
"""Regression: the transcript scan must run off the event-loop thread, or a
|
|
|
|
|
multi-second ~/.claude/projects scan wedges the proxy every poll interval."""
|
|
|
|
|
monkeypatch.setattr(SubscriptionTracker, "_load_persisted_state", lambda self: None)
|
|
|
|
|
tracker = SubscriptionTracker()
|
|
|
|
|
tracker.notify_active("Bearer live-oauth-token")
|
|
|
|
|
|
|
|
|
|
snapshot = _make_snapshot()
|
|
|
|
|
|
|
|
|
|
async def fetch_snapshot(token: str | None) -> SubscriptionSnapshot:
|
|
|
|
|
return snapshot
|
|
|
|
|
|
|
|
|
|
tracker._client = SimpleNamespace(fetch=fetch_snapshot)
|
|
|
|
|
|
|
|
|
|
loop_thread_id = threading.get_ident()
|
|
|
|
|
seen: dict[str, int] = {}
|
|
|
|
|
|
|
|
|
|
def recording_compute(snap: SubscriptionSnapshot) -> WindowTokens:
|
|
|
|
|
seen["thread_id"] = threading.get_ident()
|
|
|
|
|
return WindowTokens(input=7)
|
|
|
|
|
|
|
|
|
|
monkeypatch.setattr(tracker_module, "_compute_window_tokens_for_snapshot", recording_compute)
|
|
|
|
|
monkeypatch.setattr(tracker_module, "_detect_discrepancies", lambda snap, tokens: [])
|
|
|
|
|
monkeypatch.setattr(tracker, "_persist_state", lambda: None)
|
|
|
|
|
monkeypatch.setitem(
|
|
|
|
|
sys.modules,
|
|
|
|
|
"headroom.observability.metrics",
|
|
|
|
|
SimpleNamespace(
|
|
|
|
|
get_otel_metrics=lambda: SimpleNamespace(record_subscription_window=lambda state: None)
|
|
|
|
|
),
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
await tracker._maybe_poll()
|
|
|
|
|
|
|
|
|
|
# The blocking scan ran on a worker thread, not the event-loop thread.
|
|
|
|
|
assert seen["thread_id"] != loop_thread_id
|
|
|
|
|
|
|
|
|
|
|
fix(subscription): wire tokens_saved_rtk from RTK stats endpoint
PR-G2 (Realignment) — retire the dead `tokens_saved_rtk` data plane.
Previously, `SubscriptionTracker.update_contribution` silently mirrored
`tokens_saved_cli_filtering` into `tokens_saved_rtk`, making the two
counters identical at all times and hiding wrap-side RTK savings from
the dashboard.
Wiring:
- New `_last_rtk_tokens_saved` state on the tracker (init to 0).
- `update_contribution` now polls `_get_rtk_stats()` when the caller
omits an explicit `tokens_saved_rtk`, computes the delta against the
last lifetime total, and writes only the positive delta. State
advances monotonically; a counter regression re-baselines without
emitting a negative delta.
- `cli_filtering` and `rtk` are no longer aliased anywhere in the
hot path.
- Persistence: `to_dict()` exposes raw `cli_filtering_raw` and
`rtk_raw` keys (legacy dashboard `cli_filtering` / `rtk` still report
`max(cli, rtk)` for back-compat). `_load_persisted_state` reads the
raw keys when present and defaults to 0 otherwise so legacy state
cannot silently re-inflate `tokens_saved_rtk` by mirroring
`cli_filtering`.
Build constraints honoured:
- No silent fallback — transient `_get_rtk_stats()` exceptions are
caught, structured-logged (`event=subscription_rtk_stats_fetch_failed`
/ `event=subscription_rtk_stats_unavailable`), and yield zero delta.
- Configurable — `HEADROOM_RTK_WIRING={enabled,disabled}` opts the
polling out without disturbing tool selection. Unknown values raise
loudly via `_rtk_wiring_mode`.
- Comprehensive tests — 9 new unit tests in
`tests/test_subscription_tracker_rtk_wired.py` pin the wiring
(baseline, delta across two/three polls, None payload, monotonic
advance, counter regression, exception, env-var opt-out, explicit
override, decoupling from cli_filtering). Existing tracker tests
updated to reflect the no-mirror behaviour.
Refs: REALIGNMENT/09-phase-G-rtk-observability.md (PR-G2)
2026-05-22 13:48:42 -07:00
|
|
|
def test_persist_and_load_state_round_trip(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> None:
|
2026-04-23 07:39:52 -05:00
|
|
|
persist_path = tmp_path / "tracker-state.json"
|
|
|
|
|
tracker = SubscriptionTracker(persist_path=persist_path)
|
|
|
|
|
tracker.update_contribution(
|
|
|
|
|
tokens_submitted=11,
|
|
|
|
|
tokens_saved_compression=2,
|
|
|
|
|
tokens_saved_cache_reads=4,
|
|
|
|
|
compression_savings_usd=1.5,
|
|
|
|
|
cache_savings_usd=2.5,
|
|
|
|
|
)
|
|
|
|
|
tracker._state.poll_count = 7
|
|
|
|
|
tracker._persist_state()
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
loader = SubscriptionTracker(persist_path=persist_path)
|
|
|
|
|
assert loader._state.contribution.tokens_submitted == 11
|
|
|
|
|
assert loader._state.contribution.tokens_saved_compression == 2
|
|
|
|
|
assert loader._state.contribution.tokens_saved_cache_reads == 4
|
fix: remove rtk and lean-ctx CLI context tools (#2677)
## Description
Removes both third-party CLI context tools — **rtk** and **lean-ctx** —
and with them the context-tool selector itself. Headroom no longer
downloads, installs or configures either one, and there is no
replacement.
The previous pass (#2344) gated only three entry points inside
`headroom/cli/wrap.py`. That left the feature reachable in practice:
| Gap | Effect |
|---|---|
| `scripts/install.sh:1544`, `install.ps1:1681` | Ran `rtk init --global
--auto-patch` from bash/PowerShell, **bypassing the Python gate
entirely** — `curl \| sh` still wrote a Claude Code `PreToolUse` hook
regardless of `HEADROOM_RTK` |
| `wrap.py` `_setup_context_tool_for_agent` | **`wrap openhands` was
broken by default**: `rtk_required=True` met a gate returning `None` →
`SystemExit(1)`. Invisible because all 8 openhands tests patched
`_ensure_rtk_binary` to a fake path |
| `proxy/helpers.py`, `subscription/tracker.py` | Proxy shelled out to
`rtk gain` from `/stats`, the dashboard and `headroom perf`; the tracker
polled it per contribution (`_RTK_WIRING_DEFAULT = "enabled"`) |
| No cleanup path | Nothing removed artifacts an earlier default had
installed, so a machine that once ran the old default kept rtk in the
loop forever (#1669, #1955) |
Also worth noting: the rtk binary download had **no SHA or signature
verification** — only `rtk --version` as a smoke test.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [x] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [x] Code refactoring (no functional changes)
## Changes Made
**Removed** — `headroom/rtk/` and `headroom/lean_ctx/` packages,
`headroom/cli/wrap_rtk_metrics.py`, `_selected_context_tool` /
`_setup_context_tool_for_agent` / `_VALID_CONTEXT_TOOLS`, the `--rtk` /
`--no-rtk` / `--no-project-rtk` / `--keep-rtk` flags across all 18 wrap
subcommands, `HEADROOM_RTK*`, the proxy-side `rtk gain` polling, the
dashboard CLI-filtering panel (rows + all 8 `cliFiltering*` Alpine
getters), `paths.rtk_path()` / `lean_ctx_path()`, the SDK path helpers,
`benchmarks/rtk_loop_learn_eval.py`, and the `headroom/rtk/**` CI path
filters.
**Fails loudly, not silently** — `--context-tool` / `--no-context-tool`
/ `HEADROOM_CONTEXT_TOOL` are kept solely to error out. They live in
shell profiles, aliases and CI jobs, and accepting them as a no-op would
read as Headroom having quietly stopped working. The installers reject
them too, which matters more than it looks: their arg parsers forward
the first unknown flag **and everything after it** to the wrapped tool,
so a leftover `--no-rtk` would have silently swallowed a following
`--port` and then been ignored downstream.
**New `headroom/context_tool_cleanup.py`** — deleting the code cannot
help a machine that already ran the old default, since the hooks,
binaries and injected guidance are durable on disk.
`purge_context_tool_artifacts()` runs once per `wrap`/`unwrap` and
removes the registered hook entries, the generated hook scripts, the
Headroom-managed `~/.local/bin` symlinks, the vendored
`~/.headroom/bin/{rtk,lean-ctx}` binaries, the `lean-ctx` MCP server
entry and the marker-fenced instruction blocks. Deliberately
conservative: idempotent, **skips** a malformed config rather than
overwriting it, and only unlinks a symlink resolving inside Headroom's
own bin dir so a user's own build is untouched. It reports on
**stderr**, because `wrap/unwrap openclaw --prepare-only` emit
machine-readable JSON on stdout as their entire contract. Skipped for
`wrap selfheal` (runs from a SessionStart hook; must not race Claude
Code's writer for `~/.claude.json`) and for `--help`, which must stay
read-only.
**Client-config hardening** (discovered while investigating a "corrupted
Serena settings file" report) — `wrap.py` reset a settings file to `{}`
when an existing file would not parse, then wrote that back. One
hand-edited typo or a transient `EACCES`/`EINTR` on a valid file
destroyed the user's `permissions`, `env` and `hooks`, on **every
`headroom wrap claude`**. It now refuses to write. Separately,
`fsutil.write_text` is now atomic (temp file + `fsync` + `os.replace`),
fixing all 14 non-atomic client-config writes at once; it follows
symlinks rather than replacing them (dotfile managers) and preserves an
existing file's mode.
**Deliberately kept** — `rtk` stays in the wrapper-peel list in
`transforms/content_router.py`. It sits beside `sudo`/`env`/`timeout` as
shell-command grammar, so `rtk cat f` is still classified as a file read
for anyone running their own rtk install, which the purge intentionally
leaves alone.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ ruff check headroom/ tests/ e2e/ --exclude headroom/dashboard/templates
All checks passed!
$ ruff format --check headroom/ tests/ e2e/ --exclude headroom/dashboard/templates
1255 files already formatted
$ mypy headroom/
Success: no issues found in 508 source files
$ pytest tests/test_context_tool_cleanup.py -q
11 passed
$ pytest tests/test_fsutil.py -q
12 passed
$ pytest tests/test_cli/test_wrap_codex.py -q # 89 tests
89 passed in 431.68s
$ pytest tests/test_cli/test_wrap_opencode.py -q
39 passed in 257.46s
$ pytest tests/test_cli/test_wrap_helpers.py -q
45 passed
$ pytest tests/test_paths.py -q
75 passed
$ pytest tests/test_cli/test_unwrap_claude.py -q
14 passed
$ pytest tests/test_proxy_savings_history.py -q
39 passed
$ pytest tests/test_cli/test_wrap_copilot.py -q
27 passed
$ pytest tests/test_cli/test_wrap_zcode.py -q
20 passed
$ pytest tests/test_subscription_tracker.py -q
9 passed
$ pytest tests/test_proxy_dashboard_stats_cache.py -q
5 passed, 1 skipped
```
Repo-wide grep for 14 removed symbols (`headroom.rtk`,
`headroom.lean_ctx`, `_ensure_rtk_binary`, `_selected_context_tool`,
`_get_context_tool_stats`, `rtk_path`, `lean_ctx_path`,
`wrap_rtk_metrics`, `HEADROOM_RTK`, `cli_tokens_avoided`,
`tokens_saved_rtk`, …) across `*.py`, `*.ts`, `*.sh`, `*.ps1`, `*.yml`,
`*.html`: **zero hits**.
Notable test changes: `test_wrap_openhands.py` no longer patches
`_ensure_rtk_binary` and asserts `wrap openhands --prepare-only` exits 0
unpatched — the regression that was previously masked.
`test_wrap_continue.py` and `test_wrap_hintfile_agents.py` were removed
(every test drove RTK instruction injection). A new
`test_subscription_tracker.py::test_load_state_written_before_cli_context_tools_were_removed`
proves a pre-removal `subscription_state.json` still loads.
## Real Behavior Proof
- **Environment:** macOS 15.4 (darwin 25.4.0), Python 3.12.6, Headroom @
this branch, real `~/.headroom` and `~/.claude` on the dev machine.
- **Exact command / steps and observed result:**
```text
# 1. Retired flag fails loudly instead of silently no-op'ing
$ headroom wrap codex --prepare-only --context-tool rtk
Error: CLI context tools (rtk, lean-ctx) have been removed from Headroom: they
rewrote shell commands through a third-party binary Headroom no longer manages.
Drop --context-tool / --no-context-tool and unset HEADROOM_CONTEXT_TOOL;
`headroom wrap` uninstalls what they left behind on first run.
$ HEADROOM_CONTEXT_TOOL=lean-ctx headroom wrap codex --prepare-only
Error: CLI context tools (rtk, lean-ctx) have been removed from Headroom: ...
# 2. install.sh rejects the retired flags (extracted parse_wrap_args harness)
['--no-rtk', '--port', '9999'] rc=1 ERROR: CLI context tools ... Drop --no-rtk
['--context-tool=rtk'] rc=1 ERROR: CLI context tools ... Drop --context-tool
$ bash -n scripts/install.sh # syntax OK
# 3. Purge ran against the real machine, which had all the orphaned artifacts
$ python -c "from headroom.context_tool_cleanup import purge_context_tool_artifacts; ..."
removed ~/.headroom/bin/lean-ctx (51 MB)
removed ~/.headroom/bin/rtk (7.7 MB)
removed ~/.local/bin/rtk (symlink into ~/.headroom/bin)
removed ~/.claude/hooks/rtk-rewrite.sh
removed 8 lean-ctx-* hook scripts
# ~/.claude.json afterwards: 90 top-level keys, 19 projects, mcpServers unchanged
# → ~59 MB reclaimed, no unrelated key touched
# 4. stdout stays machine-readable while the purge reports (planted a fake artifact)
$ headroom wrap openclaw --prepare-only --gateway-provider-id codex >out 2>err
$ cat out
{"enabled":true,"config":{"proxyPort":8787,...}} # parses as JSON
$ cat err
Retired CLI context tool cleanup: removed /Users/tcms/.headroom/bin/rtk
# 5. --help is inert (planted artifact survives), a real run purges
$ headroom wrap codex --help → artifact survived: CORRECT
$ headroom wrap openclaw --prepare-only → purged: CORRECT
# 6. MCP purge dry-run against a copy of the real 82 KB ~/.claude.json
top-level keys 90 -> 90; projects 19 -> 19; LOST keys: none
all content outside mcpServers byte-identical: True
```
Dashboard rendered via the Playwright test after the panel removal:
"Token Savings" shows only `Proxy 0 (0.0%)` / `Of total wire: 36.86%`,
and "Token Usage" reads Before Compression → Proxy Removed → After
Compression with no "Filtered (this session)" row. Nothing below the
removed panel broke.
- **Not tested:** Windows and Linux (macOS only) — `install.ps1` is
verified by brace-balance and inspection, not executed, since no `pwsh`
is available locally. The wrap e2e suite (`e2e/wrap/run.py`) was updated
but not run; it needs the Docker e2e image. `serena project index`
interaction is exercised in the stacked base PR.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md` — it is generated by
release-please from my Conventional Commit PR title (a CI guard enforces
this)
## Additional Notes
**Stacked on #2676** (`tejas/serena-config-bootstrap`) — please merge
that first; this PR's base should then be retargeted to `main`, or it
will read as containing that fix too.
**Breaking-change migration for users:**
- Drop `--rtk`, `--no-rtk`, `--no-project-rtk`, `--keep-rtk`,
`--context-tool`, `--no-context-tool` from any alias, script or CI job,
and unset `HEADROOM_RTK*` / `HEADROOM_CONTEXT_TOOL`. They now error
rather than being ignored, so the failure is immediate and
self-explaining.
- Previously-installed artifacts are purged automatically on the next
`wrap`/`unwrap`; no manual cleanup needed.
- `headroom perf --json` no longer carries a `cli_filtering` key, and
`/stats` no longer returns a `context_tool` section.
**Docs:** `docs/rtk-architecture.md` deleted; RTK/lean-ctx removed from
`README.md`,
`docs/content/docs/{configuration,opencode,grok-build,docker-install,filesystem-contract}.mdx`,
`docs/observability.md` and the matching `wiki/` pages.
`REALIGNMENT/09-phase-G-rtk-observability.md` is marked SUPERSEDED
rather than deleted, to keep the planning record.
**Follow-ups not in scope:** `_emit_wrap_interrupted` was deleted as
dead code — its only caller was the `except KeyboardInterrupt` guarding
the binary download, so with no download there is nothing slow left to
interrupt.
2026-07-30 22:59:41 -07:00
|
|
|
assert loader._state.contribution.to_dict()["tokens_saved"]["compression"] == 2
|
2026-05-08 10:59:17 -07:00
|
|
|
assert loader._state.contribution.to_dict()["tokens_saved"]["proxy_compression"] == 2
|
2026-04-23 07:39:52 -05:00
|
|
|
assert loader._state.contribution.compression_savings_usd == 1.5
|
|
|
|
|
assert loader._state.contribution.cache_savings_usd == 2.5
|
|
|
|
|
assert loader._state.poll_count == 7
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
persist_path.write_text("{invalid", encoding="utf-8")
|
|
|
|
|
broken = SubscriptionTracker(persist_path=persist_path)
|
|
|
|
|
assert broken._state.poll_count == 0
|
2026-04-24 15:33:30 +02:00
|
|
|
|
2026-04-23 07:39:52 -05:00
|
|
|
missing = SubscriptionTracker(persist_path=tmp_path / "missing.json")
|
|
|
|
|
assert missing._state.poll_count == 0
|
fix: remove rtk and lean-ctx CLI context tools (#2677)
## Description
Removes both third-party CLI context tools — **rtk** and **lean-ctx** —
and with them the context-tool selector itself. Headroom no longer
downloads, installs or configures either one, and there is no
replacement.
The previous pass (#2344) gated only three entry points inside
`headroom/cli/wrap.py`. That left the feature reachable in practice:
| Gap | Effect |
|---|---|
| `scripts/install.sh:1544`, `install.ps1:1681` | Ran `rtk init --global
--auto-patch` from bash/PowerShell, **bypassing the Python gate
entirely** — `curl \| sh` still wrote a Claude Code `PreToolUse` hook
regardless of `HEADROOM_RTK` |
| `wrap.py` `_setup_context_tool_for_agent` | **`wrap openhands` was
broken by default**: `rtk_required=True` met a gate returning `None` →
`SystemExit(1)`. Invisible because all 8 openhands tests patched
`_ensure_rtk_binary` to a fake path |
| `proxy/helpers.py`, `subscription/tracker.py` | Proxy shelled out to
`rtk gain` from `/stats`, the dashboard and `headroom perf`; the tracker
polled it per contribution (`_RTK_WIRING_DEFAULT = "enabled"`) |
| No cleanup path | Nothing removed artifacts an earlier default had
installed, so a machine that once ran the old default kept rtk in the
loop forever (#1669, #1955) |
Also worth noting: the rtk binary download had **no SHA or signature
verification** — only `rtk --version` as a smoke test.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [x] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [x] Code refactoring (no functional changes)
## Changes Made
**Removed** — `headroom/rtk/` and `headroom/lean_ctx/` packages,
`headroom/cli/wrap_rtk_metrics.py`, `_selected_context_tool` /
`_setup_context_tool_for_agent` / `_VALID_CONTEXT_TOOLS`, the `--rtk` /
`--no-rtk` / `--no-project-rtk` / `--keep-rtk` flags across all 18 wrap
subcommands, `HEADROOM_RTK*`, the proxy-side `rtk gain` polling, the
dashboard CLI-filtering panel (rows + all 8 `cliFiltering*` Alpine
getters), `paths.rtk_path()` / `lean_ctx_path()`, the SDK path helpers,
`benchmarks/rtk_loop_learn_eval.py`, and the `headroom/rtk/**` CI path
filters.
**Fails loudly, not silently** — `--context-tool` / `--no-context-tool`
/ `HEADROOM_CONTEXT_TOOL` are kept solely to error out. They live in
shell profiles, aliases and CI jobs, and accepting them as a no-op would
read as Headroom having quietly stopped working. The installers reject
them too, which matters more than it looks: their arg parsers forward
the first unknown flag **and everything after it** to the wrapped tool,
so a leftover `--no-rtk` would have silently swallowed a following
`--port` and then been ignored downstream.
**New `headroom/context_tool_cleanup.py`** — deleting the code cannot
help a machine that already ran the old default, since the hooks,
binaries and injected guidance are durable on disk.
`purge_context_tool_artifacts()` runs once per `wrap`/`unwrap` and
removes the registered hook entries, the generated hook scripts, the
Headroom-managed `~/.local/bin` symlinks, the vendored
`~/.headroom/bin/{rtk,lean-ctx}` binaries, the `lean-ctx` MCP server
entry and the marker-fenced instruction blocks. Deliberately
conservative: idempotent, **skips** a malformed config rather than
overwriting it, and only unlinks a symlink resolving inside Headroom's
own bin dir so a user's own build is untouched. It reports on
**stderr**, because `wrap/unwrap openclaw --prepare-only` emit
machine-readable JSON on stdout as their entire contract. Skipped for
`wrap selfheal` (runs from a SessionStart hook; must not race Claude
Code's writer for `~/.claude.json`) and for `--help`, which must stay
read-only.
**Client-config hardening** (discovered while investigating a "corrupted
Serena settings file" report) — `wrap.py` reset a settings file to `{}`
when an existing file would not parse, then wrote that back. One
hand-edited typo or a transient `EACCES`/`EINTR` on a valid file
destroyed the user's `permissions`, `env` and `hooks`, on **every
`headroom wrap claude`**. It now refuses to write. Separately,
`fsutil.write_text` is now atomic (temp file + `fsync` + `os.replace`),
fixing all 14 non-atomic client-config writes at once; it follows
symlinks rather than replacing them (dotfile managers) and preserves an
existing file's mode.
**Deliberately kept** — `rtk` stays in the wrapper-peel list in
`transforms/content_router.py`. It sits beside `sudo`/`env`/`timeout` as
shell-command grammar, so `rtk cat f` is still classified as a file read
for anyone running their own rtk install, which the purge intentionally
leaves alone.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ ruff check headroom/ tests/ e2e/ --exclude headroom/dashboard/templates
All checks passed!
$ ruff format --check headroom/ tests/ e2e/ --exclude headroom/dashboard/templates
1255 files already formatted
$ mypy headroom/
Success: no issues found in 508 source files
$ pytest tests/test_context_tool_cleanup.py -q
11 passed
$ pytest tests/test_fsutil.py -q
12 passed
$ pytest tests/test_cli/test_wrap_codex.py -q # 89 tests
89 passed in 431.68s
$ pytest tests/test_cli/test_wrap_opencode.py -q
39 passed in 257.46s
$ pytest tests/test_cli/test_wrap_helpers.py -q
45 passed
$ pytest tests/test_paths.py -q
75 passed
$ pytest tests/test_cli/test_unwrap_claude.py -q
14 passed
$ pytest tests/test_proxy_savings_history.py -q
39 passed
$ pytest tests/test_cli/test_wrap_copilot.py -q
27 passed
$ pytest tests/test_cli/test_wrap_zcode.py -q
20 passed
$ pytest tests/test_subscription_tracker.py -q
9 passed
$ pytest tests/test_proxy_dashboard_stats_cache.py -q
5 passed, 1 skipped
```
Repo-wide grep for 14 removed symbols (`headroom.rtk`,
`headroom.lean_ctx`, `_ensure_rtk_binary`, `_selected_context_tool`,
`_get_context_tool_stats`, `rtk_path`, `lean_ctx_path`,
`wrap_rtk_metrics`, `HEADROOM_RTK`, `cli_tokens_avoided`,
`tokens_saved_rtk`, …) across `*.py`, `*.ts`, `*.sh`, `*.ps1`, `*.yml`,
`*.html`: **zero hits**.
Notable test changes: `test_wrap_openhands.py` no longer patches
`_ensure_rtk_binary` and asserts `wrap openhands --prepare-only` exits 0
unpatched — the regression that was previously masked.
`test_wrap_continue.py` and `test_wrap_hintfile_agents.py` were removed
(every test drove RTK instruction injection). A new
`test_subscription_tracker.py::test_load_state_written_before_cli_context_tools_were_removed`
proves a pre-removal `subscription_state.json` still loads.
## Real Behavior Proof
- **Environment:** macOS 15.4 (darwin 25.4.0), Python 3.12.6, Headroom @
this branch, real `~/.headroom` and `~/.claude` on the dev machine.
- **Exact command / steps and observed result:**
```text
# 1. Retired flag fails loudly instead of silently no-op'ing
$ headroom wrap codex --prepare-only --context-tool rtk
Error: CLI context tools (rtk, lean-ctx) have been removed from Headroom: they
rewrote shell commands through a third-party binary Headroom no longer manages.
Drop --context-tool / --no-context-tool and unset HEADROOM_CONTEXT_TOOL;
`headroom wrap` uninstalls what they left behind on first run.
$ HEADROOM_CONTEXT_TOOL=lean-ctx headroom wrap codex --prepare-only
Error: CLI context tools (rtk, lean-ctx) have been removed from Headroom: ...
# 2. install.sh rejects the retired flags (extracted parse_wrap_args harness)
['--no-rtk', '--port', '9999'] rc=1 ERROR: CLI context tools ... Drop --no-rtk
['--context-tool=rtk'] rc=1 ERROR: CLI context tools ... Drop --context-tool
$ bash -n scripts/install.sh # syntax OK
# 3. Purge ran against the real machine, which had all the orphaned artifacts
$ python -c "from headroom.context_tool_cleanup import purge_context_tool_artifacts; ..."
removed ~/.headroom/bin/lean-ctx (51 MB)
removed ~/.headroom/bin/rtk (7.7 MB)
removed ~/.local/bin/rtk (symlink into ~/.headroom/bin)
removed ~/.claude/hooks/rtk-rewrite.sh
removed 8 lean-ctx-* hook scripts
# ~/.claude.json afterwards: 90 top-level keys, 19 projects, mcpServers unchanged
# → ~59 MB reclaimed, no unrelated key touched
# 4. stdout stays machine-readable while the purge reports (planted a fake artifact)
$ headroom wrap openclaw --prepare-only --gateway-provider-id codex >out 2>err
$ cat out
{"enabled":true,"config":{"proxyPort":8787,...}} # parses as JSON
$ cat err
Retired CLI context tool cleanup: removed /Users/tcms/.headroom/bin/rtk
# 5. --help is inert (planted artifact survives), a real run purges
$ headroom wrap codex --help → artifact survived: CORRECT
$ headroom wrap openclaw --prepare-only → purged: CORRECT
# 6. MCP purge dry-run against a copy of the real 82 KB ~/.claude.json
top-level keys 90 -> 90; projects 19 -> 19; LOST keys: none
all content outside mcpServers byte-identical: True
```
Dashboard rendered via the Playwright test after the panel removal:
"Token Savings" shows only `Proxy 0 (0.0%)` / `Of total wire: 36.86%`,
and "Token Usage" reads Before Compression → Proxy Removed → After
Compression with no "Filtered (this session)" row. Nothing below the
removed panel broke.
- **Not tested:** Windows and Linux (macOS only) — `install.ps1` is
verified by brace-balance and inspection, not executed, since no `pwsh`
is available locally. The wrap e2e suite (`e2e/wrap/run.py`) was updated
but not run; it needs the Docker e2e image. `serena project index`
interaction is exercised in the stacked base PR.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md` — it is generated by
release-please from my Conventional Commit PR title (a CI guard enforces
this)
## Additional Notes
**Stacked on #2676** (`tejas/serena-config-bootstrap`) — please merge
that first; this PR's base should then be retargeted to `main`, or it
will read as containing that fix too.
**Breaking-change migration for users:**
- Drop `--rtk`, `--no-rtk`, `--no-project-rtk`, `--keep-rtk`,
`--context-tool`, `--no-context-tool` from any alias, script or CI job,
and unset `HEADROOM_RTK*` / `HEADROOM_CONTEXT_TOOL`. They now error
rather than being ignored, so the failure is immediate and
self-explaining.
- Previously-installed artifacts are purged automatically on the next
`wrap`/`unwrap`; no manual cleanup needed.
- `headroom perf --json` no longer carries a `cli_filtering` key, and
`/stats` no longer returns a `context_tool` section.
**Docs:** `docs/rtk-architecture.md` deleted; RTK/lean-ctx removed from
`README.md`,
`docs/content/docs/{configuration,opencode,grok-build,docker-install,filesystem-contract}.mdx`,
`docs/observability.md` and the matching `wiki/` pages.
`REALIGNMENT/09-phase-G-rtk-observability.md` is marked SUPERSEDED
rather than deleted, to keep the planning record.
**Follow-ups not in scope:** `_emit_wrap_interrupted` was deleted as
dead code — its only caller was the `except KeyboardInterrupt` guarding
the binary download, so with no download there is nothing slow left to
interrupt.
2026-07-30 22:59:41 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_load_state_written_before_cli_context_tools_were_removed(tmp_path: Path) -> None:
|
|
|
|
|
"""A pre-removal state file must still load; the retired keys are ignored.
|
|
|
|
|
|
|
|
|
|
Older releases persisted ``cli_filtering`` / ``cli_filtering_raw`` / ``rtk``
|
|
|
|
|
/ ``rtk_raw`` counters, and wrote the dashboard-facing ``compression`` as
|
|
|
|
|
proxy-compression *plus* CLI filtering. Those files are still on users'
|
|
|
|
|
disks, so the loader must neither raise on the extra keys nor double-count
|
|
|
|
|
the inflated ``compression`` value — it prefers ``proxy_compression``.
|
|
|
|
|
"""
|
|
|
|
|
persist_path = tmp_path / "tracker-state.json"
|
|
|
|
|
persist_path.write_text(
|
|
|
|
|
json.dumps(
|
|
|
|
|
{
|
|
|
|
|
"poll_count": 7,
|
|
|
|
|
"contribution": {
|
|
|
|
|
"tokens_submitted": 11,
|
|
|
|
|
"tokens_saved": {
|
|
|
|
|
# 2 (proxy) + 9 (retired CLI layer) as older code wrote it.
|
|
|
|
|
"compression": 11,
|
|
|
|
|
"proxy_compression": 2,
|
|
|
|
|
"cli_filtering": 9,
|
|
|
|
|
"cli_filtering_raw": 3,
|
|
|
|
|
"rtk": 9,
|
|
|
|
|
"rtk_raw": 9,
|
|
|
|
|
"cache_reads": 4,
|
|
|
|
|
"total": 15,
|
|
|
|
|
},
|
|
|
|
|
"savings_usd": {"compression": 1.5, "cache": 2.5},
|
|
|
|
|
},
|
|
|
|
|
}
|
|
|
|
|
),
|
|
|
|
|
encoding="utf-8",
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
loader = SubscriptionTracker(persist_path=persist_path)
|
|
|
|
|
|
|
|
|
|
assert loader._state.poll_count == 7
|
|
|
|
|
assert loader._state.contribution.tokens_submitted == 11
|
|
|
|
|
# The raw proxy field wins, so the retired layer's tokens are not counted.
|
|
|
|
|
assert loader._state.contribution.tokens_saved_compression == 2
|
|
|
|
|
assert loader._state.contribution.tokens_saved_cache_reads == 4
|
|
|
|
|
assert loader._state.contribution.compression_savings_usd == 1.5
|
|
|
|
|
assert loader._state.contribution.cache_savings_usd == 2.5
|
|
|
|
|
# The retired keys are gone from what the tracker now emits.
|
|
|
|
|
emitted = loader._state.contribution.to_dict()["tokens_saved"]
|
|
|
|
|
assert "cli_filtering" not in emitted
|
|
|
|
|
assert "rtk" not in emitted
|