headroom/tests/test_verbosity_learn.py

Ignoring revisions in .git-blame-ignore-revs. Click here to bypass and see the normal blame view.

359 lines
12 KiB
Python
Raw Normal View History

feat: output-token reduction — verbosity shaper, per-user learning, counterfactual savings (#965) ## Description Adds the first levers that reduce the tokens the model **writes back** (output), complementing Headroom's existing input compression. Output costs 5× input on Opus-class models and is full of waste (ceremony, restated code, deep "thinking" on routine steps). Two phases in one self-contained PR off `main`: the request-side output shaper, then per-user verbosity learning plus an honest counterfactual savings estimator and dashboard surfacing. ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [x] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **Output shaper** (`output_shaper.py`, opt-in `HEADROOM_OUTPUT_SHAPER=1`): cache-safe verbosity steering appended to the system-prompt tail (5 levels); effort routing that lowers `output_config.effort` on mechanical tool-result continuations; legacy `thinking.budget_tokens` clamp. Never injects effort where absent, never toggles `thinking.type`. - **`headroom learn --verbosity`**: mines Claude Code transcripts for behavioral signals (interrupts, length-adaptive fast-skips, echo ratio), recommends a verbosity level (heuristic + optional `--llm-judge`), and seeds the savings baseline. - **Counterfactual estimator** (`output_savings.py`): per-stratum synthetic-control (estimated) + A/B holdout (measured) with a propagated 95% CI; conversation-stable arm assignment for A/B validity and prefix-cache safety. - **AIMD verbosity controller** (`verbosity_controller.py`): additive-increase / fast-back-off state machine; live signal emission gated off by default. - **Wiring + surfaces**: shaper resolves the learned level; recording rides the existing `transforms_applied` channel through the outcome funnel (no `RequestOutcome` changes); `headroom output-savings` CLI; dashboard "Output Tokens Saved" card. - **Docs**: simple-words user guide + design doc with the counterfactual methodology. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ pytest tests/test_output_savings.py tests/test_output_savings_cli.py \ tests/test_verbosity_learn.py tests/test_verbosity_controller.py \ tests/test_output_shaper.py -q 94 passed in 0.54s $ pytest tests/test_request_outcome.py tests/test_handler_outcome_tag_invariant.py \ tests/test_proxy_dashboard_stats_cache.py -q 44 passed $ ruff format --check . 831 files already formatted $ mypy headroom --ignore-missing-imports Success: no issues found in 361 source files ``` ## Real Behavior Proof - Environment: macOS, Python 3.12 (`.venv`), `anthropic` 0.76, live API model `claude-opus-4-8`. - Exact command / steps: `HEADROOM_OUTPUT_SHAPER=1`; `headroom learn --verbosity --apply` (seeds level + baseline); `python scripts/eval_output_shaper.py A` (live before/after); simulate holdout traffic then `headroom output-savings`. - Observed result: code-review ask — baseline 1,750 output tokens → L2 1,354 (−22.7%) → L3 599 (−65.8%), same bugs found. `learn --verbosity` on 24 real sessions → 11% interrupt / 26% fast-skip → L3 (high confidence). Measured A/B path → 31.7% reduction (95% CI 27.7%–35.7%). 94 new tests + 44 existing outcome/dashboard tests green; ruff + mypy clean. - Not tested: live streaming-path recording exercised only via unit tests (the `transforms_applied` funnel is shared across paths); runtime AIMD signal emission is gated off by default and not exercised live. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable ## Additional Notes Output savings are counterfactual (we never observe what the model *would* have written), so the estimator separates **estimated** (vs a learned baseline) from **measured** (A/B holdout via `HEADROOM_OUTPUT_HOLDOUT`) and always reports a confidence band — never a single made-up number. CHANGELOG left unchecked (release-please manages it). Runtime AIMD self-tuning is intentionally a TODO (controller built/tested; live signal emission gated behind `HEADROOM_VERBOSITY_AUTOTUNE`).
2026-06-16 21:06:43 -07:00
"""Tests for headroom.learn.verbosity — behavioral signal extraction."""
from __future__ import annotations
import json
from pathlib import Path
from headroom.learn.verbosity import (
fix(learn): handle Windows UTF-8, drive-letter paths, and CLI shim fallback (#1895) ## Description `headroom learn --verbosity` is broken on Windows in three related ways: - Transcript/profile reads can use the platform default codec, so non-ASCII content can raise `UnicodeDecodeError` and collapse learning signals to empty output. - `--project <path>` can miss real Claude project directories because Windows profile junctions can raise `PermissionError` during directory walks, and escaped Claude project folder names cannot always distinguish `vibe-remote` from `vibe\remote`. - `headroom learn --agent codex` can fail with `` `claude` not found in PATH `` even when the npm-installed CLI exists, because Windows `.cmd` shims require `PATHEXT` resolution. Refs https://github.com/headroomlabs-ai/headroom/issues/1624 for the Windows learn failures. The dashboard-hint UX and third-party-provider-auth items in that issue are unrelated and out of scope for this PR. ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `headroom/learn/verbosity.py`: read and write verbosity transcripts/profiles with `encoding="utf-8"` so non-ASCII content works regardless of the Windows locale codec. - `headroom/learn/plugins/claude.py`: skip inaccessible siblings one entry at a time during greedy project path decoding, so one Windows junction no longer hides valid project directories. - `headroom/learn/plugins/claude.py`: prefer a valid `cwd` found in Claude session JSONL when discovering project paths, which resolves ambiguous escaped folder names such as `vibe-remote` versus `vibe\remote`. - `headroom/learn/analyzer.py`: resolve Windows CLI shim paths through `shutil.which()` after `FileNotFoundError`, then retry once for streaming and non-streaming CLI calls. - `CHANGELOG.md`: document the Windows learn fixes under `Unreleased`. ## Testing - [x] Unit tests pass (`uv run pytest tests/test_verbosity_learn.py::TestWindowsEncoding tests/test_learn/test_scanner.py::TestGreedyPathDecode::test_permission_denied_sibling_does_not_abort_the_walk tests/test_learn/test_analyzer.py::TestWindowsCliShimFallback tests/test_learn/test_scanner.py::TestDecodeProjectPath::test_discover_project_prefers_session_cwd_over_ambiguous_folder_name -q`) - [x] Linting passes (`uv run ruff check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`) - [x] Formatting passes (`uv run ruff format --check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`) - [ ] Type checking passes (`uv run mypy headroom`) not run; no new public type surface - [x] New tests added for the Windows `cwd` disambiguation regression - [x] Manual testing performed ### Test Output ```text uv run ruff format headroom/learn/plugins/claude.py 1 file reformatted uv run ruff check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py All checks passed! uv run ruff format --check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py 2 files already formatted uv run pytest tests/test_verbosity_learn.py::TestWindowsEncoding tests/test_learn/test_scanner.py::TestGreedyPathDecode::test_permission_denied_sibling_does_not_abort_the_walk tests/test_learn/test_analyzer.py::TestWindowsCliShimFallback tests/test_learn/test_scanner.py::TestDecodeProjectPath::test_discover_project_prefers_session_cwd_over_ambiguous_folder_name -q 9 passed in 0.25s ``` CI on current head `c6dbac40` is green. A prior `test (1)` run hit an unrelated timing-sensitive scheduler assertion; GitHub did not permit direct rerun without admin rights, so the empty commit `c6dbac40` retriggered CI and the shard passed. ## Real Behavior Proof - Environment: Windows 11, Python 3.12, real filesystem for the path-decoding reproduction. - Exact command / steps: Ran `uv run pytest tests/test_verbosity_learn.py::TestWindowsEncoding tests/test_learn/test_scanner.py::TestGreedyPathDecode::test_permission_denied_sibling_does_not_abort_the_walk tests/test_learn/test_analyzer.py::TestWindowsCliShimFallback tests/test_learn/test_scanner.py::TestDecodeProjectPath::test_discover_project_prefers_session_cwd_over_ambiguous_folder_name -q`, `uv run ruff check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`, and `uv run ruff format --check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`; the path tests create real Windows-style project directories, inaccessible siblings, ambiguous `vibe\remote` versus `vibe-remote` folders, and Claude session JSONL with `cwd` pointing at the intended project. - Observed result: The focused test command returned `9 passed in 0.25s`, Ruff check passed, and Ruff format check passed. The decoder skips the inaccessible sibling and reaches `vibe-remote`; the session-`cwd` test returns `vibe-remote` instead of trusting the ambiguous escaped folder name; the UTF-8 tests round-trip non-ASCII transcript/profile content under a non-UTF-8 Windows-style codec; the CLI shim tests retry once through `shutil.which()` after `FileNotFoundError`. - Not tested: real npm-installed `claude`/`codex` CLI shims, dashboard-hint UX, and third-party-provider auth. ## Review Readiness - [x] I have performed a self-review - [x] Retrospective review completed after opening; it found one missing `cwd` disambiguation case, now fixed in this PR - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Additional Notes - The post-open retrospective review concluded this needed targeted rework rather than only a retrospective sign-off. - The `cwd` recovery commit is `ee720bb1`; `0905be7b` contains the required formatter cleanup; current head `c6dbac40` is an empty CI-rerun commit after GitHub denied direct rerun without admin rights. --------- Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-09 13:49:38 -04:00
VerbosityProfile,
feat: output-token reduction — verbosity shaper, per-user learning, counterfactual savings (#965) ## Description Adds the first levers that reduce the tokens the model **writes back** (output), complementing Headroom's existing input compression. Output costs 5× input on Opus-class models and is full of waste (ceremony, restated code, deep "thinking" on routine steps). Two phases in one self-contained PR off `main`: the request-side output shaper, then per-user verbosity learning plus an honest counterfactual savings estimator and dashboard surfacing. ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [x] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **Output shaper** (`output_shaper.py`, opt-in `HEADROOM_OUTPUT_SHAPER=1`): cache-safe verbosity steering appended to the system-prompt tail (5 levels); effort routing that lowers `output_config.effort` on mechanical tool-result continuations; legacy `thinking.budget_tokens` clamp. Never injects effort where absent, never toggles `thinking.type`. - **`headroom learn --verbosity`**: mines Claude Code transcripts for behavioral signals (interrupts, length-adaptive fast-skips, echo ratio), recommends a verbosity level (heuristic + optional `--llm-judge`), and seeds the savings baseline. - **Counterfactual estimator** (`output_savings.py`): per-stratum synthetic-control (estimated) + A/B holdout (measured) with a propagated 95% CI; conversation-stable arm assignment for A/B validity and prefix-cache safety. - **AIMD verbosity controller** (`verbosity_controller.py`): additive-increase / fast-back-off state machine; live signal emission gated off by default. - **Wiring + surfaces**: shaper resolves the learned level; recording rides the existing `transforms_applied` channel through the outcome funnel (no `RequestOutcome` changes); `headroom output-savings` CLI; dashboard "Output Tokens Saved" card. - **Docs**: simple-words user guide + design doc with the counterfactual methodology. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ pytest tests/test_output_savings.py tests/test_output_savings_cli.py \ tests/test_verbosity_learn.py tests/test_verbosity_controller.py \ tests/test_output_shaper.py -q 94 passed in 0.54s $ pytest tests/test_request_outcome.py tests/test_handler_outcome_tag_invariant.py \ tests/test_proxy_dashboard_stats_cache.py -q 44 passed $ ruff format --check . 831 files already formatted $ mypy headroom --ignore-missing-imports Success: no issues found in 361 source files ``` ## Real Behavior Proof - Environment: macOS, Python 3.12 (`.venv`), `anthropic` 0.76, live API model `claude-opus-4-8`. - Exact command / steps: `HEADROOM_OUTPUT_SHAPER=1`; `headroom learn --verbosity --apply` (seeds level + baseline); `python scripts/eval_output_shaper.py A` (live before/after); simulate holdout traffic then `headroom output-savings`. - Observed result: code-review ask — baseline 1,750 output tokens → L2 1,354 (−22.7%) → L3 599 (−65.8%), same bugs found. `learn --verbosity` on 24 real sessions → 11% interrupt / 26% fast-skip → L3 (high confidence). Measured A/B path → 31.7% reduction (95% CI 27.7%–35.7%). 94 new tests + 44 existing outcome/dashboard tests green; ruff + mypy clean. - Not tested: live streaming-path recording exercised only via unit tests (the `transforms_applied` funnel is shared across paths); runtime AIMD signal emission is gated off by default and not exercised live. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable ## Additional Notes Output savings are counterfactual (we never observe what the model *would* have written), so the estimator separates **estimated** (vs a learned baseline) from **measured** (A/B holdout via `HEADROOM_OUTPUT_HOLDOUT`) and always reports a confidence band — never a single made-up number. CHANGELOG left unchecked (release-please manages it). Runtime AIMD self-tuning is intentionally a TODO (controller built/tested; live signal emission gated behind `HEADROOM_VERBOSITY_AUTOTUNE`).
2026-06-16 21:06:43 -07:00
VerbositySignals,
fix(learn): handle Windows UTF-8, drive-letter paths, and CLI shim fallback (#1895) ## Description `headroom learn --verbosity` is broken on Windows in three related ways: - Transcript/profile reads can use the platform default codec, so non-ASCII content can raise `UnicodeDecodeError` and collapse learning signals to empty output. - `--project <path>` can miss real Claude project directories because Windows profile junctions can raise `PermissionError` during directory walks, and escaped Claude project folder names cannot always distinguish `vibe-remote` from `vibe\remote`. - `headroom learn --agent codex` can fail with `` `claude` not found in PATH `` even when the npm-installed CLI exists, because Windows `.cmd` shims require `PATHEXT` resolution. Refs https://github.com/headroomlabs-ai/headroom/issues/1624 for the Windows learn failures. The dashboard-hint UX and third-party-provider-auth items in that issue are unrelated and out of scope for this PR. ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `headroom/learn/verbosity.py`: read and write verbosity transcripts/profiles with `encoding="utf-8"` so non-ASCII content works regardless of the Windows locale codec. - `headroom/learn/plugins/claude.py`: skip inaccessible siblings one entry at a time during greedy project path decoding, so one Windows junction no longer hides valid project directories. - `headroom/learn/plugins/claude.py`: prefer a valid `cwd` found in Claude session JSONL when discovering project paths, which resolves ambiguous escaped folder names such as `vibe-remote` versus `vibe\remote`. - `headroom/learn/analyzer.py`: resolve Windows CLI shim paths through `shutil.which()` after `FileNotFoundError`, then retry once for streaming and non-streaming CLI calls. - `CHANGELOG.md`: document the Windows learn fixes under `Unreleased`. ## Testing - [x] Unit tests pass (`uv run pytest tests/test_verbosity_learn.py::TestWindowsEncoding tests/test_learn/test_scanner.py::TestGreedyPathDecode::test_permission_denied_sibling_does_not_abort_the_walk tests/test_learn/test_analyzer.py::TestWindowsCliShimFallback tests/test_learn/test_scanner.py::TestDecodeProjectPath::test_discover_project_prefers_session_cwd_over_ambiguous_folder_name -q`) - [x] Linting passes (`uv run ruff check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`) - [x] Formatting passes (`uv run ruff format --check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`) - [ ] Type checking passes (`uv run mypy headroom`) not run; no new public type surface - [x] New tests added for the Windows `cwd` disambiguation regression - [x] Manual testing performed ### Test Output ```text uv run ruff format headroom/learn/plugins/claude.py 1 file reformatted uv run ruff check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py All checks passed! uv run ruff format --check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py 2 files already formatted uv run pytest tests/test_verbosity_learn.py::TestWindowsEncoding tests/test_learn/test_scanner.py::TestGreedyPathDecode::test_permission_denied_sibling_does_not_abort_the_walk tests/test_learn/test_analyzer.py::TestWindowsCliShimFallback tests/test_learn/test_scanner.py::TestDecodeProjectPath::test_discover_project_prefers_session_cwd_over_ambiguous_folder_name -q 9 passed in 0.25s ``` CI on current head `c6dbac40` is green. A prior `test (1)` run hit an unrelated timing-sensitive scheduler assertion; GitHub did not permit direct rerun without admin rights, so the empty commit `c6dbac40` retriggered CI and the shard passed. ## Real Behavior Proof - Environment: Windows 11, Python 3.12, real filesystem for the path-decoding reproduction. - Exact command / steps: Ran `uv run pytest tests/test_verbosity_learn.py::TestWindowsEncoding tests/test_learn/test_scanner.py::TestGreedyPathDecode::test_permission_denied_sibling_does_not_abort_the_walk tests/test_learn/test_analyzer.py::TestWindowsCliShimFallback tests/test_learn/test_scanner.py::TestDecodeProjectPath::test_discover_project_prefers_session_cwd_over_ambiguous_folder_name -q`, `uv run ruff check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`, and `uv run ruff format --check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`; the path tests create real Windows-style project directories, inaccessible siblings, ambiguous `vibe\remote` versus `vibe-remote` folders, and Claude session JSONL with `cwd` pointing at the intended project. - Observed result: The focused test command returned `9 passed in 0.25s`, Ruff check passed, and Ruff format check passed. The decoder skips the inaccessible sibling and reaches `vibe-remote`; the session-`cwd` test returns `vibe-remote` instead of trusting the ambiguous escaped folder name; the UTF-8 tests round-trip non-ASCII transcript/profile content under a non-UTF-8 Windows-style codec; the CLI shim tests retry once through `shutil.which()` after `FileNotFoundError`. - Not tested: real npm-installed `claude`/`codex` CLI shims, dashboard-hint UX, and third-party-provider auth. ## Review Readiness - [x] I have performed a self-review - [x] Retrospective review completed after opening; it found one missing `cwd` disambiguation case, now fixed in this PR - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Additional Notes - The post-open retrospective review concluded this needed targeted rework rather than only a retrospective sign-off. - The `cwd` recovery commit is `ee720bb1`; `0905be7b` contains the required formatter cleanup; current head `c6dbac40` is an empty CI-rerun commit after GitHub denied direct rerun without admin rights. --------- Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-09 13:49:38 -04:00
_parse_session,
feat: output-token reduction — verbosity shaper, per-user learning, counterfactual savings (#965) ## Description Adds the first levers that reduce the tokens the model **writes back** (output), complementing Headroom's existing input compression. Output costs 5× input on Opus-class models and is full of waste (ceremony, restated code, deep "thinking" on routine steps). Two phases in one self-contained PR off `main`: the request-side output shaper, then per-user verbosity learning plus an honest counterfactual savings estimator and dashboard surfacing. ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [x] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **Output shaper** (`output_shaper.py`, opt-in `HEADROOM_OUTPUT_SHAPER=1`): cache-safe verbosity steering appended to the system-prompt tail (5 levels); effort routing that lowers `output_config.effort` on mechanical tool-result continuations; legacy `thinking.budget_tokens` clamp. Never injects effort where absent, never toggles `thinking.type`. - **`headroom learn --verbosity`**: mines Claude Code transcripts for behavioral signals (interrupts, length-adaptive fast-skips, echo ratio), recommends a verbosity level (heuristic + optional `--llm-judge`), and seeds the savings baseline. - **Counterfactual estimator** (`output_savings.py`): per-stratum synthetic-control (estimated) + A/B holdout (measured) with a propagated 95% CI; conversation-stable arm assignment for A/B validity and prefix-cache safety. - **AIMD verbosity controller** (`verbosity_controller.py`): additive-increase / fast-back-off state machine; live signal emission gated off by default. - **Wiring + surfaces**: shaper resolves the learned level; recording rides the existing `transforms_applied` channel through the outcome funnel (no `RequestOutcome` changes); `headroom output-savings` CLI; dashboard "Output Tokens Saved" card. - **Docs**: simple-words user guide + design doc with the counterfactual methodology. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ pytest tests/test_output_savings.py tests/test_output_savings_cli.py \ tests/test_verbosity_learn.py tests/test_verbosity_controller.py \ tests/test_output_shaper.py -q 94 passed in 0.54s $ pytest tests/test_request_outcome.py tests/test_handler_outcome_tag_invariant.py \ tests/test_proxy_dashboard_stats_cache.py -q 44 passed $ ruff format --check . 831 files already formatted $ mypy headroom --ignore-missing-imports Success: no issues found in 361 source files ``` ## Real Behavior Proof - Environment: macOS, Python 3.12 (`.venv`), `anthropic` 0.76, live API model `claude-opus-4-8`. - Exact command / steps: `HEADROOM_OUTPUT_SHAPER=1`; `headroom learn --verbosity --apply` (seeds level + baseline); `python scripts/eval_output_shaper.py A` (live before/after); simulate holdout traffic then `headroom output-savings`. - Observed result: code-review ask — baseline 1,750 output tokens → L2 1,354 (−22.7%) → L3 599 (−65.8%), same bugs found. `learn --verbosity` on 24 real sessions → 11% interrupt / 26% fast-skip → L3 (high confidence). Measured A/B path → 31.7% reduction (95% CI 27.7%–35.7%). 94 new tests + 44 existing outcome/dashboard tests green; ruff + mypy clean. - Not tested: live streaming-path recording exercised only via unit tests (the `transforms_applied` funnel is shared across paths); runtime AIMD signal emission is gated off by default and not exercised live. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable ## Additional Notes Output savings are counterfactual (we never observe what the model *would* have written), so the estimator separates **estimated** (vs a learned baseline) from **measured** (A/B holdout via `HEADROOM_OUTPUT_HOLDOUT`) and always reports a confidence band — never a single made-up number. CHANGELOG left unchecked (release-please manages it). Runtime AIMD self-tuning is intentionally a TODO (controller built/tested; live signal emission gated behind `HEADROOM_VERBOSITY_AUTOTUNE`).
2026-06-16 21:06:43 -07:00
analyze,
extract_signals,
recommend_level,
)
def _write_session(tmp_path: Path, name: str, lines: list[dict]) -> Path:
p = tmp_path / f"{name}.jsonl"
p.write_text("\n".join(json.dumps(line) for line in lines))
return p
def _assistant(
text: str,
*,
ts: str,
out_tokens: int = 100,
model: str = "claude-opus-4-8",
in_tokens: int = 5000,
) -> dict:
return {
"type": "assistant",
"timestamp": ts,
"message": {
"model": model,
"content": [{"type": "text", "text": text}],
"usage": {"input_tokens": in_tokens, "output_tokens": out_tokens},
},
}
def _user(text: str, *, ts: str) -> dict:
return {"type": "user", "timestamp": ts, "message": {"role": "user", "content": text}}
def _tool_result(*, ts: str, content: str = "ok") -> dict:
return {
"type": "user",
"timestamp": ts,
"message": {
"role": "user",
"content": [{"type": "tool_result", "tool_use_id": "t1", "content": content}],
},
}
fix(learn): don't desync verbosity pairing on empty assistant turns (#2123) ## Description `verbosity._ordered_events` and `_parse_session` disagree about empty assistant turns, which desyncs the response list and produces spurious fast-skips. `_parse_session` only creates a `_Response` when an assistant message actually said something: ```python if words > 0 or out_tok > 0: responses.append(_Response(...)) ``` But `_ordered_events` consumes one `responses[ri]` for **every** assistant line, with no matching filter: ```python if ltype == "assistant" and ri < len(responses): out.append((responses[ri].ts, "assistant", responses[ri])) ri += 1 ``` So an assistant turn with no text and no output tokens — for example a pure `tool_use` turn where `usage` is absent — creates no `_Response` at parse time, yet still consumes a slot in `_ordered_events`. That slot actually belongs to a *later* real response, so the two lists drift by one. A human reply that follows the real answer is then paired with the next answer's (future) timestamp, `ts - last_resp.ts` goes negative, and since a negative gap is always below the read-fraction threshold, a spurious `fast_skip` is recorded. That inflates `fast_skip_rate`, which feeds `pressure`, which lowers the recommended verbosity level. The user side of `_ordered_events` already replicates its parse-site filter (`_human_text(...) is None -> continue`); only the assistant side was missing the equivalent guard. That asymmetry is the bug. ## Fix In `_ordered_events`, compute `words`/`out_tok` for the assistant line the same way `_parse_session` does and only consume a response when `words > 0 or out_tok > 0`, keeping the two functions in lockstep. Closes # ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `headroom/learn/verbosity.py`: `_ordered_events` applies the `words > 0 or out_tok > 0` guard on the assistant branch before consuming a response, with a comment explaining the desync. - `tests/test_verbosity_learn.py`: add `_empty_assistant` helper and `test_empty_assistant_message_does_not_desync_fast_skip` (an empty assistant turn before a real answer + a slow reply must not record a fast skip). - `CHANGELOG.md`: Bug Fixes entry. ## Testing - [x] Unit tests pass (`uv run --extra dev pytest tests/test_verbosity_learn.py::TestSignalExtraction::test_empty_assistant_message_does_not_desync_fast_skip -q`) - [x] Linting passes (`uvx ruff@0.15.17 check headroom/learn/verbosity.py tests/test_verbosity_learn.py headroom/memory/factory.py`) - [x] Type checking passes (`uvx mypy==1.20.2 headroom/memory/factory.py`) - [x] New tests added for new functionality - [ ] Manual testing performed ### Test Output ```text $ uvx ruff@0.15.17 check headroom/learn/verbosity.py tests/test_verbosity_learn.py All checks passed! $ python -m py_compile headroom/learn/verbosity.py tests/test_verbosity_learn.py OK ``` ## Real Behavior Proof - Environment: Windows 11, Python 3.12, `uvx ruff@0.15.17`. Importing `headroom` pulls in the torch/transformers stack and a full `pytest` gets OOM-killed on this box, so I verified the alignment with a dependency-free script that models the parse-site filter, the old vs new `_ordered_events` consume, and the resulting human-to-response pairing, and left the full pytest to CI. - Exact command / steps: built an event stream `[empty assistant, real answer #1, fast human reply, real answer #2, reply]`, computed the response list from the parse filter, then walked the old (unfiltered) and new (filtered) consume to find the gap between the first human and the response paired before it. - Observed result: old consume pairs the reply with answer #2 (a future timestamp) -> gap `-8` (spurious fast_skip); new consume keeps alignment and pairs it with answer #1 -> gap `+1`. The new test builds a session with an empty assistant turn and a genuinely slow reply and asserts `fast_skips == 0`. - Not tested: a real Claude Code transcript end to end; full local `pytest` deferred to CI (OOM, per above). ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [ ] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Additional Notes Merged current `main` to pick up the repository-wide mypy cache-key annotation fix, then verified the focused regression locally. the change adds the existing parse-site filter to one branch of a pure file-parsing function, verified by the standalone alignment proof and the new regression test for CI. Co-authored-by: JerrettDavis <mxjerrett@gmail.com> Co-authored-by: Tejas Chopra <chopratejas@gmail.com>
2026-07-14 21:34:32 +05:30
def _empty_assistant(*, ts: str) -> dict:
"""A pure tool_use assistant turn with no text and no output tokens.
`_parse_session` creates no `_Response` for it (words == 0 and out_tok == 0).
"""
return {
"type": "assistant",
"timestamp": ts,
"message": {
"model": "claude-opus-4-8",
"content": [{"type": "tool_use", "id": "t1", "name": "Bash", "input": {}}],
"usage": {"input_tokens": 100, "output_tokens": 0},
},
}
feat: output-token reduction — verbosity shaper, per-user learning, counterfactual savings (#965) ## Description Adds the first levers that reduce the tokens the model **writes back** (output), complementing Headroom's existing input compression. Output costs 5× input on Opus-class models and is full of waste (ceremony, restated code, deep "thinking" on routine steps). Two phases in one self-contained PR off `main`: the request-side output shaper, then per-user verbosity learning plus an honest counterfactual savings estimator and dashboard surfacing. ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [x] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **Output shaper** (`output_shaper.py`, opt-in `HEADROOM_OUTPUT_SHAPER=1`): cache-safe verbosity steering appended to the system-prompt tail (5 levels); effort routing that lowers `output_config.effort` on mechanical tool-result continuations; legacy `thinking.budget_tokens` clamp. Never injects effort where absent, never toggles `thinking.type`. - **`headroom learn --verbosity`**: mines Claude Code transcripts for behavioral signals (interrupts, length-adaptive fast-skips, echo ratio), recommends a verbosity level (heuristic + optional `--llm-judge`), and seeds the savings baseline. - **Counterfactual estimator** (`output_savings.py`): per-stratum synthetic-control (estimated) + A/B holdout (measured) with a propagated 95% CI; conversation-stable arm assignment for A/B validity and prefix-cache safety. - **AIMD verbosity controller** (`verbosity_controller.py`): additive-increase / fast-back-off state machine; live signal emission gated off by default. - **Wiring + surfaces**: shaper resolves the learned level; recording rides the existing `transforms_applied` channel through the outcome funnel (no `RequestOutcome` changes); `headroom output-savings` CLI; dashboard "Output Tokens Saved" card. - **Docs**: simple-words user guide + design doc with the counterfactual methodology. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ pytest tests/test_output_savings.py tests/test_output_savings_cli.py \ tests/test_verbosity_learn.py tests/test_verbosity_controller.py \ tests/test_output_shaper.py -q 94 passed in 0.54s $ pytest tests/test_request_outcome.py tests/test_handler_outcome_tag_invariant.py \ tests/test_proxy_dashboard_stats_cache.py -q 44 passed $ ruff format --check . 831 files already formatted $ mypy headroom --ignore-missing-imports Success: no issues found in 361 source files ``` ## Real Behavior Proof - Environment: macOS, Python 3.12 (`.venv`), `anthropic` 0.76, live API model `claude-opus-4-8`. - Exact command / steps: `HEADROOM_OUTPUT_SHAPER=1`; `headroom learn --verbosity --apply` (seeds level + baseline); `python scripts/eval_output_shaper.py A` (live before/after); simulate holdout traffic then `headroom output-savings`. - Observed result: code-review ask — baseline 1,750 output tokens → L2 1,354 (−22.7%) → L3 599 (−65.8%), same bugs found. `learn --verbosity` on 24 real sessions → 11% interrupt / 26% fast-skip → L3 (high confidence). Measured A/B path → 31.7% reduction (95% CI 27.7%–35.7%). 94 new tests + 44 existing outcome/dashboard tests green; ruff + mypy clean. - Not tested: live streaming-path recording exercised only via unit tests (the `transforms_applied` funnel is shared across paths); runtime AIMD signal emission is gated off by default and not exercised live. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable ## Additional Notes Output savings are counterfactual (we never observe what the model *would* have written), so the estimator separates **estimated** (vs a learned baseline) from **measured** (A/B holdout via `HEADROOM_OUTPUT_HOLDOUT`) and always reports a confidence band — never a single made-up number. CHANGELOG left unchecked (release-please manages it). Runtime AIMD self-tuning is intentionally a TODO (controller built/tested; live signal emission gated behind `HEADROOM_VERBOSITY_AUTOTUNE`).
2026-06-16 21:06:43 -07:00
LONG = " ".join(["word"] * 400) # well above the long-output floor
class TestSignalExtraction:
def test_interrupt_counted(self, tmp_path):
p = _write_session(
tmp_path,
"s",
[
_user("do a thing", ts="2026-01-01T00:00:00Z"),
_assistant(LONG, ts="2026-01-01T00:00:10Z"),
_user("[Request interrupted by user]", ts="2026-01-01T00:00:12Z"),
],
)
sig, _ = extract_signals([p])
assert sig.interrupts == 1
assert sig.human_msgs == 1 # the initial ask
def test_fast_skip_detected_length_adaptive(self, tmp_path):
# 400-word answer needs ~96s to read; reply after 5s = fast skip.
p = _write_session(
tmp_path,
"s",
[
_user("explain", ts="2026-01-01T00:00:00Z"),
_assistant(LONG, ts="2026-01-01T00:00:00Z"),
_user("ok next", ts="2026-01-01T00:00:05Z"),
],
)
sig, _ = extract_signals([p])
assert sig.skip_eligible == 1
assert sig.fast_skips == 1
def test_slow_reply_is_not_a_skip(self, tmp_path):
# Reply 120s after a 400-word answer (>read time) = read, not skipped.
p = _write_session(
tmp_path,
"s",
[
_user("explain", ts="2026-01-01T00:00:00Z"),
_assistant(LONG, ts="2026-01-01T00:00:00Z"),
_user("ok next", ts="2026-01-01T00:02:00Z"),
],
)
sig, _ = extract_signals([p])
assert sig.skip_eligible == 1
assert sig.fast_skips == 0
fix(learn): don't desync verbosity pairing on empty assistant turns (#2123) ## Description `verbosity._ordered_events` and `_parse_session` disagree about empty assistant turns, which desyncs the response list and produces spurious fast-skips. `_parse_session` only creates a `_Response` when an assistant message actually said something: ```python if words > 0 or out_tok > 0: responses.append(_Response(...)) ``` But `_ordered_events` consumes one `responses[ri]` for **every** assistant line, with no matching filter: ```python if ltype == "assistant" and ri < len(responses): out.append((responses[ri].ts, "assistant", responses[ri])) ri += 1 ``` So an assistant turn with no text and no output tokens — for example a pure `tool_use` turn where `usage` is absent — creates no `_Response` at parse time, yet still consumes a slot in `_ordered_events`. That slot actually belongs to a *later* real response, so the two lists drift by one. A human reply that follows the real answer is then paired with the next answer's (future) timestamp, `ts - last_resp.ts` goes negative, and since a negative gap is always below the read-fraction threshold, a spurious `fast_skip` is recorded. That inflates `fast_skip_rate`, which feeds `pressure`, which lowers the recommended verbosity level. The user side of `_ordered_events` already replicates its parse-site filter (`_human_text(...) is None -> continue`); only the assistant side was missing the equivalent guard. That asymmetry is the bug. ## Fix In `_ordered_events`, compute `words`/`out_tok` for the assistant line the same way `_parse_session` does and only consume a response when `words > 0 or out_tok > 0`, keeping the two functions in lockstep. Closes # ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `headroom/learn/verbosity.py`: `_ordered_events` applies the `words > 0 or out_tok > 0` guard on the assistant branch before consuming a response, with a comment explaining the desync. - `tests/test_verbosity_learn.py`: add `_empty_assistant` helper and `test_empty_assistant_message_does_not_desync_fast_skip` (an empty assistant turn before a real answer + a slow reply must not record a fast skip). - `CHANGELOG.md`: Bug Fixes entry. ## Testing - [x] Unit tests pass (`uv run --extra dev pytest tests/test_verbosity_learn.py::TestSignalExtraction::test_empty_assistant_message_does_not_desync_fast_skip -q`) - [x] Linting passes (`uvx ruff@0.15.17 check headroom/learn/verbosity.py tests/test_verbosity_learn.py headroom/memory/factory.py`) - [x] Type checking passes (`uvx mypy==1.20.2 headroom/memory/factory.py`) - [x] New tests added for new functionality - [ ] Manual testing performed ### Test Output ```text $ uvx ruff@0.15.17 check headroom/learn/verbosity.py tests/test_verbosity_learn.py All checks passed! $ python -m py_compile headroom/learn/verbosity.py tests/test_verbosity_learn.py OK ``` ## Real Behavior Proof - Environment: Windows 11, Python 3.12, `uvx ruff@0.15.17`. Importing `headroom` pulls in the torch/transformers stack and a full `pytest` gets OOM-killed on this box, so I verified the alignment with a dependency-free script that models the parse-site filter, the old vs new `_ordered_events` consume, and the resulting human-to-response pairing, and left the full pytest to CI. - Exact command / steps: built an event stream `[empty assistant, real answer #1, fast human reply, real answer #2, reply]`, computed the response list from the parse filter, then walked the old (unfiltered) and new (filtered) consume to find the gap between the first human and the response paired before it. - Observed result: old consume pairs the reply with answer #2 (a future timestamp) -> gap `-8` (spurious fast_skip); new consume keeps alignment and pairs it with answer #1 -> gap `+1`. The new test builds a session with an empty assistant turn and a genuinely slow reply and asserts `fast_skips == 0`. - Not tested: a real Claude Code transcript end to end; full local `pytest` deferred to CI (OOM, per above). ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [ ] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Additional Notes Merged current `main` to pick up the repository-wide mypy cache-key annotation fix, then verified the focused regression locally. the change adds the existing parse-site filter to one branch of a pure file-parsing function, verified by the standalone alignment proof and the new regression test for CI. Co-authored-by: JerrettDavis <mxjerrett@gmail.com> Co-authored-by: Tejas Chopra <chopratejas@gmail.com>
2026-07-14 21:34:32 +05:30
def test_empty_assistant_message_does_not_desync_fast_skip(self, tmp_path):
# An empty assistant turn (pure tool_use, no output) creates no response
# at parse time, so _ordered_events must not consume a response slot for
# it. Otherwise a later real answer's slot is consumed early, the slow
# reply below is paired with a future-timestamped response, the gap goes
# negative, and a spurious fast_skip is recorded.
p = _write_session(
tmp_path,
"s",
[
_user("explain", ts="2026-01-01T00:00:00Z"),
_empty_assistant(ts="2026-01-01T00:00:00Z"),
_assistant(LONG, ts="2026-01-01T00:00:01Z"), # real answer #1
# +199s: a slow, considered reply — NOT a fast skip.
_user("here is my careful follow-up", ts="2026-01-01T00:03:20Z"),
_assistant(LONG, ts="2026-01-01T00:03:21Z"), # real answer #2
_user("thanks", ts="2026-01-01T00:07:00Z"),
],
)
sig, _ = extract_signals([p])
assert sig.fast_skips == 0
feat: output-token reduction — verbosity shaper, per-user learning, counterfactual savings (#965) ## Description Adds the first levers that reduce the tokens the model **writes back** (output), complementing Headroom's existing input compression. Output costs 5× input on Opus-class models and is full of waste (ceremony, restated code, deep "thinking" on routine steps). Two phases in one self-contained PR off `main`: the request-side output shaper, then per-user verbosity learning plus an honest counterfactual savings estimator and dashboard surfacing. ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [x] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **Output shaper** (`output_shaper.py`, opt-in `HEADROOM_OUTPUT_SHAPER=1`): cache-safe verbosity steering appended to the system-prompt tail (5 levels); effort routing that lowers `output_config.effort` on mechanical tool-result continuations; legacy `thinking.budget_tokens` clamp. Never injects effort where absent, never toggles `thinking.type`. - **`headroom learn --verbosity`**: mines Claude Code transcripts for behavioral signals (interrupts, length-adaptive fast-skips, echo ratio), recommends a verbosity level (heuristic + optional `--llm-judge`), and seeds the savings baseline. - **Counterfactual estimator** (`output_savings.py`): per-stratum synthetic-control (estimated) + A/B holdout (measured) with a propagated 95% CI; conversation-stable arm assignment for A/B validity and prefix-cache safety. - **AIMD verbosity controller** (`verbosity_controller.py`): additive-increase / fast-back-off state machine; live signal emission gated off by default. - **Wiring + surfaces**: shaper resolves the learned level; recording rides the existing `transforms_applied` channel through the outcome funnel (no `RequestOutcome` changes); `headroom output-savings` CLI; dashboard "Output Tokens Saved" card. - **Docs**: simple-words user guide + design doc with the counterfactual methodology. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ pytest tests/test_output_savings.py tests/test_output_savings_cli.py \ tests/test_verbosity_learn.py tests/test_verbosity_controller.py \ tests/test_output_shaper.py -q 94 passed in 0.54s $ pytest tests/test_request_outcome.py tests/test_handler_outcome_tag_invariant.py \ tests/test_proxy_dashboard_stats_cache.py -q 44 passed $ ruff format --check . 831 files already formatted $ mypy headroom --ignore-missing-imports Success: no issues found in 361 source files ``` ## Real Behavior Proof - Environment: macOS, Python 3.12 (`.venv`), `anthropic` 0.76, live API model `claude-opus-4-8`. - Exact command / steps: `HEADROOM_OUTPUT_SHAPER=1`; `headroom learn --verbosity --apply` (seeds level + baseline); `python scripts/eval_output_shaper.py A` (live before/after); simulate holdout traffic then `headroom output-savings`. - Observed result: code-review ask — baseline 1,750 output tokens → L2 1,354 (−22.7%) → L3 599 (−65.8%), same bugs found. `learn --verbosity` on 24 real sessions → 11% interrupt / 26% fast-skip → L3 (high confidence). Measured A/B path → 31.7% reduction (95% CI 27.7%–35.7%). 94 new tests + 44 existing outcome/dashboard tests green; ruff + mypy clean. - Not tested: live streaming-path recording exercised only via unit tests (the `transforms_applied` funnel is shared across paths); runtime AIMD signal emission is gated off by default and not exercised live. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable ## Additional Notes Output savings are counterfactual (we never observe what the model *would* have written), so the estimator separates **estimated** (vs a learned baseline) from **measured** (A/B holdout via `HEADROOM_OUTPUT_HOLDOUT`) and always reports a confidence band — never a single made-up number. CHANGELOG left unchecked (release-please manages it). Runtime AIMD self-tuning is intentionally a TODO (controller built/tested; live signal emission gated behind `HEADROOM_VERBOSITY_AUTOTUNE`).
2026-06-16 21:06:43 -07:00
def test_short_answer_not_skip_eligible(self, tmp_path):
p = _write_session(
tmp_path,
"s",
[
_user("hi", ts="2026-01-01T00:00:00Z"),
_assistant("short reply", ts="2026-01-01T00:00:00Z"),
_user("ok", ts="2026-01-01T00:00:01Z"),
],
)
sig, _ = extract_signals([p])
assert sig.skip_eligible == 0
def test_baseline_captures_output_tokens_by_stratum(self, tmp_path):
p = _write_session(
tmp_path,
"s",
[
_user("task", ts="2026-01-01T00:00:00Z"),
_assistant("a reply", ts="2026-01-01T00:00:01Z", out_tokens=420, in_tokens=5000),
],
)
_, baseline = extract_signals([p])
assert baseline.total_samples == 1
# new_user_ask, input bucket "s" (5000), opus, no tools in this session
mean, _, n = baseline.lookup("opus|new_user_ask|s|notools")
assert n == 1
assert mean == 420.0
def test_tool_result_makes_session_have_tools(self, tmp_path):
p = _write_session(
tmp_path,
"s",
[
_user("task", ts="2026-01-01T00:00:00Z"),
_assistant("reading", ts="2026-01-01T00:00:01Z", out_tokens=50),
_tool_result(ts="2026-01-01T00:00:02Z"),
_assistant("done", ts="2026-01-01T00:00:03Z", out_tokens=200),
],
)
_, baseline = extract_signals([p])
# Every response in a tool-using session is stratified as has_tools.
assert any("|tools" in k for k in baseline.strata)
assert not any("|notools" in k for k in baseline.strata)
def test_tool_result_reply_not_counted_as_human(self, tmp_path):
p = _write_session(
tmp_path,
"s",
[
_user("task", ts="2026-01-01T00:00:00Z"),
_assistant("reading", ts="2026-01-01T00:00:01Z"),
_tool_result(ts="2026-01-01T00:00:02Z"),
],
)
sig, _ = extract_signals([p])
assert sig.human_msgs == 1 # only the real ask, not the tool_result
class TestRecommendLevel:
def _sig(self, *, human, interrupts, skip_eligible, fast_skips) -> VerbositySignals:
s = VerbositySignals()
s.human_msgs = human
s.interrupts = interrupts
s.skip_eligible = skip_eligible
s.fast_skips = fast_skips
return s
def test_too_few_turns_defaults_l2_low(self):
level, conf, _ = recommend_level(
self._sig(human=3, interrupts=0, skip_eligible=0, fast_skips=0)
)
assert level == 2
assert conf == "low"
def test_low_pressure_user_gets_l1(self):
# 100 turns, almost no interrupts/skips.
s = self._sig(human=100, interrupts=1, skip_eligible=100, fast_skips=2)
level, conf, _ = recommend_level(s)
assert level == 1
assert conf == "high"
def test_moderate_pressure_gets_l2(self):
s = self._sig(human=80, interrupts=8, skip_eligible=80, fast_skips=12)
level, _, _ = recommend_level(s)
assert level == 2
def test_high_pressure_gets_l3(self):
# Mirrors the real measured user: ~11% interrupt, ~26% skip.
s = self._sig(human=200, interrupts=29, skip_eligible=119, fast_skips=31)
level, conf, _ = recommend_level(s)
assert level == 3
assert conf == "high"
class TestAnalyze:
def test_llm_judge_overrides_heuristic(self, tmp_path):
p = _write_session(
tmp_path,
"s",
[_user("x", ts="2026-01-01T00:00:00Z"), _assistant("y", ts="2026-01-01T00:00:01Z")]
* 20,
)
def judge(signals_dict):
return 4, "LLM says this user wants caveman mode"
profile, _ = analyze([p], "/proj", llm_judge=judge)
assert profile.level == 4
assert profile.source == "llm"
assert "caveman" in profile.rationale
def test_llm_judge_failure_falls_back_to_heuristic(self, tmp_path):
p = _write_session(
tmp_path,
"s",
[_user("x", ts="2026-01-01T00:00:00Z"), _assistant("y", ts="2026-01-01T00:00:01Z")]
* 20,
)
def bad_judge(signals_dict):
raise RuntimeError("no api key")
profile, _ = analyze([p], "/proj", llm_judge=bad_judge)
assert profile.source == "heuristic"
def test_profile_roundtrip(self, tmp_path):
from headroom.learn.verbosity import VerbosityProfile
prof = VerbosityProfile(
project_path="/proj",
level=3,
confidence="high",
source="heuristic",
rationale="because",
signals={"interrupt_rate": 0.11},
)
path = tmp_path / "verbosity.json"
prof.save(path)
loaded = VerbosityProfile.load(path)
assert loaded is not None
assert loaded.level == 3
assert loaded.confidence == "high"
def test_load_missing_returns_none(self, tmp_path):
from headroom.learn.verbosity import VerbosityProfile
assert VerbosityProfile.load(tmp_path / "nope.json") is None
fix(learn): handle Windows UTF-8, drive-letter paths, and CLI shim fallback (#1895) ## Description `headroom learn --verbosity` is broken on Windows in three related ways: - Transcript/profile reads can use the platform default codec, so non-ASCII content can raise `UnicodeDecodeError` and collapse learning signals to empty output. - `--project <path>` can miss real Claude project directories because Windows profile junctions can raise `PermissionError` during directory walks, and escaped Claude project folder names cannot always distinguish `vibe-remote` from `vibe\remote`. - `headroom learn --agent codex` can fail with `` `claude` not found in PATH `` even when the npm-installed CLI exists, because Windows `.cmd` shims require `PATHEXT` resolution. Refs https://github.com/headroomlabs-ai/headroom/issues/1624 for the Windows learn failures. The dashboard-hint UX and third-party-provider-auth items in that issue are unrelated and out of scope for this PR. ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `headroom/learn/verbosity.py`: read and write verbosity transcripts/profiles with `encoding="utf-8"` so non-ASCII content works regardless of the Windows locale codec. - `headroom/learn/plugins/claude.py`: skip inaccessible siblings one entry at a time during greedy project path decoding, so one Windows junction no longer hides valid project directories. - `headroom/learn/plugins/claude.py`: prefer a valid `cwd` found in Claude session JSONL when discovering project paths, which resolves ambiguous escaped folder names such as `vibe-remote` versus `vibe\remote`. - `headroom/learn/analyzer.py`: resolve Windows CLI shim paths through `shutil.which()` after `FileNotFoundError`, then retry once for streaming and non-streaming CLI calls. - `CHANGELOG.md`: document the Windows learn fixes under `Unreleased`. ## Testing - [x] Unit tests pass (`uv run pytest tests/test_verbosity_learn.py::TestWindowsEncoding tests/test_learn/test_scanner.py::TestGreedyPathDecode::test_permission_denied_sibling_does_not_abort_the_walk tests/test_learn/test_analyzer.py::TestWindowsCliShimFallback tests/test_learn/test_scanner.py::TestDecodeProjectPath::test_discover_project_prefers_session_cwd_over_ambiguous_folder_name -q`) - [x] Linting passes (`uv run ruff check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`) - [x] Formatting passes (`uv run ruff format --check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`) - [ ] Type checking passes (`uv run mypy headroom`) not run; no new public type surface - [x] New tests added for the Windows `cwd` disambiguation regression - [x] Manual testing performed ### Test Output ```text uv run ruff format headroom/learn/plugins/claude.py 1 file reformatted uv run ruff check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py All checks passed! uv run ruff format --check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py 2 files already formatted uv run pytest tests/test_verbosity_learn.py::TestWindowsEncoding tests/test_learn/test_scanner.py::TestGreedyPathDecode::test_permission_denied_sibling_does_not_abort_the_walk tests/test_learn/test_analyzer.py::TestWindowsCliShimFallback tests/test_learn/test_scanner.py::TestDecodeProjectPath::test_discover_project_prefers_session_cwd_over_ambiguous_folder_name -q 9 passed in 0.25s ``` CI on current head `c6dbac40` is green. A prior `test (1)` run hit an unrelated timing-sensitive scheduler assertion; GitHub did not permit direct rerun without admin rights, so the empty commit `c6dbac40` retriggered CI and the shard passed. ## Real Behavior Proof - Environment: Windows 11, Python 3.12, real filesystem for the path-decoding reproduction. - Exact command / steps: Ran `uv run pytest tests/test_verbosity_learn.py::TestWindowsEncoding tests/test_learn/test_scanner.py::TestGreedyPathDecode::test_permission_denied_sibling_does_not_abort_the_walk tests/test_learn/test_analyzer.py::TestWindowsCliShimFallback tests/test_learn/test_scanner.py::TestDecodeProjectPath::test_discover_project_prefers_session_cwd_over_ambiguous_folder_name -q`, `uv run ruff check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`, and `uv run ruff format --check headroom/learn/plugins/claude.py tests/test_learn/test_scanner.py`; the path tests create real Windows-style project directories, inaccessible siblings, ambiguous `vibe\remote` versus `vibe-remote` folders, and Claude session JSONL with `cwd` pointing at the intended project. - Observed result: The focused test command returned `9 passed in 0.25s`, Ruff check passed, and Ruff format check passed. The decoder skips the inaccessible sibling and reaches `vibe-remote`; the session-`cwd` test returns `vibe-remote` instead of trusting the ambiguous escaped folder name; the UTF-8 tests round-trip non-ASCII transcript/profile content under a non-UTF-8 Windows-style codec; the CLI shim tests retry once through `shutil.which()` after `FileNotFoundError`. - Not tested: real npm-installed `claude`/`codex` CLI shims, dashboard-hint UX, and third-party-provider auth. ## Review Readiness - [x] I have performed a self-review - [x] Retrospective review completed after opening; it found one missing `cwd` disambiguation case, now fixed in this PR - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Additional Notes - The post-open retrospective review concluded this needed targeted rework rather than only a retrospective sign-off. - The `cwd` recovery commit is `ee720bb1`; `0905be7b` contains the required formatter cleanup; current head `c6dbac40` is an empty CI-rerun commit after GitHub denied direct rerun without admin rights. --------- Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-09 13:49:38 -04:00
class TestWindowsEncoding:
"""Windows defaults text I/O without an explicit encoding to a locale
codec (e.g. GBK, cp1252) rather than UTF-8. Transcripts containing
non-ASCII text then raised UnicodeDecodeError, which was silently
swallowed and produced "Sessions: 0, human turns: 0" (issue #1624).
``locale.getpreferredencoding`` is monkeypatched to simulate that
non-UTF-8 default on any platform, including the UTF-8-default CI/dev
machines this suite normally runs on.
"""
def _write_utf8_session(self, tmp_path: Path, name: str, lines: list[dict]) -> Path:
p = tmp_path / f"{name}.jsonl"
p.write_text(
"\n".join(json.dumps(line, ensure_ascii=False) for line in lines),
encoding="utf-8",
)
return p
def test_parse_session_reads_non_ascii_under_non_utf8_locale(self, tmp_path, monkeypatch):
monkeypatch.setattr("locale.getpreferredencoding", lambda do_setlocale=True: "cp1252")
p = self._write_utf8_session(
tmp_path,
"s",
[
_user("你好,请帮我写代码", ts="2026-01-01T00:00:00Z"),
_assistant("好的," + LONG, ts="2026-01-01T00:00:01Z"),
],
)
responses, humans, _ = _parse_session(p)
assert len(responses) == 1
assert len(humans) == 1
def test_extract_signals_not_empty_under_non_utf8_locale(self, tmp_path, monkeypatch):
monkeypatch.setattr("locale.getpreferredencoding", lambda do_setlocale=True: "cp1252")
p = self._write_utf8_session(
tmp_path,
"s",
[
_user("開始してください", ts="2026-01-01T00:00:00Z"),
_assistant(LONG, ts="2026-01-01T00:00:01Z"),
],
)
sig, _ = extract_signals([p])
assert sig.sessions == 1
assert sig.asst_responses == 1
def test_profile_save_load_roundtrip_non_ascii_under_non_utf8_locale(
self, tmp_path, monkeypatch
):
monkeypatch.setattr("locale.getpreferredencoding", lambda do_setlocale=True: "cp1252")
prof = VerbosityProfile(
project_path="D:\\work\\DPJ",
level=2,
confidence="high",
source="heuristic",
rationale="用户回复很快,倾向于更简短的回答",
signals={},
)
path = tmp_path / "verbosity.json"
prof.save(path)
loaded = VerbosityProfile.load(path)
assert loaded is not None
assert loaded.rationale == prof.rationale
assert loaded.project_path == prof.project_path