headroom/tests/test_install
Tejas Chopra b121223ec9
fix(install): default to cache mode, matching headroom proxy (#1893 follow-up) (#2563)
## Description

`headroom install` and `headroom deploy` defaulted `--mode` to
**token**, while `headroom proxy` and the server env default both
resolve to **cache**. Because `install/planner.py:155` writes
`"HEADROOM_MODE": proxy_mode` into the install base env, installing
Headroom did not merely differ from running it directly — it **actively
overrode** the good server default with the cache-busting one.

| Entry point | Effective default | Where |
|---|---|---|
| `headroom proxy` | **cache** | `cli/proxy.py:1129` — `mode or
HEADROOM_MODE or PROXY_MODE_CACHE` |
| `proxy/server.py` env | **cache** | `server.py:4962`, commented
*"delta-only compression at ~0 prefix-cache busts"* |
| `headroom install` / `deploy` | **token**  | `cli/install.py:455,615`
|

Cache mode freezes prior turns and compresses only the newest delta, so
the cached prefix stays byte-identical. Token mode rewrites frozen
history, which moves the bytes the provider hashed for its cache key and
forces a full cold re-write of the entire prefix.

Why that is expensive — measured on 35 local Claude Code sessions
(23,018 turns, 8,985M prompt tokens): cache **writes** are ~46% of input
spend from just 6.3% of tokens, and 714 warm turns that each re-wrote
>100K tokens carried 83% of all warm-path write tokens (~26% of total
input spend) at ~452K tokens per event. Full-prefix re-writes are the
dominant cost in this workload, and token mode makes them more likely.

**This is an oversight, not a deliberate divergence.** #1893 ("ship the
coding profile as Headroom's out-of-box default posture") introduced the
cache default but its diff touched only `agent_savings.py`,
`cli/proxy.py`, and `proxy/server.py` — verified with `git show 68676daa
--stat`. Neither `cli/install.py` nor `install/` was in it. The install
default predates it (#1404 and the persistent-install lifecycle work).

Closes #

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)

## Changes Made

- `cli/install.py` — `--mode` default `token` -> `cache` on **both**
commands (`install_apply`, `deploy`), with the help text stating what
cache mode buys.
- `install/models.py` — `DeploymentManifest.proxy_mode` default `token`
-> `cache`, so a manifest that omits the field no longer falls back to
token either.
- New `tests/test_install/test_proxy_mode_default.py` (5 tests) pinning
the agreement between the two entry points — the regression guard that
was missing when #1893 landed.

`--mode token` remains fully available for anyone who wants maximum
compression and accepts the prefix-cache busts. The option type is
unchanged (free text through `normalize_proxy_mode_value`, aliases
intact), and a test asserts token stays reachable.

## ⚠️ Existing installs are not migrated

A manifest already on disk has `proxy_mode: "token"` serialized
explicitly, so it keeps token until it is re-applied. This PR fixes the
default going forward only. Immediate remedy for affected users:

```bash
export HEADROOM_MODE=cache          # or re-run: headroom install --mode cache
```

Deliberately out of scope here: manifest migration, and a `doctor` check
that would flag an installed-but-token-mode deployment. Happy to follow
up with either.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ python -m pytest tests/test_install/ tests/test_proxy_mode_policy.py -q
146 passed, 2 skipped in 1.19s

$ python -m pytest tests/test_install/test_proxy_mode_default.py -q
5 passed in 0.45s

$ ruff check headroom/ tests/test_install/ --exclude headroom/dashboard
All checks passed!

$ mypy headroom/cli/install.py headroom/install/models.py
Success: no issues found in 2 source files
```

## Real Behavior Proof

- **Environment:** macOS 25.4.0, Python 3.12 (`.venv`), branch off
`upstream/main` @ 58555c5b, run in an isolated `git worktree` with
`PYTHONPATH` pinned to it.
- **Exact command / steps:**
1. Traced the divergence: `grep -n proxy_mode headroom/cli/install.py`
(two `--mode` options, one shared manifest builder) and `grep -n
HEADROOM_MODE headroom/install/planner.py` (line 155 writes it into base
env).
2. Confirmed intent with `git log -S 'PROXY_MODE_CACHE' --
headroom/cli/proxy.py` (-> #1893) and `git show 68676daa --stat`
(install not in the diff).
  3. Ran the suites above.
- **Observed result:** both `--mode` option defaults now report `cache`;
`DeploymentManifest().proxy_mode == "cache"`;
`normalize_proxy_mode_value("token")` still returns token, so the
opt-out path is intact. 146 install/mode tests pass.
- **Not tested:**
- **No end-to-end install performed.** I did not run `headroom install`
against a real system and inspect the written manifest/systemd unit; the
change is verified at the option-default and dataclass-default level
plus the existing install unit suites.
- **The cost claim is measured on Claude Code traffic only**, from
transcripts — not from an A/B of token-vs-cache mode on identical
workloads. Cache mode's "~0 prefix-cache busts" is the repo's own
existing characterization (`server.py:4962`), not something this PR
benchmarked.
  - No migration path for existing manifests is included or tested.
- A broad local `-k "mode"` run accidentally matched every test
containing "**model**" (~1,100 tests) and surfaced 13 failures; the ones
I could identify are pre-existing or environmental —
`test_model_uses_memory_id_to_call_memory_delete` fails identically on
clean `main`, the two `test_langchain_live` errors need live API keys,
and `test_unload_when_no_model` passes in isolation on this branch
(global-state ordering). My captured log was truncated, so I did not
account for all 13 individually; the full sharded suite in CI is the
authoritative check.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md`
2026-07-25 19:26:15 -07:00
..
test_codex_install.py fix(codex): detect keyring-backed ChatGPT auth (#2478) 2026-07-22 06:09:31 -07:00
test_health.py test: expand persistent install coverage 2026-04-11 18:24:15 -05:00
test_native_installers.py fix: B1 — retire ICM, RollingWindow, scoring, relevance + dependents 2026-05-02 12:23:17 -07:00
test_paths.py ci: restore green lint (reformat for ruff 0.15.17, fix mypy no-any-return, pin linters) (#1295) 2026-06-22 15:14:40 -05:00
test_planner.py feat(install): add apply flag parity, --env passthrough, and EIO retry (#2152) 2026-07-14 14:10:39 -04:00
test_providers.py fix(codex): rewrite config.toml properly so Codex will route through … (#2102) 2026-07-14 11:53:24 -04:00
test_proxy_mode_default.py fix(install): default to cache mode, matching headroom proxy (#1893 follow-up) (#2563) 2026-07-25 19:26:15 -07:00
test_runtime.py fix(install): use CREATE_NO_WINDOW instead of DETACHED_PROCESS on Windows (#2527) 2026-07-25 19:24:53 -07:00
test_state.py fix(install): migrate deployments off the retired chopratejas image repo (#2427) 2026-07-19 22:15:01 -07:00
test_supervisors.py feat(install): add apply flag parity, --env passthrough, and EIO retry (#2152) 2026-07-14 14:10:39 -04:00