headroom/REALIGNMENT/12-decisions-needed.md
Tejas Chopra e0ce4b1d48
fix: remove rtk and lean-ctx CLI context tools (#2677)
## Description

Removes both third-party CLI context tools — **rtk** and **lean-ctx** —
and with them the context-tool selector itself. Headroom no longer
downloads, installs or configures either one, and there is no
replacement.

The previous pass (#2344) gated only three entry points inside
`headroom/cli/wrap.py`. That left the feature reachable in practice:

| Gap | Effect |
|---|---|
| `scripts/install.sh:1544`, `install.ps1:1681` | Ran `rtk init --global
--auto-patch` from bash/PowerShell, **bypassing the Python gate
entirely** — `curl \| sh` still wrote a Claude Code `PreToolUse` hook
regardless of `HEADROOM_RTK` |
| `wrap.py` `_setup_context_tool_for_agent` | **`wrap openhands` was
broken by default**: `rtk_required=True` met a gate returning `None` →
`SystemExit(1)`. Invisible because all 8 openhands tests patched
`_ensure_rtk_binary` to a fake path |
| `proxy/helpers.py`, `subscription/tracker.py` | Proxy shelled out to
`rtk gain` from `/stats`, the dashboard and `headroom perf`; the tracker
polled it per contribution (`_RTK_WIRING_DEFAULT = "enabled"`) |
| No cleanup path | Nothing removed artifacts an earlier default had
installed, so a machine that once ran the old default kept rtk in the
loop forever (#1669, #1955) |

Also worth noting: the rtk binary download had **no SHA or signature
verification** — only `rtk --version` as a smoke test.

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [x] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [x] Code refactoring (no functional changes)

## Changes Made

**Removed** — `headroom/rtk/` and `headroom/lean_ctx/` packages,
`headroom/cli/wrap_rtk_metrics.py`, `_selected_context_tool` /
`_setup_context_tool_for_agent` / `_VALID_CONTEXT_TOOLS`, the `--rtk` /
`--no-rtk` / `--no-project-rtk` / `--keep-rtk` flags across all 18 wrap
subcommands, `HEADROOM_RTK*`, the proxy-side `rtk gain` polling, the
dashboard CLI-filtering panel (rows + all 8 `cliFiltering*` Alpine
getters), `paths.rtk_path()` / `lean_ctx_path()`, the SDK path helpers,
`benchmarks/rtk_loop_learn_eval.py`, and the `headroom/rtk/**` CI path
filters.

**Fails loudly, not silently** — `--context-tool` / `--no-context-tool`
/ `HEADROOM_CONTEXT_TOOL` are kept solely to error out. They live in
shell profiles, aliases and CI jobs, and accepting them as a no-op would
read as Headroom having quietly stopped working. The installers reject
them too, which matters more than it looks: their arg parsers forward
the first unknown flag **and everything after it** to the wrapped tool,
so a leftover `--no-rtk` would have silently swallowed a following
`--port` and then been ignored downstream.

**New `headroom/context_tool_cleanup.py`** — deleting the code cannot
help a machine that already ran the old default, since the hooks,
binaries and injected guidance are durable on disk.
`purge_context_tool_artifacts()` runs once per `wrap`/`unwrap` and
removes the registered hook entries, the generated hook scripts, the
Headroom-managed `~/.local/bin` symlinks, the vendored
`~/.headroom/bin/{rtk,lean-ctx}` binaries, the `lean-ctx` MCP server
entry and the marker-fenced instruction blocks. Deliberately
conservative: idempotent, **skips** a malformed config rather than
overwriting it, and only unlinks a symlink resolving inside Headroom's
own bin dir so a user's own build is untouched. It reports on
**stderr**, because `wrap/unwrap openclaw --prepare-only` emit
machine-readable JSON on stdout as their entire contract. Skipped for
`wrap selfheal` (runs from a SessionStart hook; must not race Claude
Code's writer for `~/.claude.json`) and for `--help`, which must stay
read-only.

**Client-config hardening** (discovered while investigating a "corrupted
Serena settings file" report) — `wrap.py` reset a settings file to `{}`
when an existing file would not parse, then wrote that back. One
hand-edited typo or a transient `EACCES`/`EINTR` on a valid file
destroyed the user's `permissions`, `env` and `hooks`, on **every
`headroom wrap claude`**. It now refuses to write. Separately,
`fsutil.write_text` is now atomic (temp file + `fsync` + `os.replace`),
fixing all 14 non-atomic client-config writes at once; it follows
symlinks rather than replacing them (dotfile managers) and preserves an
existing file's mode.

**Deliberately kept** — `rtk` stays in the wrapper-peel list in
`transforms/content_router.py`. It sits beside `sudo`/`env`/`timeout` as
shell-command grammar, so `rtk cat f` is still classified as a file read
for anyone running their own rtk install, which the purge intentionally
leaves alone.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ ruff check headroom/ tests/ e2e/ --exclude headroom/dashboard/templates
All checks passed!

$ ruff format --check headroom/ tests/ e2e/ --exclude headroom/dashboard/templates
1255 files already formatted

$ mypy headroom/
Success: no issues found in 508 source files

$ pytest tests/test_context_tool_cleanup.py -q
11 passed

$ pytest tests/test_fsutil.py -q
12 passed

$ pytest tests/test_cli/test_wrap_codex.py -q            # 89 tests
89 passed in 431.68s
$ pytest tests/test_cli/test_wrap_opencode.py -q
39 passed in 257.46s
$ pytest tests/test_cli/test_wrap_helpers.py -q
45 passed
$ pytest tests/test_paths.py -q
75 passed
$ pytest tests/test_cli/test_unwrap_claude.py -q
14 passed
$ pytest tests/test_proxy_savings_history.py -q
39 passed
$ pytest tests/test_cli/test_wrap_copilot.py -q
27 passed
$ pytest tests/test_cli/test_wrap_zcode.py -q
20 passed
$ pytest tests/test_subscription_tracker.py -q
9 passed
$ pytest tests/test_proxy_dashboard_stats_cache.py -q
5 passed, 1 skipped
```

Repo-wide grep for 14 removed symbols (`headroom.rtk`,
`headroom.lean_ctx`, `_ensure_rtk_binary`, `_selected_context_tool`,
`_get_context_tool_stats`, `rtk_path`, `lean_ctx_path`,
`wrap_rtk_metrics`, `HEADROOM_RTK`, `cli_tokens_avoided`,
`tokens_saved_rtk`, …) across `*.py`, `*.ts`, `*.sh`, `*.ps1`, `*.yml`,
`*.html`: **zero hits**.

Notable test changes: `test_wrap_openhands.py` no longer patches
`_ensure_rtk_binary` and asserts `wrap openhands --prepare-only` exits 0
unpatched — the regression that was previously masked.
`test_wrap_continue.py` and `test_wrap_hintfile_agents.py` were removed
(every test drove RTK instruction injection). A new
`test_subscription_tracker.py::test_load_state_written_before_cli_context_tools_were_removed`
proves a pre-removal `subscription_state.json` still loads.

## Real Behavior Proof

- **Environment:** macOS 15.4 (darwin 25.4.0), Python 3.12.6, Headroom @
this branch, real `~/.headroom` and `~/.claude` on the dev machine.
- **Exact command / steps and observed result:**

```text
# 1. Retired flag fails loudly instead of silently no-op'ing
$ headroom wrap codex --prepare-only --context-tool rtk
Error: CLI context tools (rtk, lean-ctx) have been removed from Headroom: they
rewrote shell commands through a third-party binary Headroom no longer manages.
Drop --context-tool / --no-context-tool and unset HEADROOM_CONTEXT_TOOL;
`headroom wrap` uninstalls what they left behind on first run.

$ HEADROOM_CONTEXT_TOOL=lean-ctx headroom wrap codex --prepare-only
Error: CLI context tools (rtk, lean-ctx) have been removed from Headroom: ...

# 2. install.sh rejects the retired flags (extracted parse_wrap_args harness)
['--no-rtk', '--port', '9999']   rc=1  ERROR: CLI context tools ... Drop --no-rtk
['--context-tool=rtk']           rc=1  ERROR: CLI context tools ... Drop --context-tool
$ bash -n scripts/install.sh   # syntax OK

# 3. Purge ran against the real machine, which had all the orphaned artifacts
$ python -c "from headroom.context_tool_cleanup import purge_context_tool_artifacts; ..."
  removed ~/.headroom/bin/lean-ctx        (51 MB)
  removed ~/.headroom/bin/rtk             (7.7 MB)
  removed ~/.local/bin/rtk                (symlink into ~/.headroom/bin)
  removed ~/.claude/hooks/rtk-rewrite.sh
  removed 8 lean-ctx-* hook scripts
# ~/.claude.json afterwards: 90 top-level keys, 19 projects, mcpServers unchanged
# → ~59 MB reclaimed, no unrelated key touched

# 4. stdout stays machine-readable while the purge reports (planted a fake artifact)
$ headroom wrap openclaw --prepare-only --gateway-provider-id codex >out 2>err
$ cat out
{"enabled":true,"config":{"proxyPort":8787,...}}     # parses as JSON
$ cat err
Retired CLI context tool cleanup: removed /Users/tcms/.headroom/bin/rtk

# 5. --help is inert (planted artifact survives), a real run purges
$ headroom wrap codex --help   → artifact survived: CORRECT
$ headroom wrap openclaw --prepare-only → purged: CORRECT

# 6. MCP purge dry-run against a copy of the real 82 KB ~/.claude.json
top-level keys 90 -> 90;  projects 19 -> 19;  LOST keys: none
all content outside mcpServers byte-identical: True
```

Dashboard rendered via the Playwright test after the panel removal:
"Token Savings" shows only `Proxy 0 (0.0%)` / `Of total wire: 36.86%`,
and "Token Usage" reads Before Compression → Proxy Removed → After
Compression with no "Filtered (this session)" row. Nothing below the
removed panel broke.

- **Not tested:** Windows and Linux (macOS only) — `install.ps1` is
verified by brace-balance and inspection, not executed, since no `pwsh`
is available locally. The wrap e2e suite (`e2e/wrap/run.py`) was updated
but not run; it needs the Docker e2e image. `serena project index`
interaction is exercised in the stacked base PR.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md` — it is generated by
release-please from my Conventional Commit PR title (a CI guard enforces
this)

## Additional Notes

**Stacked on #2676** (`tejas/serena-config-bootstrap`) — please merge
that first; this PR's base should then be retargeted to `main`, or it
will read as containing that fix too.

**Breaking-change migration for users:**
- Drop `--rtk`, `--no-rtk`, `--no-project-rtk`, `--keep-rtk`,
`--context-tool`, `--no-context-tool` from any alias, script or CI job,
and unset `HEADROOM_RTK*` / `HEADROOM_CONTEXT_TOOL`. They now error
rather than being ignored, so the failure is immediate and
self-explaining.
- Previously-installed artifacts are purged automatically on the next
`wrap`/`unwrap`; no manual cleanup needed.
- `headroom perf --json` no longer carries a `cli_filtering` key, and
`/stats` no longer returns a `context_tool` section.

**Docs:** `docs/rtk-architecture.md` deleted; RTK/lean-ctx removed from
`README.md`,
`docs/content/docs/{configuration,opencode,grok-build,docker-install,filesystem-contract}.mdx`,
`docs/observability.md` and the matching `wiki/` pages.
`REALIGNMENT/09-phase-G-rtk-observability.md` is marked SUPERSEDED
rather than deleted, to keep the planning record.

**Follow-ups not in scope:** `_emit_wrap_interrupted` was deleted as
dead code — its only caller was the `except KeyboardInterrupt` guarding
the binary download, so with no download there is nothing slow left to
interrupt.
2026-07-30 22:59:41 -07:00

11 KiB
Raw Permalink Blame History

12 — Decisions Needed

Open questions the realignment can't resolve unilaterally. Greenlight or alternative each before the corresponding PR lands.


Q1. Phase A timing — land tonight or wait?

Recommendation: Land PR-A1 tonight. It's a small diff (-180/+30) that eliminates the worst cache-killer cluster (P0-3, P0-4, P0-5 stop firing immediately). The proxy goes to passthrough on /v1/messages; compression returns in Phase B. Net positive because today's compression is actively destroying cache hit rate.

PR-A2 through PR-A8 land over the rest of the week.

Alternative: Hold all of Phase A until the synthesis is "perfect." Risk: cache hit rate stays poor.


Q2. ICM removal scope — Tier 1+2, or include Tier 3?

Recommendation: Tier 1 + Tier 2 in Phase B PR-B1 (~10K LOC).

  • Tier 1 (ICM proper): intelligent_context.py, manager.rs, icm.rs, the proxy call site.
  • Tier 2 (subsystems whose only consumer is ICM): RollingWindow, ProgressiveSummarizer, scoring.py, tool_crusher.py, MessageScorer, all of crates/headroom-core/src/scoring/ and relevance/, most of context/ (keep safety.rs).
  • Tier 3 (separable cleanup): CacheAligner rewrite path is in Phase A PR-A2 (already scheduled). Memory _inject_system_context paths in Phase A PR-A2 + Phase B PR-B6 (already scheduled).

So "Tier 1 + Tier 2" is the right scope for the Phase B big-delete PR; Tier 3 is already covered by Phase A and Phase B's other PRs.

Alternative: Stop at Tier 1 (just ICM proper). Risk: ~6 K LOC of dead-but-still-imported scoring/relevance machinery; future contributors won't know it's dead.


Q3. MessageScorer Rust port — delete?

Recommendation: Delete.

The PR #338 / #343 port (April 2026) was investment in the wrong abstraction (per Agent G's audit: scoring's only consumer is DropByScoreStrategy::try_fit, which Phase B retires). Keeping it as a dead crate creates maintenance debt and confusion. Sunk cost stays sunk; the parity-harness scaffolding learnings carry forward to live-zone work where they actually matter.

Folded into Phase B PR-B1.

Alternative: Keep the crate around as off-path "in case scoring is needed later." Risk: dead-code review burden every PR.


Q4. Stage 3g (lossless-first compression pipeline, issue #315) — re-scope or close?

Context: Per project memory ~/.claude/projects/-Users-tchopra-claude-projects-headroom/memory/project_lossless_first_pipeline.md, Stage 3g was queued to formalize "lossless-then-lossy-then-CCR ordering as a CompressionPipeline orchestrator + LosslessTransform/LossyTransform traits." The plan assumed an ICM-style orchestrator over the messages array.

Recommendation: Re-scope issue #315 to "live-zone-only pipeline orchestrator." The traits stay (LosslessTransform/LossyTransform); the scope changes from "history compactor" to "live-zone block dispatcher." This is what Phase B PR-B2 builds. Update issue #315's body to reflect the realignment.

Alternative: Close issue #315 and treat Phase B PR-B2 as fulfilling its intent. Risk: history of the decision is lost.


Q5. Headroom Loop / AWS Marketplace BYOC — affected scope?

Context: Per project memory project_headroom_loop.md (enterprise paid product) and project_headroom_aws_marketplace.md (BYOC CFN stack in customer VPC). Both depend on the OSS proxy.

Recommendation: The realignment strengthens both:

  • Headroom Loop's value proposition is "trace stream + enterprise compression policy"; Phase F's auth-mode policy is exactly the surface Loop wants to gate on.
  • AWS Marketplace BYOC's pitch is "context compression in front of Bedrock"; Phase D's native Bedrock support makes that pitch real (today's LiteLLM-converted Bedrock path was fake; Phase D fixes it).

No re-scoping needed; revisit after Phase D lands.

Alternative: Pause Headroom Loop / Marketplace work until Phase D completes. Recommended if their roadmap conflicts with Phase D timing.


Q6. make test-parity per-PR gate — enable now or wait?

Recommendation: Enable now (Phase I PR-I6) with the existing stubs. Skipped is permitted; Diff fails the build. As Phase I PR-I5 promotes stubs to real comparators, the per-PR gate gradually tightens.

Alternative: Wait until all stubs are real. Risk: parity divergence merges silently for the next month.


Q7. Operator config switch — explicit HEADROOM_PROXY_BACKEND env var, or implicit?

Context: During Phase H rollout, operators need a way to choose Python vs Rust proxy.

Recommendation: Add HEADROOM_PROXY_BACKEND={python|rust} env var in Phase H PR-H1; default to rust once the canary in Phase I PR-I4 confirms ≥99.9% byte-equality. Keep the Python proxy alive in the codebase for 30 days post-Phase-H as an explicit rollback target. After 30 days of stable Rust operation, run Phase H PR-H2/H3 to delete Python.

Alternative: Cut over implicitly (headroom proxy start always uses Rust after Phase H). Riskier; no clean rollback path.


Q8. Container image strategy — single binary or multi-stage?

Recommendation: Single binary (headroom-proxy Rust). Container is FROM scratch or FROM gcr.io/distroless/static. Image size drops from ~500 MB (with Python + LiteLLM + ONNX models) to ~50 MB.

Alternative: Multi-stage Docker with Rust binary + Python sidecar (for evals/learn/memory writers). Recommended only if those subsystems become production-relevant; today they're CLI tools.


Q9. RTK proxy-side invocation — ever revisit?

Resolved — moot. RTK was removed from Headroom outright (see 09-phase-G-rtk-observability.md), so there is no proxy-side invocation to revisit. The original recommendation was "no, document the decision in docs/rtk-architecture.md" (that doc was deleted with the feature). The argument is kept because reasons 13 apply to any future shell-output rewriter:

  1. Cache hot zone risk: shell-out + buffer per tool result is correctness-fragile.
  2. Parallel implementation: crates/headroom-core/src/transforms/log_compressor.rs covers post-hoc log/output compression; RTK rewrites commands (different value).
  3. RTK itself is a third-party binary the team doesn't control; an upstream version change silently busts cache.

If a future requirement emerges (e.g., "Headroom must compress shell output for users who don't run wrap"), reconsider with explicit cache-safety design.

Alternative: Build proxy-side RTK as a feature-flagged opt-in. Recommended only if the wrap-CLI breadth (PR-G1) doesn't cover enough surface.


Q10. Bedrock/Vertex priority — parallel with proxy port (Phase D in calendar) or after Phase H?

Recommendation: Parallel. Phase D blocks H2 (Python LiteLLM retirement) but not H1 (Python proxy retirement). Run Phase D and Phase C/E/F/G concurrently.

Alternative: Sequential, Phase D after Phase H. Risk: Bedrock/Vertex users stay on the broken Python LiteLLM path for an extra month.


Q11. Memory subsystem — auto-tail mode default, or tool-only?

Recommendation: Auto-tail mode default in Phase B PR-B6, with tool-only mode behind a flag. Migrate users to tool-only over the next 6 months once docs and tooling are mature. Auto-tail is byte-deterministic (per the cache-safety invariant) and matches existing UX.

Alternative: Force tool-only immediately. Risk: breaks customers' existing memory-augmented prompts.


Q12. Parity harness post-Phase-H — keep or delete?

Context: After Phase H deletes Python, crates/headroom-parity/ no longer has a Python side to compare against. Per Phase H PR-H3, this is a decision point.

Recommendation: Repurpose, don't delete. Rename to crates/headroom-version-parity/ and use it to compare current-Rust-version vs previous-Rust-version on the recorded fixtures. Catches Rust-vs-Rust regressions during future ML compressor variants (e.g., when Kompress is ported to Rust via ort).

Alternative: Delete entirely. Save ~2K LOC. Risk: no automated regression test for compressor changes.


Q13. Auth-mode UA detection list — which CLIs to recognize?

Phase F PR-F1 starts with this list:

  • claude-cli/ (Anthropic CLI)
  • claude-code/ (Claude Code)
  • codex-cli/ (Codex CLI)
  • cursor/ (Cursor IDE)
  • claude-vscode/
  • github-copilot/
  • anthropic-cli/
  • antigravity/ (Cloudcode Antigravity)

Recommendation: Extend over time as new CLIs emerge. Alphabetic sort for determinism. Document in docs/auth-modes.md.

Alternative: Start with a smaller list; expand reactively. Risk: subscription users mis-classified as PAYG and fingerprint-leaked.


Q14. The ICM removal blast radius — confirm acceptable

Counts:

  • Lines deleted (Python): ~3,300
  • Lines deleted (Rust): ~4,500
  • Files deleted: ~30
  • Tests deleted: ~50
  • PRs that recently merged but become wasted work: PR #338, PR #343 (MessageScorer Rust port)
  • Project memory updates needed: 1 (the "53270 lines" content_router.py figure was wrong by 25× — already corrected in MEMORY.md).

Recommendation: Acceptable. The cache-killer bugs cost more than the deleted code's hypothetical future value.


Q15. Calendar + capacity — sequential or parallel?

Sequential calendar: ~13 weeks. One contributor working through phases A→I. Parallel calendar: ~8 weeks with 2-3 contributors splitting along these natural boundaries:

  • Lead: Phase A (lockdown), Phase B (live-zone), Phase H (retirement) — the critical path.
  • Contributor 2: Phase C (Rust proxy paths), Phase D (Bedrock/Vertex). Self-contained.
  • Contributor 3 (optional): Phase E (cache stabilization), Phase F (auth-mode), Phase G (RTK + obs), Phase I (test infra). Mostly independent.

Recommendation: Parallel. The bug list is real and the cache hit rate is hemorrhaging in production today.


Quick answer template

For decision sign-off, fill in this block:

Q1 (Phase A timing): [ ] tonight  [ ] wait
Q2 (ICM scope):       [ ] Tier 1+2  [ ] Tier 1 only  [ ] all 3 tiers
Q3 (MessageScorer):   [ ] delete  [ ] keep
Q4 (issue #315):      [ ] re-scope  [ ] close
Q5 (Loop/Marketplace):[ ] proceed unchanged  [ ] pause until D
Q6 (parity gate):     [ ] enable now  [ ] wait
Q7 (operator switch): [ ] env var w/ default rust  [ ] implicit cutover
Q8 (container):       [ ] single binary  [ ] multi-stage
Q9 (RTK proxy-side):  [ ] document never  [ ] feature-flag for future
Q10 (Bedrock priority):[ ] parallel  [ ] sequential after H
Q11 (memory mode):    [ ] auto-tail default  [ ] tool-only force
Q12 (parity harness): [ ] repurpose  [ ] delete
Q13 (UA list):        [ ] approve list  [ ] revise: ___________
Q14 (ICM blast radius): [ ] accept  [ ] reduce scope
Q15 (calendar):       [ ] parallel (2-3 contributors)  [ ] sequential