headroom/docs/proposals
Tejas Chopra f4bd2fe68f
docs(vertex): Claude Code + Vertex via Headroom guide (validated) (#1180)
## Description

Documents the **validated** way to run **Claude Code** against **Claude
models on Google Vertex AI** with **Headroom compressing the context**.
Corrects the prior review's assumption that the "Vertex-mode redirect"
approach would work — Claude Code's client-side `probeVertexModel`
blocks it — and documents the working **Anthropic-mode + LiteLLM
`vertex_ai`** path, verified end-to-end against live Vertex quota (~22%
context compression observed).

Closes # <!-- n/a -->

## Type of Change

- [ ] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [x] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- **`docs/claude-code-vertex-headroom.md`** (new) — copy-paste runbook:
prerequisites (GCP ADC, `google-cloud-aiplatform`, Vertex quota),
two-terminal setup (proxy `--backend litellm-vertex_ai --region <loc>
--code-aware`; Claude Code in normal Anthropic mode via
`ANTHROPIC_BASE_URL`), verification, a troubleshooting table, and a
section on what `--code-aware` does and what it never touches (local
files / protected `Read`/`Glob`/`Grep`/`Write`/`Edit` output).
- **`wiki/vertex.md`** — new "Claude Code with Headroom compression"
section pointing at the runbook, with the two ⚠️ caveats (Vertex-mode
probe rejects custom URLs; `vertexai` dep + `--code-aware` required).
- **`docs/proposals/vertex-claude-compression-review.md`** — corrected
TL;DR: Setup A is blocked by Claude Code's probe; Setup B is the
validated path.

## Testing

- [ ] Unit tests pass (`pytest`) — **N/A (docs-only, no code changed)**
- [ ] Linting passes (`ruff check .`) — **N/A (no Python changed)**
- [ ] Type checking passes (`mypy headroom`) — **N/A (no Python
changed)**
- [ ] New tests added for new functionality — **N/A (docs)**
- [x] Manual testing performed (live Vertex validation — see below)

### Test Output

```text
# 1) Direct Vertex quota check (global)
POST .../locations/global/publishers/anthropic/models/claude-sonnet-4-6:rawPredict
  -> HTTP 200  {"content":[{"text":"VERTEX OK"}], "model":"claude-sonnet-4-6"}

# 2) Headroom in Anthropic mode -> LiteLLM(vertex_ai) -> Vertex global
POST http://127.0.0.1:8787/v1/messages  (model=claude-sonnet-4-6)
  -> HTTP 200  {"content":[{"text":"LITELLM VERTEX OK"}], "model":"claude-sonnet-4-6"}

# 3) Real Claude Code session (normal mode) through Headroom, --code-aware ON
claude -p "...run two Bash source dumps + summarize..."  (ANTHROPIC_BASE_URL=proxy)
  -> is_error: False, modelUsage: ['claude-sonnet-4-6']
  request_log: orig=9353  saved=2029 (21.7%)  transforms=['router:tool_result:mixed']

# 4) Compressors loaded (GET /debug/warmup)
{'kompress':'loaded', 'code_aware':'loaded', 'tree_sitter':'loaded', 'smart_crusher':'loaded'}
```

## Real Behavior Proof

- **Environment:** macOS (arm64); Claude Code 2.1.181; Headroom 0.27.0;
venv Python 3.12; LiteLLM `vertex_ai` via `google-cloud-aiplatform`
1.158.0; GCP project `eternal-sunset-495505-t0`; Vertex location
`global`; model `claude-sonnet-4-6` (only model with quota on this
project); auth via `gcloud auth application-default login` (ADC).
- **Exact command / steps:** the two-terminal setup in
`docs/claude-code-vertex-headroom.md` — proxy `headroom proxy --port
8787 --backend litellm-vertex_ai --region global --code-aware`; client
`ANTHROPIC_BASE_URL=http://127.0.0.1:8787` +
`ANTHROPIC_MODEL=claude-sonnet-4-6` in normal mode (no
`CLAUDE_CODE_USE_VERTEX`).
- **Observed result:** Claude Code answered via Vertex (`modelUsage:
claude-sonnet-4-6`); ~22% context compression
(`router:tool_result:mixed`) on a code-heavy request forwarded to Vertex
`global`; all compressors loaded.
- **Not tested:** cumulative savings over long multi-turn sessions;
non-global regions (no quota on this project); Opus 4.8 (not enabled in
this project — 404); automated tests for the LiteLLM-vertex path (still
absent — pre-existing gap).

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines (docs)
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
(N/A — docs)
- [x] I have made corresponding changes to the documentation (this *is*
the documentation)
- [x] My changes generate no new warnings
- [ ] I have added tests that prove my fix is effective or that my
feature works — **N/A (docs-only)**
- [ ] New and existing unit tests pass locally with my changes — **N/A
(no code changed)**
- [ ] I have updated the CHANGELOG.md if applicable — **N/A
(docs-only)**

## Additional Notes

- **Docs-only PR** — no Python changed, so `ruff` / `mypy` / `pytest`
are N/A.
- **Base:** branched from latest `origin/main`; clean 3-file diff (the
prerequisite review doc and Vertex wiki content are already on `main`).
- **Follow-ups:** optional `headroom wrap claude` Vertex turnkey; add
automated tests for the LiteLLM-vertex path; consider defaulting
`--code-aware` (or warning when code content is detected but code-aware
is off), since its default-off state makes compression silently no-op on
coding sessions.
2026-06-19 18:14:14 -07:00
..
output-token-reduction.md feat: output-token reduction — verbosity shaper, per-user learning, counterfactual savings (#965) 2026-06-16 21:06:43 -07:00
vertex-claude-compression-review.md docs(vertex): Claude Code + Vertex via Headroom guide (validated) (#1180) 2026-06-19 18:14:14 -07:00