From f4bd2fe68f0fe8071082380dad0a2eedbfbeb5cd Mon Sep 17 00:00:00 2001 From: Tejas Chopra Date: Fri, 19 Jun 2026 18:14:14 -0700 Subject: [PATCH] docs(vertex): Claude Code + Vertex via Headroom guide (validated) (#1180) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ## Description Documents the **validated** way to run **Claude Code** against **Claude models on Google Vertex AI** with **Headroom compressing the context**. Corrects the prior review's assumption that the "Vertex-mode redirect" approach would work — Claude Code's client-side `probeVertexModel` blocks it — and documents the working **Anthropic-mode + LiteLLM `vertex_ai`** path, verified end-to-end against live Vertex quota (~22% context compression observed). Closes # ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [x] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **`docs/claude-code-vertex-headroom.md`** (new) — copy-paste runbook: prerequisites (GCP ADC, `google-cloud-aiplatform`, Vertex quota), two-terminal setup (proxy `--backend litellm-vertex_ai --region --code-aware`; Claude Code in normal Anthropic mode via `ANTHROPIC_BASE_URL`), verification, a troubleshooting table, and a section on what `--code-aware` does and what it never touches (local files / protected `Read`/`Glob`/`Grep`/`Write`/`Edit` output). - **`wiki/vertex.md`** — new "Claude Code with Headroom compression" section pointing at the runbook, with the two ⚠️ caveats (Vertex-mode probe rejects custom URLs; `vertexai` dep + `--code-aware` required). - **`docs/proposals/vertex-claude-compression-review.md`** — corrected TL;DR: Setup A is blocked by Claude Code's probe; Setup B is the validated path. ## Testing - [ ] Unit tests pass (`pytest`) — **N/A (docs-only, no code changed)** - [ ] Linting passes (`ruff check .`) — **N/A (no Python changed)** - [ ] Type checking passes (`mypy headroom`) — **N/A (no Python changed)** - [ ] New tests added for new functionality — **N/A (docs)** - [x] Manual testing performed (live Vertex validation — see below) ### Test Output ```text # 1) Direct Vertex quota check (global) POST .../locations/global/publishers/anthropic/models/claude-sonnet-4-6:rawPredict -> HTTP 200 {"content":[{"text":"VERTEX OK"}], "model":"claude-sonnet-4-6"} # 2) Headroom in Anthropic mode -> LiteLLM(vertex_ai) -> Vertex global POST http://127.0.0.1:8787/v1/messages (model=claude-sonnet-4-6) -> HTTP 200 {"content":[{"text":"LITELLM VERTEX OK"}], "model":"claude-sonnet-4-6"} # 3) Real Claude Code session (normal mode) through Headroom, --code-aware ON claude -p "...run two Bash source dumps + summarize..." (ANTHROPIC_BASE_URL=proxy) -> is_error: False, modelUsage: ['claude-sonnet-4-6'] request_log: orig=9353 saved=2029 (21.7%) transforms=['router:tool_result:mixed'] # 4) Compressors loaded (GET /debug/warmup) {'kompress':'loaded', 'code_aware':'loaded', 'tree_sitter':'loaded', 'smart_crusher':'loaded'} ``` ## Real Behavior Proof - **Environment:** macOS (arm64); Claude Code 2.1.181; Headroom 0.27.0; venv Python 3.12; LiteLLM `vertex_ai` via `google-cloud-aiplatform` 1.158.0; GCP project `eternal-sunset-495505-t0`; Vertex location `global`; model `claude-sonnet-4-6` (only model with quota on this project); auth via `gcloud auth application-default login` (ADC). - **Exact command / steps:** the two-terminal setup in `docs/claude-code-vertex-headroom.md` — proxy `headroom proxy --port 8787 --backend litellm-vertex_ai --region global --code-aware`; client `ANTHROPIC_BASE_URL=http://127.0.0.1:8787` + `ANTHROPIC_MODEL=claude-sonnet-4-6` in normal mode (no `CLAUDE_CODE_USE_VERTEX`). - **Observed result:** Claude Code answered via Vertex (`modelUsage: claude-sonnet-4-6`); ~22% context compression (`router:tool_result:mixed`) on a code-heavy request forwarded to Vertex `global`; all compressors loaded. - **Not tested:** cumulative savings over long multi-turn sessions; non-global regions (no quota on this project); Opus 4.8 (not enabled in this project — 404); automated tests for the LiteLLM-vertex path (still absent — pre-existing gap). ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines (docs) - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas (N/A — docs) - [x] I have made corresponding changes to the documentation (this *is* the documentation) - [x] My changes generate no new warnings - [ ] I have added tests that prove my fix is effective or that my feature works — **N/A (docs-only)** - [ ] New and existing unit tests pass locally with my changes — **N/A (no code changed)** - [ ] I have updated the CHANGELOG.md if applicable — **N/A (docs-only)** ## Additional Notes - **Docs-only PR** — no Python changed, so `ruff` / `mypy` / `pytest` are N/A. - **Base:** branched from latest `origin/main`; clean 3-file diff (the prerequisite review doc and Vertex wiki content are already on `main`). - **Follow-ups:** optional `headroom wrap claude` Vertex turnkey; add automated tests for the LiteLLM-vertex path; consider defaulting `--code-aware` (or warning when code content is detected but code-aware is off), since its default-off state makes compression silently no-op on coding sessions. --- docs/claude-code-vertex-headroom.md | 163 ++++++++++++++++++ .../vertex-claude-compression-review.md | 35 ++-- wiki/vertex.md | 20 +++ 3 files changed, 206 insertions(+), 12 deletions(-) create mode 100644 docs/claude-code-vertex-headroom.md diff --git a/docs/claude-code-vertex-headroom.md b/docs/claude-code-vertex-headroom.md new file mode 100644 index 000000000..c8b0b655e --- /dev/null +++ b/docs/claude-code-vertex-headroom.md @@ -0,0 +1,163 @@ +# Claude Code + Google Vertex AI, with Headroom compression + +*Validated end-to-end on 2026-06-19 (Claude Code 2.1.181, Headroom 0.27.0).* + +This is the **working, tested** way to run **Claude Code** against **Claude models on +Google Vertex AI** with **Headroom compressing the context** in the middle. + +## TL;DR + +Run Claude Code in **normal Anthropic mode** (NOT Vertex mode) pointed at a local +Headroom proxy, and let **Headroom** be the thing that talks to Vertex: + +``` +Claude Code ──ANTHROPIC_BASE_URL──▶ Headroom proxy ──LiteLLM (vertex_ai)──▶ Vertex AI + (normal mode) (plain http) (compresses) (your GCP ADC) (Claude) +``` + +Two non-obvious requirements make the difference between "works" and "silently does nothing": + +1. **`pip install "google-cloud-aiplatform>=1.38"`** into the proxy's environment — + LiteLLM's `vertex_ai` provider needs it, or every request 500s with + `No module named 'vertexai'`. +2. **Start the proxy with `--code-aware`** — coding sessions are mostly *source code*, + which routes to the AST/code-aware compressor. It is **disabled by default**, so + without this flag compression no-ops on code and you see `tokens_saved: 0`. + +## Why not "just point Claude Code's Vertex URL at Headroom"? + +That approach (Vertex mode + `ANTHROPIC_VERTEX_BASE_URL`=proxy) **does not work** with +Claude Code today. In Vertex mode Claude Code runs a **client-side `probeVertexModel` +check before any request**. When `ANTHROPIC_VERTEX_BASE_URL` points at a non-Google +host, that probe fails *instantly* (no network call is made) with a misleading +`"The model … is not available on your vertex deployment"`, and the proxy never +receives a byte. This is a Claude Code limitation, not a Headroom bug. The native +`:rawPredict` passthrough in Headroom is correct and compresses (verified by direct +curl) — but the client won't route to it. So we use the Anthropic-mode path below. + +## Prerequisites + +- **Google Cloud auth (ADC).** Run once: `gcloud auth application-default login` + (and `gcloud config set project `). The proxy uses ADC to call Vertex; no + API key is held by Headroom. A service-account JSON via + `GOOGLE_APPLICATION_CREDENTIALS` works too. +- **Vertex Claude quota** for the model + location you intend to use. Confirm with a + direct call before involving Headroom: + ```bash + ACCESS_TOKEN="$(gcloud auth application-default print-access-token)" + curl -sS -X POST \ + -H "Authorization: Bearer ${ACCESS_TOKEN}" -H "Content-Type: application/json" \ + "https://aiplatform.googleapis.com/v1/projects//locations/global/publishers/anthropic/models/claude-sonnet-4-6:rawPredict" \ + -d '{"anthropic_version":"vertex-2023-10-16","max_tokens":20,"messages":[{"role":"user","content":"hi"}]}' + ``` + HTTP 200 → good. 429 → model exists but no quota in that location. 404 → model not + enabled in that project/location. +- **Headroom ML extra** for compression: `pip install "google-cloud-aiplatform>=1.38"` + plus the Kompress ML stack (`torch`, `transformers`, `onnxruntime` — the + `headroom-ai[ml]` extra). The `kompress-v2-base` model downloads from Hugging Face + on first use. + +## Terminal 1 — start the Headroom proxy (Vertex backend) + +```bash +cd /path/to/headroom +source .venv/bin/activate + +export VERTEXAI_PROJECT= +export GOOGLE_CLOUD_PROJECT= +export VERTEXAI_LOCATION=global # match where your quota lives + +headroom proxy --port 8787 \ + --backend litellm-vertex_ai \ # NOTE: the _ai suffix is required + --region global \ # becomes LiteLLM vertex_location + --code-aware # REQUIRED for code compression +``` + +On startup, confirm components loaded: `curl -s localhost:8787/debug/warmup` should +show `kompress: loaded`, `code_aware: loaded`, `tree_sitter: loaded`, +`smart_crusher: loaded`. + +## Terminal 2 — run Claude Code (normal Anthropic mode) against the proxy + +```bash +cd /path/to/your/project + +export ANTHROPIC_BASE_URL=http://127.0.0.1:8787 +export ANTHROPIC_API_KEY=sk-placeholder-not-used # Claude Code needs *a* key to start +export ANTHROPIC_MODEL=claude-sonnet-4-6 # sent to Headroom, mapped to vertex_ai/claude-sonnet-4-6 +export ANTHROPIC_SMALL_FAST_MODEL=claude-sonnet-4-6 # pin background model to one you have quota for + +# Do NOT set CLAUDE_CODE_USE_VERTEX or ANTHROPIC_VERTEX_BASE_URL — those put Claude +# Code into Vertex mode and trigger the broken probe described above. + +claude +``` + +Claude Code now talks plain Anthropic `/v1/messages` to Headroom; Headroom compresses +and forwards to Vertex via LiteLLM, then translates the answer back. + +## Verify compression is happening + +- Dashboard: — "tokens saved" climbs as you work. +- `curl -s localhost:8787/stats` → `tokens.saved`, and `request_logs[].transforms_applied` + (look for `router:tool_result:mixed`, `kompress:*`, `code_aware:*`). +- Savings appear on **large tool outputs** (Bash/Grep/web fetches) and accumulate over + turns. Note: **`Read`/`Glob`/`Grep`/`Write`/`Edit` outputs are protected from the + ContentRouter by default** (safest for coding agents); stale `Read`s are handled + separately by the Read-lifecycle system. So the biggest wins come from non-excluded + large outputs and multi-turn sessions, not single one-shot reads. + +## What `--code-aware` does — and what it never touches + +**What it does.** Code-Aware is an **AST (tree-sitter) compressor for source code that +passes through the proxy inside a request**. It parses the code, **keeps the structure +that matters** — imports, function/class signatures, type annotations, error handlers — +and **shrinks the less-important function bodies**, always emitting **syntactically +valid code** (the output still parses). Languages: Python, JS, TS (tier 1); Go, Rust, +Java, C, C++ (tier 2). The original is **stored for retrieval (CCR, ~5-minute TTL)**, so +if the model needs the exact bytes it can pull them back via the `headroom_retrieve` +tool — compression is reversible, not destructive. + +**What it does NOT touch:** + +- **Your files on disk.** Headroom is a network proxy: it only rewrites the *request + body* in flight on the way to Vertex. It never reads, writes, or modifies any local + file. Code-Aware operates on text that is *already inside the API request*, not on + your repository. +- **Claude Code's `Read` tool output.** `Read`, `Glob`, `Grep`, `Write`, and `Edit` + are in `DEFAULT_EXCLUDE_TOOLS` and are **protected from the ContentRouter by default** + (`protect_recent_reads_fraction = 0.0` ⇒ protect-all). So when Claude Code opens a + file the normal way, **the model sees it verbatim** — Code-Aware does not alter it. + (Stale `Read`s — files you later edit — are handled separately and reversibly by the + Read-lifecycle, replacing the superseded copy with a retrievable marker.) + +**Where it actually applies:** code that reaches the model through *other* channels — +most commonly **`Bash` output that prints code** (`cat file.py`, `sed`, `nl`, build +logs with snippets) or large code in results from non-excluded/custom tools. That is +the content that gets AST-compressed. In the validation run, the ~22% savings came +exactly from two `Bash` commands that dumped source files — not from `Read`. + +**Net:** with the default config your real file reads and edits go to the model +untouched; Code-Aware only trims bulky *incidental* code (shell dumps, logs, pasted +snippets) and keeps the originals retrievable. Omit `--code-aware` if you want zero +code transformation at all (you lose code compression but keep everything else). + +## Model-string notes (Vertex) + +- Headroom maps clean ids to Vertex publisher ids: `claude-sonnet-4-6` → + `vertex_ai/claude-sonnet-4-6` (see `headroom/backends/litellm.py`). Newer models use + the bare alias; older ones are date-pinned (e.g. `claude-sonnet-4-5@20250929`). +- `--region global` works (LiteLLM targets the `aiplatform.googleapis.com` global + endpoint). Use a specific region (`us-east5`, `europe-west1`, …) only if that's where + your quota is. + +## Troubleshooting + +| Symptom | Cause | Fix | +|---|---|---| +| `500 … No module named 'vertexai'` | LiteLLM vertex provider dep missing | `pip install "google-cloud-aiplatform>=1.38"`, restart proxy | +| `tokens_saved: 0` on code | Code-Aware disabled | start proxy with `--code-aware` | +| `tokens_saved: 0` everywhere, `/debug/warmup` shows `kompress: not installed` | ML extra missing | install `headroom-ai[ml]` (torch/transformers/onnxruntime) | +| `"model … not available on your vertex deployment"`, proxy logs nothing | Claude Code is in Vertex mode (probe) | unset `CLAUDE_CODE_USE_VERTEX` / `ANTHROPIC_VERTEX_BASE_URL`; use Anthropic mode above | +| `429 RESOURCE_EXHAUSTED` | no quota in that location | switch `--region`/`VERTEXAI_LOCATION` to where your quota is | +| `404 Publisher Model not found` | model not enabled in project/location | enable it in Vertex Model Garden / request quota | diff --git a/docs/proposals/vertex-claude-compression-review.md b/docs/proposals/vertex-claude-compression-review.md index 1ee28043e..e4486b7cd 100644 --- a/docs/proposals/vertex-claude-compression-review.md +++ b/docs/proposals/vertex-claude-compression-review.md @@ -4,23 +4,34 @@ ## TL;DR (read this first) -**Yes, Headroom can compress Claude traffic that goes to Vertex AI — but only if you -set it up by hand, and only if you work around one real bug.** There is **no -turnkey path today**: `headroom wrap claude` does not know anything about Vertex, -and the one backend flag everybody is told to use (`--backend litellm-vertex`) is -broken. +**Yes — and the working path is now validated end-to-end (2026-06-19). See the +copy-paste runbook: [`docs/claude-code-vertex-headroom.md`](../claude-code-vertex-headroom.md).** -Two setups actually work (both need manual environment variables): +> **2026-06-19 update — tested against live Vertex quota (Claude Code 2.1.181):** +> Of the two setups below, **only Setup B works in practice.** +> +> - **Setup A (Vertex mode + `ANTHROPIC_VERTEX_BASE_URL`→proxy) is blocked by Claude +> Code itself.** In Vertex mode Claude Code runs a client-side `probeVertexModel` +> check *before any request*; pointing its Vertex URL at a non-Google host makes that +> probe fail instantly ("model … not available on your vertex deployment") and the +> proxy never receives a byte. Not a Headroom bug — the native `:rawPredict` +> passthrough is correct and compresses (verified by direct curl), but the client +> won't route to it. +> - **Setup B (normal Anthropic mode + `--backend litellm-vertex_ai`) is the working +> path.** Verified: Claude Code → Headroom → LiteLLM → Vertex (`global`), real +> answers, and **~22% context compression on a code-heavy request**. Two gotchas: +> (1) `pip install "google-cloud-aiplatform>=1.38"` or requests 500 with +> `No module named 'vertexai'`; (2) start the proxy with **`--code-aware`** or code +> content silently no-ops (it is disabled by default). -- **Setup A — keep Claude Code in Vertex mode**, and point its Vertex URL at - Headroom. Headroom compresses, then forwards to real Vertex. -- **Setup B — run Claude Code in normal Anthropic mode**, and let Headroom be the - translator that talks to Vertex (`--backend litellm-vertex_ai`). +There is still **no `headroom wrap claude` turnkey** for Vertex, and the one backend +flag older help text advertises (`--backend litellm-vertex`) is broken — use +`--backend litellm-vertex_ai`. Everything in the *middle* (compression, request/response translation, streaming, tool calls) is implemented correctly. The gaps are all at the **edges**: how the -client is pointed at Headroom, one mis-named backend, and a few env vars Headroom -never sets for you. +client is pointed at Headroom, one mis-named backend, a missing pip extra, a +default-off compressor, and a few env vars Headroom never sets for you. > Correction to an earlier claim: it is **not** true that "the Python proxy just > passes Vertex through without compressing." For the Anthropic publisher it runs diff --git a/wiki/vertex.md b/wiki/vertex.md index 5e5e2cbe2..806e9a087 100644 --- a/wiki/vertex.md +++ b/wiki/vertex.md @@ -65,3 +65,23 @@ Headroom also forwards Anthropic publisher calls on Vertex: The Python proxy preserves caller-supplied Google bearer auth. The native Rust proxy path additionally resolves GCP ADC and injects the bearer token for the Anthropic publisher route. + +## Claude Code with Headroom compression (validated) + +To run **Claude Code** against Claude-on-Vertex **with Headroom compressing the +context**, use the dedicated, tested runbook: + +➡️ **[Claude Code + Vertex + Headroom](../docs/claude-code-vertex-headroom.md)** + +Short version: run Claude Code in **normal Anthropic mode** (`ANTHROPIC_BASE_URL` +→ the proxy) and start the proxy with `--backend litellm-vertex_ai --region +--code-aware`; Headroom holds the GCP ADC creds and calls Vertex. + +> ⚠️ Do **not** put Claude Code into Vertex mode and point `ANTHROPIC_VERTEX_BASE_URL` +> at the proxy. Claude Code's client-side model probe rejects any non-Google Vertex +> URL before sending a request ("model … not available on your vertex deployment"), +> so the proxy is never reached. Use the Anthropic-mode runbook above instead. +> +> ⚠️ Two easy-to-miss requirements: `pip install "google-cloud-aiplatform>=1.38"` +> (LiteLLM `vertex_ai` provider) and the `--code-aware` flag (code compression is +> off by default). Without them you get a 500 or `tokens_saved: 0`.