mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
fix(proxy/vertex): route google-publisher requests to the request region (#2069)
## Description
The Vertex `publisher=google` routes forward to a **fixed** upstream
host, ignoring the
request's region. In `headroom/providers/proxy_routes.py`,
`vertex_generate_content`,
`vertex_stream_generate_content`, and `vertex_count_tokens` all do:
```python
del api_version, project, location # <-- location discarded
if publisher == "google":
return await proxy.handle_gemini_generate_content(
request, model,
_api_target(proxy, "vertex"), # <-- single fixed host (default us-central1)
"vertex:google",
)
```
The sibling Anthropic `rawPredict` route already does this correctly —
it keeps `location` and
passes `_vertex_target_for_location(proxy, location)`, which derives the
regional host from the
path.
So a request to
`.../locations/europe-west1/publishers/google/models/gemini-2.0-flash:generateContent`
(with the proxy left at the default Vertex URL) is forwarded to
`https://us-central1-aiplatform.googleapis.com/...europe-west1...` — a
`us-central1` host serving a
`europe-west1` path. Vertex requires the host region to match the path
location, so it rejects the
request. `_vertex_target_for_location` and the region-aware Anthropic
routing landed together in
`0e059150`; the three google routes were the missed spot.
Closes: no issue filed — found while auditing Vertex routing.
## Fix
In all three `publisher == "google"` branches, keep `location` and pass
`_vertex_target_for_location(proxy, location)` instead of
`_api_target(proxy, "vertex")`. That
helper honors an operator-pinned non-default upstream (private gateway)
and otherwise derives the
host from the request's `location` (`global` → the unprefixed host).
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
## Changes Made
- `headroom/providers/proxy_routes.py`: region-aware host for the google
generateContent / streamGenerateContent / countTokens routes.
- `tests/test_vertex_claude_compression.py`: add route-level tests that
the google generateContent and countTokens routes forward a
`europe-west1` request to
`https://europe-west1-aiplatform.googleapis.com` (default config),
mirroring the existing anthropic-route test.
## Testing
- [x] New regression tests added
(`tests/test_vertex_claude_compression.py`)
- [x] Linting/formatting clean — run with the CI-pinned `ruff==0.15.17`
- [ ] Full `pytest` deferred to CI (local-OOM reason below).
```text
$ uvx ruff@0.15.17 check headroom/providers/proxy_routes.py tests/test_vertex_claude_compression.py
All checks passed!
```
## Real Behavior Proof
- Environment: Windows 11, Python 3.10, headroom from this branch.
Importing `headroom` pulls in the torch/transformers stack and a full
`pytest` gets OOM-killed on this box, so I verified the host-derivation
with a dependency-free script and left the full pytest to CI.
- Exact command / steps: ran a `europe-west1` request through the old
fixed `_api_target` host and the new `_vertex_target_for_location`, plus
the `us-central1`/`global`/operator-pinned cases.
- Observed result: the old path sends europe-west1 to the us-central1
host (rejected); the new path derives the correct region and still
honors a pinned upstream:
```text
europe-west1: OLD host=https://us-central1-aiplatform.googleapis.com
europe-west1: NEW host=https://europe-west1-aiplatform.googleapis.com
VERTEX REGION ROUTING FIX VERIFIED (old = fixed us-central1; new = per-request region)
```
- Not tested: a live GCP/Vertex round-trip (handlers stubbed, as the
existing tests do). The existing tests that pin a non-default
`vertex_api_url="https://vertex.test"` still pass, since
`_vertex_target_for_location` honors the pinned upstream. Full local
`pytest` deferred to CI (OOM, per above).
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [ ] New and existing unit tests pass locally with my changes — ran
lint + a standalone logic check; full pytest deferred to CI (local OOM,
disclosed above)
- [x] I have updated the CHANGELOG.md if applicable
## Additional Notes
- Reuses the in-file `_vertex_target_for_location` helper the anthropic
route already uses; no new dependencies. (The
non-`google`/non-`anthropic` publisher passthrough is still fixed-host —
a separate, lower-priority follow-up.)
- @JerrettDavis tagging you — non-`us-central1` Vertex Gemini requests
currently fail on a host/region mismatch; this brings the google routes
in line with the anthropic one you reviewed. Thanks!
Co-authored-by: JerrettDavis <mxjerrett@gmail.com>
This commit is contained in:
parent
19201e842f
commit
1843346283
3 changed files with 53 additions and 6 deletions
|
|
@ -102,6 +102,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|||
### Bug Fixes
|
||||
|
||||
* **ccr:** don't crash `parse_tool_call` on a CCR tool call whose arguments aren't an object. For the OpenAI/`openai_responses` shape the arguments are `json.loads`-decoded and only `JSONDecodeError` was caught, so a model that emitted `arguments='[]'`/`'"abc"'`/`'123'` (decoding to a list/str/number) — or a non-dict Anthropic `input` — reached `input_data.get("hash")` and raised `AttributeError`; a null `arguments` raised an uncaught `TypeError` from `json.loads(None)`. Both are now handled: the decode also catches `TypeError`, and a non-dict `input_data` returns `None` (not a valid CCR call) instead of crashing CCR response processing.
|
||||
* **proxy/vertex:** route Vertex `publisher=google` (Gemini) requests to the region matching the request path. `vertex_generate_content`, `vertex_stream_generate_content`, and `vertex_count_tokens` discarded the path's `location` and forwarded to the single fixed host from `_api_target(proxy, "vertex")` (default `us-central1`), instead of the region-aware `_vertex_target_for_location` the sibling Anthropic `rawPredict` route already uses. So a request to `.../locations/europe-west1/publishers/google/...` was sent to a `us-central1` host, which Vertex rejects on the region/host mismatch. The three google routes now derive the host from the request's `location` (operator-pinned upstreams are still honored).
|
||||
* **proxy/anthropic:** give each Anthropic conversation its own session id. `SessionTrackerStore.compute_session_id` derived its fallback id from `model` + system text harvested only from `role:"system"` entries inside `messages` — but Anthropic carries the system prompt as a top-level `body["system"]` field, so genuine Anthropic requests (which never carry `x-headroom-session-id`) collapsed to `md5(model:[])` and every conversation on the same model shared one `PrefixCacheTracker`. That let session-sticky state cross-contaminate: conversation A's sticky `headroom_retrieve`/memory tools and `anthropic-beta` headers were injected into conversation B, and frozen-prefix/compression-cache state mixed across conversations. The Anthropic handler now folds the top-level `system` into the session-id inputs (prepending a synthetic `role:"system"` message used only to derive the id), giving distinct conversations distinct ids.
|
||||
* **cache/semantic:** key entries by the full-context hash, not the trailing query text. `SemanticCache.put` stored each response under `sha256(query)[:16]` where `query` is only the last user message, and the exact-match branch of `get` returned the slot without checking the stored entry's `messages_hash`. Two requests that share a trailing message ("continue", "yes", "run the tests") but differ in earlier context therefore collided on one slot — the second overwrote the first, and the first's hash then resolved to the second's cached response (wrong data served). Entries are now keyed by `messages_hash` when present, and `get` verifies `entry.messages_hash` before returning.
|
||||
* **proxy/openai:** stop overriding an explicit client `stream_options.include_usage` on the streaming chat path. To count tokens from the trailing usage chunk, the handler set `include_usage: True` unconditionally — including flipping an explicit client `false` to `true`. The upstream then appended a usage-only chunk (`choices: []`) the client never requested, and the common `chunk.choices[0].delta` loop raised `IndexError`. The option is now only filled in when the client left the choice open (no `stream_options`, or a dict without `include_usage`); an explicit `true`/`false` is respected.
|
||||
|
|
|
|||
|
|
@ -254,12 +254,12 @@ def register_provider_routes(app: FastAPI, proxy: Any) -> None:
|
|||
publisher: str,
|
||||
model: str,
|
||||
):
|
||||
del api_version, project, location
|
||||
del api_version, project
|
||||
if is_vertex_google_publisher(publisher):
|
||||
return await proxy.handle_gemini_generate_content(
|
||||
request,
|
||||
model,
|
||||
_api_target(proxy, "vertex"),
|
||||
_vertex_target_for_location(proxy, location),
|
||||
VERTEX_GOOGLE_PROVIDER_NAME,
|
||||
)
|
||||
return await vertex_publisher_passthrough(request, publisher, VERTEX_GENERATE_CONTENT.name)
|
||||
|
|
@ -275,12 +275,12 @@ def register_provider_routes(app: FastAPI, proxy: Any) -> None:
|
|||
publisher: str,
|
||||
model: str,
|
||||
):
|
||||
del api_version, project, location
|
||||
del api_version, project
|
||||
if is_vertex_google_publisher(publisher):
|
||||
return await proxy.handle_gemini_generate_content(
|
||||
request,
|
||||
model,
|
||||
_api_target(proxy, "vertex"),
|
||||
_vertex_target_for_location(proxy, location),
|
||||
VERTEX_GOOGLE_PROVIDER_NAME,
|
||||
)
|
||||
return await vertex_publisher_passthrough(
|
||||
|
|
@ -300,12 +300,12 @@ def register_provider_routes(app: FastAPI, proxy: Any) -> None:
|
|||
publisher: str,
|
||||
model: str,
|
||||
):
|
||||
del api_version, project, location
|
||||
del api_version, project
|
||||
if is_vertex_google_publisher(publisher):
|
||||
return await proxy.handle_gemini_count_tokens(
|
||||
request,
|
||||
model,
|
||||
_api_target(proxy, "vertex"),
|
||||
_vertex_target_for_location(proxy, location),
|
||||
VERTEX_GOOGLE_PROVIDER_NAME,
|
||||
)
|
||||
return await vertex_publisher_passthrough(request, publisher, VERTEX_COUNT_TOKENS.name)
|
||||
|
|
|
|||
|
|
@ -142,6 +142,52 @@ def test_vertex_rawpredict_anthropic_runs_compression_handler(monkeypatch) -> No
|
|||
assert captured["model"] == "claude-sonnet-4-6"
|
||||
|
||||
|
||||
def test_vertex_google_generate_content_uses_region_derived_host(monkeypatch) -> None:
|
||||
"""The google-publisher generateContent route must derive the upstream host
|
||||
from the request's location (like the anthropic route), not send a
|
||||
europe-west1 request to the fixed us-central1 host."""
|
||||
captured: dict[str, str] = {}
|
||||
|
||||
# The route calls handle_gemini_generate_content(request, model, base_url, provider).
|
||||
async def fake(self, request, model, base_url, provider, *rest): # type: ignore[no-untyped-def]
|
||||
captured.update(base_url=str(base_url), provider=str(provider), model=str(model))
|
||||
return JSONResponse({"ok": True})
|
||||
|
||||
monkeypatch.setattr(HeadroomProxy, "handle_gemini_generate_content", fake)
|
||||
|
||||
with TestClient(_default_vertex_app()) as client:
|
||||
resp = client.post(
|
||||
"/v1/projects/p/locations/europe-west1/publishers/google/models/"
|
||||
"gemini-2.0-flash:generateContent",
|
||||
json={"contents": []},
|
||||
)
|
||||
assert resp.status_code == 200
|
||||
assert captured["provider"] == "vertex:google"
|
||||
assert captured["base_url"] == "https://europe-west1-aiplatform.googleapis.com"
|
||||
assert captured["model"] == "gemini-2.0-flash"
|
||||
|
||||
|
||||
def test_vertex_google_count_tokens_uses_region_derived_host(monkeypatch) -> None:
|
||||
"""The google-publisher countTokens route is region-aware too."""
|
||||
captured: dict[str, str] = {}
|
||||
|
||||
async def fake(self, request, model, base_url, provider, *rest): # type: ignore[no-untyped-def]
|
||||
captured.update(base_url=str(base_url), provider=str(provider))
|
||||
return JSONResponse({"ok": True})
|
||||
|
||||
monkeypatch.setattr(HeadroomProxy, "handle_gemini_count_tokens", fake)
|
||||
|
||||
with TestClient(_default_vertex_app()) as client:
|
||||
resp = client.post(
|
||||
"/v1/projects/p/locations/europe-west1/publishers/google/models/"
|
||||
"gemini-2.0-flash:countTokens",
|
||||
json={"contents": []},
|
||||
)
|
||||
assert resp.status_code == 200
|
||||
assert captured["provider"] == "vertex:google"
|
||||
assert captured["base_url"] == "https://europe-west1-aiplatform.googleapis.com"
|
||||
|
||||
|
||||
def test_vertex_rawpredict_versionless_anthropic_rewrites_to_v1(monkeypatch) -> None:
|
||||
captured: dict[str, Any] = {}
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue