Commit graph

7 commits

Author SHA1 Message Date
Ingmar Krusch
3c1a5cdb1c
fix(backends/litellm): drop oversized tool names before Bedrock Converse (#2129)
## Description

The Bedrock Converse API hard-rejects any request containing a tool name
over 64 characters (`toolConfig.tools.N.member.toolSpec.name`). Claude
Code includes every globally-added claude.ai MCP connector tool in every
request it sends, even connectors the user hasn't enabled locally. One
org-wide connector with a 65-char tool name is enough to fail every
single request routed through this backend's Bedrock path, with no way
to remove or disable the connector client-side.

Direct Bedrock mode (`CLAUDE_CODE_USE_BEDROCK=1`, bypassing this proxy)
is unaffected: it hits Bedrock's native Anthropic-compatible endpoint,
which has no such length limit. Only the Converse API, which this
LiteLLM-backed `bedrock` provider path uses, enforces it.


## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- `headroom/backends/litellm.py`: `send_message` and `stream_message`
both filter tools with names over 64 characters out of the payload
before converting/forwarding, but only for `self.provider == "bedrock"`.
Other providers are untouched.
- `tests/test_backend_bugs.py`: new
`TestBedrockOversizedToolNameFiltering` covering both `send_message` and
`stream_message` — an oversized (65-char) name is dropped on `bedrock`,
a name at exactly the 64-char boundary is kept, and non-`bedrock`
providers forward oversized names unfiltered (the limit is a Bedrock
Converse constraint, not a general one).
- `CHANGELOG.md`: added a `### Fixed` entry under `Unreleased`.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ uv run pytest tests/test_backend_bugs.py tests/test_backend_anyllm.py -q
============================= test session starts ==============================
platform darwin -- Python 3.13.13, pytest-9.0.3, pluggy-1.6.0
collected 57 items

tests/test_backend_bugs.py ..........................................    [ 73%]
tests/test_backend_anyllm.py ...............                             [100%]

============================== 57 passed in 1.42s ==============================

$ uv run ruff check headroom/backends/litellm.py tests/test_backend_bugs.py
All checks passed!

$ uv run mypy headroom/backends/litellm.py
Success: no issues found in 1 source file
```

## Real Behavior Proof

- **Environment:** personal fork deployed as a real proxy (macOS launchd
service, `headroom install apply`) with `--backend bedrock --mode token
--code-aware --bedrock-profile sso-bedrock`, fronting a live Claude Code
session with a globally-added-but-not-locally-enabled claude.ai MCP
connector (`TopCounsel`) whose tool name is 65 characters.
- **Exact command / steps:** run any Claude Code request through this
deployment while the org-wide `TopCounsel` connector is present (it is
included in the tool list on every request regardless of local
enablement).
- **Observed result:** before the fix, every request failed with a
LiteLLM `BedrockException`: `1 validation error detected: Value
'mcp__claude_ai_TopCounsel_by_The_L_Suite__complete_authentication' at
'toolConfig.tools.N.member.toolSpec.name' failed to satisfy constraint:
Member must have length less than or equal to 64`. After applying the
fix (filtering the oversized tool out before the LiteLLM call), the same
session proceeds normally with no validation error, confirmed live
against this deployment.
- **Not tested:** truncating the name instead of dropping it was tried
and discarded during investigation — the model echoes the truncated name
back in `tool_use` blocks, and Claude Code matches tool calls by the
original full name, so truncation breaks routing on the return path.
This PR drops the tool entirely rather than truncating, which is why it
is not present as an alternative in the diff.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable

## Screenshots (if applicable)

N/A — backend logic change, no UI surface.

## Additional Notes

- "I have made corresponding changes to the documentation" is unchecked:
no file in `docs/`, `README.md`, or `CONTRIBUTING.md` documents
backend-specific tool-list filtering behavior, so there is no existing
section to update.
- No linked issue number: this was found via independent investigation
of a personal deployment (a live Bedrock validation failure), not filed
as a `headroomlabs-ai/headroom` issue first. Checked `gh pr list`/`gh
issue list` for existing coverage of "Bedrock Converse 64-char tool
name" and found none open or merged.
- A native Bedrock Anthropic-compatible endpoint backend (avoiding
Converse's tool-name limit entirely) would be the more complete
long-term fix, but is out of scope for this PR.

---------

Co-authored-by: Ingmar Krusch <ingmar.krusch@gmail.com>
Co-authored-by: Tejas Chopra <chopratejas@gmail.com>
2026-07-13 19:53:49 -04:00
chopratejas
d5ca50cd03 fix(tests): stop module-level dotenv loaders from polluting os.environ during pytest collection
# The bug

Several test modules and two production modules loaded the project `.env`
at *import time*. During pytest collection (where every test module is
imported once), this populated `os.environ` with API keys from `.env`.

The skipif guards in `test_proxy_passthrough_integration.py` (and
others) evaluate at collection time:

    @pytest.mark.skipif(not os.environ.get("OPENAI_API_KEY"), reason="...")

If the polluter module was collected *before* the guard, the guard saw
the leaked key, decided not to skip, and the integration tests ran
live against a fake key and failed. In a fresh local-dev venv with
`.env` + full `[dev]` extras, this manifested as ~16 spurious test
failures plus a misleading test runtime of 6+ minutes (live HTTP).

# Why now

CI does not see this (no `.env`). It only manifests when:
1. `litellm` (and friends) are installed — they run `dotenv.load_dotenv()`
   on import, populating `os.environ` from `.env`.
2. A `.env` file with real API keys exists locally.

Until the venv was provisioned with the full `[dev]` extras during
recent test work, `pytest.importorskip("litellm")` and
`from headroom.pricing import litellm_pricing` both silently no-op'd
(via try/except ImportError → `LITELLM_AVAILABLE=False`), so the leak
never triggered. With litellm now installed, the latent bug surfaced.

# The fix — three patterns

1. **Production modules** (`headroom/pricing/litellm_pricing.py`,
   `headroom/backends/litellm.py`): wrap the eager `import litellm` with
   a snapshot/restore of `os.environ`. Any keys litellm's bundled
   `python-dotenv` adds during import are deleted immediately. The
   module is fully imported and cached in `sys.modules` so subsequent
   imports hit the cache without re-running the side effect.

2. **Test modules using `pytest.importorskip("litellm")`**
   (`test_backend_bugs.py`, `test_bedrock_region.py`,
   `test_cost_tracker_counterfactual.py`): replace with
   `tests._dotenv.importorskip_no_env_leak("litellm")`, which does the
   same snapshot/restore around `importlib.import_module`.

3. **Test modules that intentionally need `.env` values for skipif
   guards** (`test_compression_summary_*.py`, `test_query_echo.py`,
   `test_cost_tracker_counterfactual.py`, `test_memory_usage_integration.py`,
   `test_bundled_tools_savings.py`): replace module-level
   `os.environ.setdefault(...)` / `dotenv.load_dotenv()` with
   `tests._dotenv.load_env_overrides()` (returns a local dict — does
   NOT mutate `os.environ`) plus `autouse_apply_env(...)` (function-
   scoped fixture that applies via `monkeypatch.setenv`, auto-cleaned
   at teardown). The skipif still works because
   `ANTHROPIC_KEY = os.environ.get(...) or _env_overrides.get(...)`
   reads from the local dict as fallback.

# Helper module

New `tests/_dotenv.py` exposes:
- `load_env_overrides() -> dict[str, str]` — read `.env` into a dict.
- `autouse_apply_env(overrides) -> fixture` — function-scoped autouse
  fixture that applies via `monkeypatch.setenv`.
- `importorskip_no_env_leak(module) -> module` — drop-in
  `pytest.importorskip` substitute that quarantines env mutations.

# Results

Local full-suite (excluding live-LLM and live-feed tests):
- Before: 46 failed, 4830 passed, 387s
- After:   2 failed, 4672 passed, 134s

The remaining 2 failures are unrelated environment-dependent tests
(missing `PIL` / Docker daemon).
2026-04-26 09:15:37 -07:00
chopratejas
feb163d8c8 Fix Bedrock/Vertex auth regression: don't forward api_key for env-auth providers (#105)
The 0.5.18 refactor added api_key forwarding from request headers to
LiteLLM kwargs in all 4 handler methods. This breaks Bedrock (AWS SigV4)
and Vertex AI (Google ADC) which authenticate via env vars, not API keys.
Forwarding a dummy key like sk-ant-dummy overrides AWS credentials.

Fix: skip api_key forwarding for bedrock, vertex_ai, vertex_ai_beta,
and sagemaker providers. Applied to all 4 occurrences.
2026-04-07 14:24:19 -07:00
chopratejas
7d25c87286 Fix tool_result ordering: no stray user message between tool_calls and tool
Bedrock requires role=tool messages immediately after assistant tool_calls.
The previous fix inserted a user text message in between when the message
contained both text and tool_result blocks, breaking the pairing.

Drop text alongside tool_result (Claude Code never sends it in practice).
Added ordering regression tests for the Bedrock constraint.
2026-03-30 16:25:22 -07:00
chopratejas
8228f0edfb Fix streaming tool calls, compressed request bodies, beacon field names; bump to 0.5.13
- Fix LiteLLM stream_message: emit tool_use blocks from delta.tool_calls,
  set stop_reason from finish_reason (fixes silent MCP tool call failures)
- Fix _convert_messages_for_litellm: convert Anthropic tool_result/tool_use
  to OpenAI role=tool/tool_calls format (fixes 500 on tool round-trips)
- Fix proxy request body parsing: decompress zstd/gzip/deflate/brotli
  Content-Encoding before JSON decode (fixes Codex UnicodeDecodeError crash)
- Fix telemetry beacon field names to match /stats endpoint
  (tokens.total_before_compression, not tokens.original)
- Add dedupe_telemetry.py script for hourly Supabase row deduplication
2026-03-30 14:48:48 -07:00
chopratejas
93d41b66b1 Commit message:
Fix OpenAI streaming with backends and /v1 double-path bug

Add stream_openai_message() to LiteLLM and any-llm backends so
/v1/chat/completions with stream:true returns SSE events instead
of a JSON blob. Clients (Kilo Code, Cursor, etc.) were hanging
because the proxy ignored the stream flag when routing through
a backend.

Also strip trailing /v1 from OPENAI_TARGET_API_URL to prevent
double-path URLs like /v1/v1/models.
2026-03-13 16:49:18 -07:00
chopratejas
4bea17ba8a Fix proxy backend bugs: env vars, tool forwarding, and provider support
- CLI now reads OPENAI_TARGET_API_URL, GEMINI_TARGET_API_URL, and
HEADROOM_ANYLLM_PROVIDER environment variables
- Add --openai-api-url and --gemini-api-url CLI flags
- Remove --backend choices restriction so litellm-* backends work
- Forward tools/tool_choice through LiteLLM and any-llm backends
- Parse tool call arguments from JSON string to dict (Anthropic format)
- Forward top_p, stop_sequences, tools in streaming paths
- Update Vertex AI model map with Claude 3 through 4.6 (from official docs)
- Bump version to 0.4.1
2026-03-12 20:55:26 -07:00