mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
## Description Makes the compression-only `POST /v1/compress` endpoint usable as a **network compression sidecar** behind an API gateway (Kong, LiteLLM, ...), and fixes a latent content-detector hang that silently zeroed compression on non-Windows hosts. Motivated by a LiteLLM-sidecar deployment whose team documented five build-time patches; this ports the ones that belong upstream, generalized so they cover any aliasing gateway (not just LiteLLM). Closes # ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [x] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **`lossy_inline` compress mode** (`config.mode="lossy_inline"`, alias `"lossless_then_lossy"`): lossless byte/data fold first, then Kompress the folded remainder, with `ccr_inject_marker=False` so every compressor emits **inline, marker-free** output — no `<<ccr:…>>` markers and no CCR store write, so the result is safe to forward straight to a provider with no retrieval round-trip. The mode inherits the deployment's `enable_kompress`. - **`HEADROOM_COMPRESS_ALLOW_REMOTE`** opt-in: drops the loopback dependency on the `/v1/compress` route **only** so an authorized in-network gateway can reach it. Default is unchanged (loopback-only); inbound `HEADROOM_PROXY_TOKEN` auth still applies. - **`HEADROOM_MODEL_ALIAS_MAP`** (gateway-agnostic, fail-soft): one shared resolver in `pricing/litellm_pricing.py` reduces a gateway-aliased model name (e.g. `claude-opus`) to a priced `litellm.model_cost` key, trying the mapped target as-is and with a `bedrock/` / `vertex_ai/` prefix stripped. `proxy/savings_tracker.py` now delegates to it, so the live (`/stats`) and persisted (`/stats-history`) dollar figures price identically. - **`get_context_limit`**: an operator-configured limit (`HEADROOM_MODEL_LIMITS` / `~/.headroom/models.json`) now wins **before** the dynamic LiteLLM lookup, so an aliased name no longer falls through to the 128K default and skews compression. - **fix(content_router): first-call detector watchdog on all platforms.** The native content detector can deadlock on first use (#575, previously flagged Windows-only). The watchdog was `win32`-only, so on macOS/Linux a first-use hang was unbounded → `_detect_content` never returned → the `/v1/compress` executor timeout fired → fail-open → **`tokens_before=0`, silent zero compression**. Now the native detector runs under the watchdog on the first call on every platform; once it returns it is marked verified and the direct fast path is used (zero steady-state overhead). A hang degrades to pure-Python detection with a clear warning. `win32` behavior is unchanged. - Thread `waste_signals` / `pipeline_timing` into the already-present `/v1/compress` outcome record so the guardrail path populates the dashboard panels like the forward-proxy paths. Deliberately **not** ported: the sidecar's LiteLLM-specific `GET /model/info` HTTP fetch (urllib/ssl/threading/TTL). Kong has no such endpoint; the static `HEADROOM_MODEL_ALIAS_MAP` covers any gateway with no network dependency on the pricing path. ## Testing - [x] Unit tests pass (targeted — see output) - [x] Linting passes (`ruff check`) - [x] Type checking passes (`mypy`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ ruff check <changed files> All checks passed! $ mypy <changed source files> Success: no issues found in 6 source files $ pytest tests/test_gateway_sidecar_ports.py tests/test_proxy_compress_endpoint.py -q tests/test_gateway_sidecar_ports.py ........ [ 34%] tests/test_proxy_compress_endpoint.py ............... [100%] ============================= 23 passed in 20.62s ============================== ``` ## Real Behavior Proof - **Environment:** macOS (darwin/arm64), Python 3.12, `.venv`; Kompress offloaded to a Modal endpoint via `HEADROOM_KOMPRESS_ENDPOINT`. - **Exact command / steps:** posted typical tool-output payloads to `POST /v1/compress` (via the FastAPI `TestClient`, loopback) in both `default` and `lossy_inline` modes; separately reproduced the detector hang with `faulthandler.dump_traceback_later`. - **Observed result:** - Real savings through the endpoint (structural/lossless, Kompress off): **JSON 150 records 13,982→9,514 (32.0%)**, **logs 314 lines 12,240→9,549 (22.0%)**, **search 200 hits 5,231→3,471 (33.6%)**. `lossy_inline` emits **zero** CCR markers. - `faulthandler` pinned the pre-fix hang to `content_router.py:_detect_content` → native `_rust_detect`. With the fix, the first call degrades at the 5s watchdog with `"Native content detector hung … using pure-Python detection"` and compression proceeds (previously it hung and the endpoint returned `tokens_before=0`). - Modal Kompress warm latency measured ~0.8s/call; the learned pass compresses prose further (62→56 words on a sample). - **Not tested:** full `pytest` suite (ran the two affected test files only); the native-detector hang was reproduced on a local macOS/arm64 build — the fix's degrade path is verified, but a healthy-native CI Linux run should confirm the fast (verified) path there. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I did **not** edit `CHANGELOG.md` ## Additional Notes - The five-item context comes from a downstream LiteLLM sidecar's `PATCHES.md`; item #3 (record an outcome from the guardrail path) was already upstreamed — this PR only adds the missing `waste_signals`/`pipeline_timing` threading. Item #2 (observability read-only exemption when `HEADROOM_PROXY_TOKEN` is set) is not addressed here. - All new config is opt-in and fail-soft; with nothing set, behavior is byte-identical to today.
134 lines
4.6 KiB
Python
134 lines
4.6 KiB
Python
"""Ports of the LiteLLM/Kong sidecar patches (see the sidecar PATCHES.md #1/#4/#5/#6).
|
|
|
|
- #1 HEADROOM_COMPRESS_ALLOW_REMOTE opt-in drops the loopback guard on
|
|
/v1/compress so an authorized in-network gateway (Kong, LiteLLM) can reach it.
|
|
- #4/#5 HEADROOM_MODEL_ALIAS_MAP reduces a gateway-aliased model name (e.g.
|
|
"claude-opus") to a priced litellm.model_cost key — one shared resolver for
|
|
the live (cost.py) and persisted (savings_tracker) price paths.
|
|
- #6 an operator-configured context limit (HEADROOM_MODEL_LIMITS) wins over the
|
|
128K default for an aliased name.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import json
|
|
|
|
import pytest
|
|
|
|
pytest.importorskip("fastapi")
|
|
from fastapi.testclient import TestClient
|
|
|
|
from headroom.proxy.server import ProxyConfig, create_app
|
|
|
|
|
|
def _clear_pricing_cache() -> None:
|
|
from headroom.pricing import litellm_pricing as lp
|
|
|
|
lp._resolved_model_cache.clear()
|
|
|
|
|
|
def _first_priced_opus_key() -> str:
|
|
litellm = pytest.importorskip("litellm")
|
|
for key, val in litellm.model_cost.items():
|
|
if "opus" in key.lower() and val.get("input_cost_per_token") is not None:
|
|
return key
|
|
pytest.skip("no priced opus key in this litellm build")
|
|
|
|
|
|
# ----- #4/#5 pricing: HEADROOM_MODEL_ALIAS_MAP -> priced key -----
|
|
|
|
|
|
def test_alias_map_resolves_gateway_name_to_priced_key(monkeypatch):
|
|
litellm = pytest.importorskip("litellm")
|
|
from headroom.pricing.litellm_pricing import resolve_litellm_model
|
|
|
|
key = _first_priced_opus_key()
|
|
monkeypatch.setenv("HEADROOM_MODEL_ALIAS_MAP", json.dumps({"claude-opus": key}))
|
|
_clear_pricing_cache()
|
|
|
|
resolved = resolve_litellm_model("claude-opus")
|
|
info = litellm.model_cost.get(resolved)
|
|
assert info and info.get("input_cost_per_token") is not None
|
|
|
|
|
|
def test_alias_map_strips_bedrock_prefix(monkeypatch):
|
|
pytest.importorskip("litellm")
|
|
from headroom.pricing.litellm_pricing import resolve_litellm_model
|
|
|
|
key = _first_priced_opus_key()
|
|
monkeypatch.setenv("HEADROOM_MODEL_ALIAS_MAP", json.dumps({"claude-opus": f"bedrock/{key}"}))
|
|
_clear_pricing_cache()
|
|
assert resolve_litellm_model("claude-opus") == key
|
|
|
|
|
|
def test_unpriced_alias_falls_through_soft(monkeypatch):
|
|
from headroom.pricing.litellm_pricing import resolve_litellm_model
|
|
|
|
monkeypatch.setenv("HEADROOM_MODEL_ALIAS_MAP", json.dumps({"claude-opus": "not-a-real-model"}))
|
|
_clear_pricing_cache()
|
|
# No crash; falls back to bare-prefix resolution (never returns the bogus target).
|
|
assert resolve_litellm_model("claude-opus") != "not-a-real-model"
|
|
|
|
|
|
def test_unset_env_is_unchanged(monkeypatch):
|
|
from headroom.pricing.litellm_pricing import resolve_litellm_model
|
|
|
|
monkeypatch.delenv("HEADROOM_MODEL_ALIAS_MAP", raising=False)
|
|
_clear_pricing_cache()
|
|
assert isinstance(resolve_litellm_model("gpt-4o"), str)
|
|
|
|
|
|
def test_savings_tracker_delegates_to_shared_resolver(monkeypatch):
|
|
pytest.importorskip("litellm")
|
|
from headroom.proxy.savings_tracker import _resolve_litellm_model
|
|
|
|
key = _first_priced_opus_key()
|
|
monkeypatch.setenv("HEADROOM_MODEL_ALIAS_MAP", json.dumps({"claude-opus": key}))
|
|
_clear_pricing_cache()
|
|
# Persisted funnel prices the alias identically to the live path.
|
|
assert _resolve_litellm_model("claude-opus") == key
|
|
|
|
|
|
# ----- #6 context limit: configured alias wins over the 128K default -----
|
|
|
|
|
|
def test_configured_context_limit_wins_over_default(monkeypatch):
|
|
monkeypatch.setenv(
|
|
"HEADROOM_MODEL_LIMITS",
|
|
json.dumps({"openai": {"context_limits": {"claude-opus": 200000}}}),
|
|
)
|
|
from headroom.providers.openai import OpenAIProvider
|
|
|
|
provider = OpenAIProvider()
|
|
assert provider.get_context_limit("claude-opus") == 200000 # not the 128000 default
|
|
|
|
|
|
# ----- #1 loopback opt-in on /v1/compress -----
|
|
|
|
|
|
def _fast_app():
|
|
return create_app(
|
|
ProxyConfig(
|
|
optimize=False,
|
|
cache_enabled=False,
|
|
rate_limit_enabled=False,
|
|
cost_tracking_enabled=False,
|
|
)
|
|
)
|
|
|
|
|
|
_BODY = {"messages": [{"role": "user", "content": "hi"}], "model": "gpt-4"}
|
|
|
|
|
|
def test_compress_blocks_non_loopback_by_default(monkeypatch):
|
|
monkeypatch.delenv("HEADROOM_COMPRESS_ALLOW_REMOTE", raising=False)
|
|
# A vanilla TestClient presents client.host="testclient" (non-loopback).
|
|
client = TestClient(_fast_app())
|
|
assert client.post("/v1/compress", json=_BODY).status_code == 404
|
|
|
|
|
|
def test_compress_allows_non_loopback_with_flag(monkeypatch):
|
|
monkeypatch.setenv("HEADROOM_COMPRESS_ALLOW_REMOTE", "1")
|
|
client = TestClient(_fast_app())
|
|
resp = client.post("/v1/compress", json=_BODY)
|
|
assert resp.status_code == 200, resp.text
|