mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
🤖 I have created a release *beep* *boop* --- <details><summary>0.35.0</summary> ## [0.35.0](https://github.com/headroomlabs-ai/headroom/compare/v0.34.0...v0.35.0) (2026-08-12) ### Features * **beacon:** allowlist the routing summary key ([#2818](https://github.com/headroomlabs-ai/headroom/issues/2818)) ([7940c05](7940c05ebf)) * **beacon:** hourly R2 compaction, per-strategy savings, and a stack that reports ([#2853](https://github.com/headroomlabs-ai/headroom/issues/2853)) ([e0870ef](e0870ef931)) * **cli,pricing:** add CLI extension seam and prompt-cache TTL pricing ([#2802](https://github.com/headroomlabs-ai/headroom/issues/2802)) ([6ec3e34](6ec3e3478a)) ### Bug Fixes * **anthropic:** strip first-party tool search on custom upstreams ([#2539](https://github.com/headroomlabs-ai/headroom/issues/2539)) ([7f6950b](7f6950be34)) * **backends/anyllm:** convert Anthropic tools and tool_choice to OpenAI shape ([0d6866b](0d6866b91a)) * **backends/anyllm:** stream tool_use blocks and map finish_reason on the streaming path ([e4904e2](e4904e23a6)) * **backends/litellm:** None-guard core token counts in OpenAI usage block ([#2324](https://github.com/headroomlabs-ai/headroom/issues/2324)) ([12f9f58](12f9f58cb3)) * **beacon:** report all-layers savings, not context-compression only ([#2796](https://github.com/headroomlabs-ai/headroom/issues/2796)) ([e9a24f3](e9a24f3ec1)) * **beacon:** split session failures by status code ([#2815](https://github.com/headroomlabs-ai/headroom/issues/2815)) ([2954e37](2954e37048)) * **cache:** bound compression cache bookkeeping ([0ae948c](0ae948c151)) * **cache:** enforce Anthropic's 1h-before-5m cache_control ordering before forwarding ([#2941](https://github.com/headroomlabs-ai/headroom/issues/2941)) ([3752458](3752458022)) * **cache:** mirror client cache_control positions instead of single-marker consolidation ([def3d76](def3d76e5a)) * **cache:** stabilize Anthropic block-growing lineages ([#2917](https://github.com/headroomlabs-ai/headroom/issues/2917)) ([1a04c95](1a04c957f5)) * **ccr:** avoid injecting tool on chat streaming ([d0c1f5b](d0c1f5b8ad)) * **ccr:** preserve exact SQLite TTL boundary ([#2669](https://github.com/headroomlabs-ai/headroom/issues/2669)) ([d0a86d4](d0a86d409f)) * **ccr:** report embedded hashes from compress endpoint ([#717](https://github.com/headroomlabs-ai/headroom/issues/717)) ([685ebe4](685ebe457d)) * **ccr:** resolve <<ccr:...>> markers inline when no retrieve-tool path exists ([#2512](https://github.com/headroomlabs-ai/headroom/issues/2512)) ([ce8ce83](ce8ce8313f)) * **ccr:** tolerate null/malformed OpenAI data in response handling ([#2467](https://github.com/headroomlabs-ai/headroom/issues/2467)) ([e583e08](e583e082d8)) * **ci:** publish latest from the root Docker manifest ([#2252](https://github.com/headroomlabs-ai/headroom/issues/2252)) ([5568d73](5568d738af)) * **claude:** stop forcing tool search on Foundry ([#2477](https://github.com/headroomlabs-ai/headroom/issues/2477)) ([7981396](798139608c)) * **cli/update:** let install ownership win over bare /.dockerenv so venv installs self-update ([#2830](https://github.com/headroomlabs-ai/headroom/issues/2830)) ([7092b53](7092b53c46)) * **codex:** route alpha search through the Codex backend ([#2538](https://github.com/headroomlabs-ai/headroom/issues/2538)) ([a540eb2](a540eb2c61)) * **content-router:** protect custom-tag blocks before mixed-content section split ([d7bc1e2](d7bc1e275f)) * **deps:** bump h2 to 4.4.1 for CVE-2026-71554 ([#2839](https://github.com/headroomlabs-ai/headroom/issues/2839)) ([564e0a8](564e0a8d0f)) * **deps:** enforce audited transitive dependency floors ([#2791](https://github.com/headroomlabs-ai/headroom/issues/2791)) ([64e2039](64e203931b)) * **doctor:** flag `ollama launch claude` proxy bypass instead of misdirecting ([#2566](https://github.com/headroomlabs-ai/headroom/issues/2566)) ([7f24d69](7f24d695ee)) * emit SSE ping before message_start on Bedrock streaming path (issue [#902](https://github.com/headroomlabs-ai/headroom/issues/902)) ([#1080](https://github.com/headroomlabs-ai/headroom/issues/1080)) ([4dab254](4dab254d52)) * **gemini:** resolve native CCR retrieval calls ([#2253](https://github.com/headroomlabs-ai/headroom/issues/2253)) ([2483f57](2483f57002)) * **health:** label kompress as degraded/optional when not yet loaded ([#2865](https://github.com/headroomlabs-ai/headroom/issues/2865)) ([8949371](89493714d2)) * **image:** decouple routing types from trained_router so importing the compressor doesn't import torch ([#2513](https://github.com/headroomlabs-ai/headroom/issues/2513)) ([#2537](https://github.com/headroomlabs-ai/headroom/issues/2537)) ([d7cf981](d7cf981093)) * **install/windows:** register persistent-task from S4U hidden XML ([#2453](https://github.com/headroomlabs-ai/headroom/issues/2453)) ([#2459](https://github.com/headroomlabs-ai/headroom/issues/2459)) ([1edaeb8](1edaeb8b76)) * **install:** don't crash the PowerShell installer when $PROFILE is unset ([#2469](https://github.com/headroomlabs-ai/headroom/issues/2469)) ([fc5c4e2](fc5c4e239c)) * **install:** trust Docker bridge for dashboard metadata ([e044139](e044139001)) * **install:** use --userns=keep-id under Podman so bind-mount writes don't fail ([#2846](https://github.com/headroomlabs-ai/headroom/issues/2846)) ([3488f8d](3488f8d4b5)) * **learn/gemini:** stop double-counting session tokens ([#2230](https://github.com/headroomlabs-ai/headroom/issues/2230)) ([29d8a5e](29d8a5e563)) * **learn/grok:** detect a Windows absolute project path ([#2283](https://github.com/headroomlabs-ai/headroom/issues/2283)) ([e240df2](e240df2b69)) * **learn:** stop classifying a successful exit code 0 as an error ([#2289](https://github.com/headroomlabs-ai/headroom/issues/2289)) ([a24fe7d](a24fe7dcbf)) * **litellm:** add async_post_call_success_hook to HeadroomCallback ([#1322](https://github.com/headroomlabs-ai/headroom/issues/1322)) ([3107994](3107994aed)) * **litellm:** don't forward a caller key the target cannot accept ([#2883](https://github.com/headroomlabs-ai/headroom/issues/2883)) ([2f2950a](2f2950a626)) * **memory:** bound the TrafficLearner pending-pattern accumulator (memory leak) ([#2579](https://github.com/headroomlabs-ai/headroom/issues/2579)) ([1f5feff](1f5fefffd3)) * **memory:** close DirectMem0 resources ([6596182](65961827cf)) * **memory:** close MCP backend on shutdown ([4bd8ecd](4bd8ecd1e3)) * **memory:** don't crash inline memory extraction on a non-object <memory> block ([#2470](https://github.com/headroomlabs-ai/headroom/issues/2470)) ([e00c6ff](e00c6ff81c)) * **memory:** keep vector metadata in sync ([#2295](https://github.com/headroomlabs-ai/headroom/issues/2295)) ([c471800](c471800e8e)) * **memory:** make explicit-project and user store keys collision-resistant ([#2231](https://github.com/headroomlabs-ai/headroom/issues/2231)) ([f840d5f](f840d5f2fe)) * **memory:** skip <system-reminder> blocks when building the retrieval query ([#2195](https://github.com/headroomlabs-ai/headroom/issues/2195)) ([#2541](https://github.com/headroomlabs-ai/headroom/issues/2541)) ([4e5a67a](4e5a67a342)) * **memory:** sync FTS5 and vector indexes on CLI delete/edit/prune/purge ([fd4628d](fd4628d821)) * **oauth2:** make repository lint checks pass ([c85abf7](c85abf7a87)) * **observability:** aggregate tool savings in OTEL ([#2936](https://github.com/headroomlabs-ai/headroom/issues/2936)) ([941c25d](941c25d31e)) * **onnx:** stop ONNX thread pools from spinning idle cores ([#2495](https://github.com/headroomlabs-ai/headroom/issues/2495)) ([#2540](https://github.com/headroomlabs-ai/headroom/issues/2540)) ([5c561bd](5c561bd913)) * **openai:** skip Responses tool-search deferral for clients that cannot execute it ([#2696](https://github.com/headroomlabs-ai/headroom/issues/2696)) ([54ea28d](54ea28d983)) * **opencode:** ship the transport hook-shim so wheel installs route Node child traffic ([702dbc5](702dbc5902)) * **providers/anthropic:** don't crash token estimation on null tool_calls ([#2472](https://github.com/headroomlabs-ai/headroom/issues/2472)) ([08466f3](08466f3cae)) * **providers/openai:** bound tiktoken vocab loads with the guarded loader ([#2554](https://github.com/headroomlabs-ai/headroom/issues/2554)) ([0805e8e](0805e8e410)) * **proxy/anthropic:** inject headroom_retrieve whenever a CCR marker is present, not only for new markers ([#2848](https://github.com/headroomlabs-ai/headroom/issues/2848)) ([3808f60](3808f60ca6)) * **proxy/anthropic:** None-guard usage token counts on the direct buffered path ([#2434](https://github.com/headroomlabs-ai/headroom/issues/2434)) ([2b5ee7c](2b5ee7cde8)) * **proxy/anthropic:** run tool-search history repair after turn hooks ([c6f9948](c6f99482e1)) * **proxy/batch:** don't crash an OpenAI batch on a valid-JSON non-object line ([#2316](https://github.com/headroomlabs-ai/headroom/issues/2316)) ([1f2c681](1f2c681c0b)) * **proxy/bedrock:** report uncached input tokens from backend usage, not the live-zone count ([#2318](https://github.com/headroomlabs-ai/headroom/issues/2318)) ([c19e412](c19e412b33)) * **proxy/gemini:** keep streaming-parity baseline so eligible_pct can't exceed 100 ([#2824](https://github.com/headroomlabs-ai/headroom/issues/2824)) ([b97c7c6](b97c7c6e99)) * **proxy/metrics:** cap client-supplied model label cardinality ([#2480](https://github.com/headroomlabs-ai/headroom/issues/2480)) ([e24a7e6](e24a7e66b9)) * **proxy/metrics:** escape label values in the Prometheus export ([#2463](https://github.com/headroomlabs-ai/headroom/issues/2463)) ([6a53861](6a53861063)) * **proxy/openai:** don't crash the Responses memory tool loops on null arguments ([#2273](https://github.com/headroomlabs-ai/headroom/issues/2273)) ([a30db2c](a30db2cae4)) * **proxy/openai:** feed Codex WS traffic into the traffic learner ([#2334](https://github.com/headroomlabs-ai/headroom/issues/2334)) ([f669149](f669149769)) * **proxy/openai:** run response hooks on Responses, and bill their re-drives ([#2872](https://github.com/headroomlabs-ai/headroom/issues/2872)) ([675d13f](675d13f08d)) * **proxy:** allow settings routes for trusted gateway/dashboard clients ([#2491](https://github.com/headroomlabs-ai/headroom/issues/2491)) ([a5b0a8f](a5b0a8f4cc)) * **proxy:** cache litellm model resolution to stop repeated Provider List spam ([99f07e7](99f07e7bbd)) * **proxy:** cancel periodic TOIN task on shutdown ([739fdef](739fdef423)) * **proxy:** close the upstream stream when a streaming body is never consumed ([0951663](0951663562)) * **proxy:** compress cache-mode cold starts and tag prefix-mismatch passthrough ([#2365](https://github.com/headroomlabs-ai/headroom/issues/2365)) ([aaeba0a](aaeba0a319)) * **proxy:** emit request log timestamps in UTC ([620028f](620028fa18)) * **proxy:** enable tool search by default and repair poisoned transcripts ([#2807](https://github.com/headroomlabs-ai/headroom/issues/2807)) ([0237cbf](0237cbffbb)) * **proxy:** gate mid-turn message coalescing to Claude Code clients ([#1643](https://github.com/headroomlabs-ai/headroom/issues/1643)) ([a4bd2e6](a4bd2e62a5)) * **proxy:** give each Codex /v1/responses WS turn a unique request_id ([#2164](https://github.com/headroomlabs-ai/headroom/issues/2164)) ([d02df10](d02df10758)) * **proxy:** graceful shutdown and reliable Ctrl+C exit ([#621](https://github.com/headroomlabs-ai/headroom/issues/621)) ([17cdb18](17cdb185bc)) * **proxy:** guard telemetry and TOIN endpoints ([cde1513](cde1513c91)) * **proxy:** include tool_search_deferral savings in the savings ledger ([12149f7](12149f7446)) * **proxy:** pass through cross-region prefixed Bedrock model IDs directly ([#2330](https://github.com/headroomlabs-ai/headroom/issues/2330)) ([64cb46e](64cb46e24b)) * **proxy:** port session-sticky beta headers to the Rust proxy ([#2381](https://github.com/headroomlabs-ai/headroom/issues/2381)) ([f6398a6](f6398a6476)) * **proxy:** preserve merged session and quarantine contracts ([#2943](https://github.com/headroomlabs-ai/headroom/issues/2943)) ([039cd24](039cd2431a)) * **proxy:** preserve signed Anthropic thinking blocks on outbound re-serialize ([#2254](https://github.com/headroomlabs-ai/headroom/issues/2254)) ([dc163bc](dc163bcd1c)) * **proxy:** stop discarding compressed Codex WS later-frame payloads ([#2823](https://github.com/headroomlabs-ai/headroom/issues/2823)) ([4ec416d](4ec416df88)) * **proxy:** time-cap the compression timeout-debt quarantine ([#2360](https://github.com/headroomlabs-ai/headroom/issues/2360)) ([#2412](https://github.com/headroomlabs-ai/headroom/issues/2412)) ([c5a08d2](c5a08d22e0)) * **proxy:** unwrap Hermes tool_call bridge in tool name map ([#2717](https://github.com/headroomlabs-ai/headroom/issues/2717)) ([a97b824](a97b82413b)) * publish headroom-opencode in release workflow ([#2372](https://github.com/headroomlabs-ai/headroom/issues/2372)) ([7859154](78591545ce)) * **settings:** accept documented HEADROOM_* env names as settings keys ([#2833](https://github.com/headroomlabs-ai/headroom/issues/2833)) ([de9e052](de9e0523da)) * **subscription:** dedup transcript usage by message id ([#2340](https://github.com/headroomlabs-ai/headroom/issues/2340) token inflation) ([#2408](https://github.com/headroomlabs-ai/headroom/issues/2408)) ([74275b7](74275b7c3e)) * **toin:** bound private query and pattern retention ([8cd1380](8cd138039e)) * **tokenizer:** coerce non-string tool_call fields before counting ([#2801](https://github.com/headroomlabs-ai/headroom/issues/2801)) ([b6f9877](b6f9877c78)) * **tokenizer:** price CJK in the Rust fixed-ratio estimator (Python parity) ([#2260](https://github.com/headroomlabs-ai/headroom/issues/2260)) ([6840153](6840153473)) * **transforms/adaptive-sizer:** honor max_k on small-input fast path ([#2319](https://github.com/headroomlabs-ai/headroom/issues/2319)) ([8a90523](8a90523209)) * **transforms/smart_crusher:** don't crash on a tool call with a null function ([#2232](https://github.com/headroomlabs-ai/headroom/issues/2232)) ([3bb02f8](3bb02f8f75)) * Vertex model pricing shows $0.00 for versioned model names and vertex:anthropic provider ([#2517](https://github.com/headroomlabs-ai/headroom/issues/2517)) ([eb5b5e4](eb5b5e4198)) * **wrap/claude:** keep --1m effective when an explicit --model is passed through ([c093bf1](c093bf11eb)) * **wrap/opencode:** verify the opencode binary before mutating config ([ae38486](ae384862a4)) * **wrap/serena:** install Serena from the serena-agent PyPI wheel, not the git source ([d7b25ae](d7b25ae3bb)) * **wrap:** honor Copilot OAuth wire-api override and model default ([#2387](https://github.com/headroomlabs-ai/headroom/issues/2387)) ([1db6d88](1db6d88ab4)) * **wrap:** serialize shared proxy startup ([#2946](https://github.com/headroomlabs-ai/headroom/issues/2946)) ([e540d64](e540d64feb)) * **wrap:** stop the launch cwd from shadowing the installed package in the proxy subprocess ([#2843](https://github.com/headroomlabs-ai/headroom/issues/2843)) ([c49be26](c49be269a1)) ### Performance Improvements * cut hot-path latency 27% (token-count memo, startup preloads, JSON scan memo) ([#2838](https://github.com/headroomlabs-ai/headroom/issues/2838)) ([53af90d](53af90d68c)) * **proxy:** bound upstream calls and hot-path costs ([#2852](https://github.com/headroomlabs-ai/headroom/issues/2852)) ([f624d3a](f624d3a00a)) * **subscription:** skip transcripts older than the window in compute_window_tokens ([#2861](https://github.com/headroomlabs-ai/headroom/issues/2861)) ([91d6bf3](91d6bf33cd)) ### Dependencies * bump brace-expansion from 5.0.7 to 5.0.9 in /docs ([#2751](https://github.com/headroomlabs-ai/headroom/issues/2751)) ([56ee57b](56ee57be98)) * bump bytesize from 1.3.3 to 2.4.2 ([#2286](https://github.com/headroomlabs-ai/headroom/issues/2286)) ([6448545](6448545a7f)) * bump hf-hub from 0.4.3 to 0.5.0 ([#2285](https://github.com/headroomlabs-ai/headroom/issues/2285)) ([4925bf6](4925bf6a82)) * bump next from 16.2.10 to 16.3.0 in /docs ([#2750](https://github.com/headroomlabs-ai/headroom/issues/2750)) ([0fd0b99](0fd0b996a4)) * bump postcss from 8.5.19 to 8.5.25 in /plugins/openclaw ([#2749](https://github.com/headroomlabs-ai/headroom/issues/2749)) ([cd60ee9](cd60ee9ae8)) * bump postcss from 8.5.19 to 8.5.25 in /plugins/opencode ([#2748](https://github.com/headroomlabs-ai/headroom/issues/2748)) ([ff4e016](ff4e0167bb)) * bump postcss from 8.5.19 to 8.5.25 in /sdk/typescript ([#2747](https://github.com/headroomlabs-ai/headroom/issues/2747)) ([267c2bd](267c2bdcb5)) * bump postcss from 8.5.19 to 8.5.26 in /docs ([#2881](https://github.com/headroomlabs-ai/headroom/issues/2881)) ([e6e5826](e6e5826423)) * bump ruff from 0.15.17 to 0.15.22 in the pip-minor-patch group ([#2501](https://github.com/headroomlabs-ai/headroom/issues/2501)) ([ecf130d](ecf130d3ac)) * bump rusqlite from 0.32.1 to 0.40.1 ([#2287](https://github.com/headroomlabs-ai/headroom/issues/2287)) ([522faa1](522faa1a59)) * bump the cargo-minor-patch group across 1 directory with 22 updates ([#2916](https://github.com/headroomlabs-ai/headroom/issues/2916)) ([148d860](148d8605e2)) </details> --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please). --------- Co-authored-by: JD Davis <mxjerrett@gmail.com> Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
627 lines
21 KiB
Python
627 lines
21 KiB
Python
"""Tests for the loopback-only /debug/* introspection endpoints (Unit 5)."""
|
|
|
|
from __future__ import annotations
|
|
|
|
import asyncio
|
|
from contextlib import contextmanager
|
|
|
|
import pytest
|
|
|
|
pytest.importorskip("fastapi")
|
|
pytest.importorskip("httpx")
|
|
|
|
from fastapi import HTTPException
|
|
from fastapi.testclient import TestClient
|
|
|
|
from headroom.proxy.debug_introspection import (
|
|
collect_tasks,
|
|
)
|
|
from headroom.proxy.loopback_guard import (
|
|
LOOPBACK_HOSTS,
|
|
is_loopback_host,
|
|
is_loopback_host_header,
|
|
require_loopback,
|
|
)
|
|
from headroom.proxy.server import ProxyConfig, create_app
|
|
from headroom.proxy.warmup import WarmupRegistry
|
|
from headroom.proxy.ws_session_registry import (
|
|
WebSocketSessionRegistry,
|
|
WSSessionHandle,
|
|
)
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Shared fixtures
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
@pytest.fixture
|
|
def client(monkeypatch):
|
|
# Debug endpoint tests must not depend on live upstream network access.
|
|
# Dedicated health-check tests cover both successful and failed upstream
|
|
# probes in tests/test_proxy_healthchecks.py.
|
|
monkeypatch.setenv("HEADROOM_SKIP_UPSTREAM_CHECK", "1")
|
|
config = ProxyConfig(
|
|
optimize=False,
|
|
cache_enabled=False,
|
|
rate_limit_enabled=False,
|
|
cost_tracking_enabled=False,
|
|
)
|
|
app = create_app(config)
|
|
# Pin the simulated client address to loopback so the /debug/* guard
|
|
# accepts the request. Without this, FastAPI's TestClient reports
|
|
# the host as ``testclient`` and the guard correctly 404s us.
|
|
# ``base_url`` pins the inbound ``Host:`` header to a loopback name
|
|
# so the DNS-rebinding gate added in 2026-06 also passes.
|
|
with TestClient(
|
|
app,
|
|
base_url="http://127.0.0.1",
|
|
client=("127.0.0.1", 12345),
|
|
) as test_client:
|
|
yield test_client
|
|
|
|
|
|
@pytest.fixture
|
|
def app_and_client():
|
|
config = ProxyConfig(
|
|
optimize=False,
|
|
cache_enabled=False,
|
|
rate_limit_enabled=False,
|
|
cost_tracking_enabled=False,
|
|
)
|
|
app = create_app(config)
|
|
with TestClient(
|
|
app,
|
|
base_url="http://127.0.0.1",
|
|
client=("127.0.0.1", 12345),
|
|
) as test_client:
|
|
yield app, test_client
|
|
|
|
|
|
@pytest.fixture
|
|
def app_and_external_client():
|
|
"""TestClient that reports a non-loopback address (to exercise 404)."""
|
|
config = ProxyConfig(
|
|
optimize=False,
|
|
cache_enabled=False,
|
|
rate_limit_enabled=False,
|
|
cost_tracking_enabled=False,
|
|
)
|
|
app = create_app(config)
|
|
with TestClient(
|
|
app,
|
|
base_url="http://127.0.0.1",
|
|
client=("10.0.0.1", 54321),
|
|
) as test_client:
|
|
yield app, test_client
|
|
|
|
|
|
@pytest.fixture
|
|
def app_and_rebinding_client():
|
|
"""TestClient that simulates a DNS-rebinding attack.
|
|
|
|
The simulated TCP peer is loopback (``request.client.host`` passes
|
|
the legacy IP check), but the inbound ``Host:`` header reads
|
|
``attacker.com`` — exactly what the browser sends after the
|
|
attacker's DNS record flips to ``127.0.0.1``.
|
|
"""
|
|
config = ProxyConfig(
|
|
optimize=False,
|
|
cache_enabled=False,
|
|
rate_limit_enabled=False,
|
|
cost_tracking_enabled=False,
|
|
)
|
|
app = create_app(config)
|
|
with TestClient(
|
|
app,
|
|
base_url="http://attacker.com",
|
|
client=("127.0.0.1", 12345),
|
|
) as test_client:
|
|
yield app, test_client
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Loopback guard unit tests
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_is_loopback_host_accepts_canonical_hosts():
|
|
for host in LOOPBACK_HOSTS:
|
|
assert is_loopback_host(host) is True
|
|
# None (TestClient with no client info) is treated as loopback.
|
|
assert is_loopback_host(None) is True
|
|
|
|
|
|
def test_is_loopback_host_rejects_external_hosts():
|
|
assert is_loopback_host("10.0.0.1") is False
|
|
assert is_loopback_host("192.168.1.100") is False
|
|
assert is_loopback_host("8.8.8.8") is False
|
|
|
|
|
|
def test_is_loopback_host_accepts_ipv6_mapped_ipv4_loopback():
|
|
# On Linux dual-stack sockets with IPV6_V6ONLY=0, an IPv4 loopback
|
|
# connection arrives as ``::ffff:127.0.0.1``. The guard must treat
|
|
# this as loopback or /debug/* silently 404s when the proxy binds
|
|
# to ``::`` / ``0.0.0.0``.
|
|
assert is_loopback_host("::ffff:127.0.0.1") is True
|
|
|
|
|
|
def test_is_loopback_host_rejects_ipv6_mapped_external_ipv4():
|
|
assert is_loopback_host("::ffff:10.0.0.1") is False
|
|
|
|
|
|
def test_is_loopback_host_rejects_non_loopback_ipv6():
|
|
assert is_loopback_host("2001:db8::1") is False
|
|
|
|
|
|
def test_is_loopback_host_rejects_malformed_input():
|
|
assert is_loopback_host("not-an-ip") is False
|
|
assert is_loopback_host("") is False
|
|
|
|
|
|
def test_require_loopback_raises_404_for_external_client():
|
|
class _FakeClient:
|
|
host = "10.0.0.1"
|
|
|
|
class _FakeRequest:
|
|
client = _FakeClient()
|
|
|
|
with pytest.raises(HTTPException) as exc_info:
|
|
require_loopback(_FakeRequest()) # type: ignore[arg-type]
|
|
|
|
assert exc_info.value.status_code == 404
|
|
# Privacy: 404 explicitly, not 403 — endpoints should be invisible.
|
|
assert exc_info.value.status_code != 403
|
|
|
|
|
|
def test_require_loopback_accepts_loopback_client():
|
|
class _FakeClient:
|
|
host = "127.0.0.1"
|
|
|
|
class _FakeRequest:
|
|
client = _FakeClient()
|
|
|
|
# Should not raise. ``headers`` is absent so the Host-header gate
|
|
# falls back to the legacy IP-only behaviour for callers that
|
|
# construct a bare request stub.
|
|
require_loopback(_FakeRequest()) # type: ignore[arg-type]
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Host-header (DNS-rebinding) guard unit tests
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_is_loopback_host_header_accepts_canonical_values():
|
|
for value in (
|
|
"127.0.0.1",
|
|
"127.0.0.1:8787",
|
|
"localhost",
|
|
"localhost:8787",
|
|
"LOCALHOST",
|
|
"Localhost:8787",
|
|
"[::1]",
|
|
"[::1]:8787",
|
|
):
|
|
assert is_loopback_host_header(value) is True, value
|
|
|
|
|
|
def test_is_loopback_host_header_rejects_external_names():
|
|
for value in (
|
|
"attacker.com",
|
|
"attacker.com:8787",
|
|
"evil.example",
|
|
"10.0.0.1",
|
|
"10.0.0.1:8787",
|
|
"8.8.8.8",
|
|
):
|
|
assert is_loopback_host_header(value) is False, value
|
|
|
|
|
|
def test_is_loopback_host_header_rejects_missing_and_malformed():
|
|
assert is_loopback_host_header(None) is False
|
|
assert is_loopback_host_header("") is False
|
|
assert is_loopback_host_header(" ") is False
|
|
# Unterminated bracketed IPv6
|
|
assert is_loopback_host_header("[::1") is False
|
|
# Hostname that merely contains a loopback substring
|
|
assert is_loopback_host_header("localhost.attacker.com") is False
|
|
|
|
|
|
def test_require_loopback_blocks_dns_rebinding_host_header():
|
|
"""Loopback IP + ``Host: attacker.com`` is the rebinding signature."""
|
|
|
|
class _FakeClient:
|
|
host = "127.0.0.1"
|
|
|
|
class _FakeHeaders:
|
|
def get(self, key, default=None):
|
|
if key.lower() == "host":
|
|
return "attacker.com"
|
|
return default
|
|
|
|
class _FakeRequest:
|
|
client = _FakeClient()
|
|
headers = _FakeHeaders()
|
|
|
|
with pytest.raises(HTTPException) as exc_info:
|
|
require_loopback(_FakeRequest()) # type: ignore[arg-type]
|
|
assert exc_info.value.status_code == 404
|
|
|
|
|
|
def test_require_loopback_accepts_loopback_host_header():
|
|
class _FakeClient:
|
|
host = "127.0.0.1"
|
|
|
|
class _FakeHeaders:
|
|
def get(self, key, default=None):
|
|
if key.lower() == "host":
|
|
return "127.0.0.1:8787"
|
|
return default
|
|
|
|
class _FakeRequest:
|
|
client = _FakeClient()
|
|
headers = _FakeHeaders()
|
|
|
|
# Should not raise — both gates pass.
|
|
require_loopback(_FakeRequest()) # type: ignore[arg-type]
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Serializer unit tests
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_warmup_registry_to_dict_returns_registry_shape():
|
|
"""Serializer equivalent of the old collect_warmup helper.
|
|
|
|
The helper was inlined at the /debug/warmup route handler in server.py
|
|
(``registry.to_dict() if registry else {}``); this test preserves
|
|
coverage of the registry's own serializer contract.
|
|
"""
|
|
registry = WarmupRegistry()
|
|
registry.kompress.mark_loaded(handle=object(), source_status="enabled")
|
|
registry.memory_backend.mark_error("boom")
|
|
|
|
payload = registry.to_dict()
|
|
|
|
assert payload["kompress"]["status"] == "loaded"
|
|
assert payload["memory_backend"]["status"] == "error"
|
|
assert payload["memory_backend"]["error"] == "boom"
|
|
# Raw handle must never leak into the serialized payload.
|
|
assert "handle" not in payload["kompress"]
|
|
|
|
|
|
def test_ws_session_registry_snapshot_returns_registered_entries():
|
|
"""Serializer equivalent of the old collect_ws_sessions helper."""
|
|
reg = WebSocketSessionRegistry()
|
|
handle = WSSessionHandle(
|
|
session_id="sess-debug-1",
|
|
request_id="req-debug-1",
|
|
client_addr="127.0.0.1:9999",
|
|
upstream_url="wss://upstream/test",
|
|
)
|
|
reg.register(handle)
|
|
|
|
payload = reg.snapshot()
|
|
assert len(payload) == 1
|
|
assert payload[0]["session_id"] == "sess-debug-1"
|
|
assert payload[0]["request_id"] == "req-debug-1"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_collect_tasks_returns_current_tasks_with_metadata():
|
|
async def _noop_task():
|
|
await asyncio.sleep(0.05)
|
|
|
|
task = asyncio.create_task(_noop_task(), name="debug-test-task")
|
|
try:
|
|
entries = collect_tasks()
|
|
matching = [e for e in entries if e["name"] == "debug-test-task"]
|
|
assert matching, "expected the named task to appear in collect_tasks output"
|
|
entry = matching[0]
|
|
assert entry["coro_qualname"] is not None
|
|
# Privacy: no frame locals, no coroutine args.
|
|
assert "locals" not in entry
|
|
assert "cr_frame" not in entry
|
|
assert "args" not in entry
|
|
assert entry["stack_depth"] is None or isinstance(entry["stack_depth"], int)
|
|
finally:
|
|
task.cancel()
|
|
try:
|
|
await task
|
|
except (asyncio.CancelledError, BaseException):
|
|
pass
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_collect_tasks_derives_age_from_ws_registry_for_codex_relays():
|
|
reg = WebSocketSessionRegistry()
|
|
sid = "relay-sess-1"
|
|
reg.register(
|
|
WSSessionHandle(
|
|
session_id=sid,
|
|
request_id="req-relay-1",
|
|
client_addr="127.0.0.1:1",
|
|
upstream_url="wss://upstream",
|
|
)
|
|
)
|
|
|
|
async def _long_relay():
|
|
await asyncio.sleep(0.2)
|
|
|
|
relay_task = asyncio.create_task(_long_relay(), name=f"codex-ws-c2u-{sid}")
|
|
try:
|
|
await asyncio.sleep(0.02) # let some age accrue
|
|
entries = collect_tasks(ws_registry=reg)
|
|
named = [e for e in entries if e["name"] == f"codex-ws-c2u-{sid}"]
|
|
assert named, "expected relay task in output"
|
|
entry = named[0]
|
|
assert entry["age_seconds"] is not None
|
|
assert entry["age_seconds"] >= 0.0
|
|
finally:
|
|
relay_task.cancel()
|
|
try:
|
|
await relay_task
|
|
except (asyncio.CancelledError, BaseException):
|
|
pass
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# HTTP endpoint tests (loopback)
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_debug_tasks_returns_json_array_for_loopback(client):
|
|
response = client.get("/debug/tasks")
|
|
assert response.status_code == 200
|
|
data = response.json()
|
|
assert isinstance(data, list)
|
|
# Each entry at least has name + coro_qualname fields.
|
|
for entry in data:
|
|
assert "name" in entry
|
|
assert "coro_qualname" in entry
|
|
|
|
|
|
def test_debug_tasks_stack_depth_is_gated_behind_query(client):
|
|
"""Default response must not compute stack_depth (P3 Fix 29 perf gate).
|
|
|
|
``?stack=true`` opts into the synchronous ``Task.get_stack`` walk; the
|
|
default stays cheap so snapshotting during a reconnect storm does
|
|
not stall the event loop.
|
|
"""
|
|
default = client.get("/debug/tasks")
|
|
assert default.status_code == 200
|
|
for entry in default.json():
|
|
assert entry["stack_depth"] is None, (
|
|
f"default /debug/tasks must not compute stack_depth; "
|
|
f"got {entry['stack_depth']!r} for {entry.get('name')!r}"
|
|
)
|
|
|
|
with_stack = client.get("/debug/tasks?stack=true")
|
|
assert with_stack.status_code == 200
|
|
entries = with_stack.json()
|
|
# At least one entry should have a computed depth (the TestClient
|
|
# itself runs under a task). Some entries may still be None if
|
|
# get_stack raised defensively — we only require that opting in
|
|
# produces at least one integer result.
|
|
integer_depths = [e["stack_depth"] for e in entries if isinstance(e["stack_depth"], int)]
|
|
assert integer_depths, (
|
|
f"expected at least one int stack_depth when ?stack=true; got entries={entries!r}"
|
|
)
|
|
|
|
|
|
def test_debug_warmup_reports_registry_slots(client):
|
|
response = client.get("/debug/warmup")
|
|
assert response.status_code == 200
|
|
data = response.json()
|
|
# Registry surfaces all canonical slot names.
|
|
assert "kompress" in data
|
|
assert "magika" in data
|
|
assert "memory_backend" in data
|
|
assert "memory_embedder" in data
|
|
assert "runtime" in data
|
|
# Each slot has at least a status field.
|
|
assert "status" in data["memory_backend"]
|
|
assert data["runtime"]["anthropic_pre_upstream"]["resolved_concurrency"] >= 0
|
|
assert data["runtime"]["websocket_sessions"]["active_relay_tasks"] == 0
|
|
|
|
|
|
class _KompressStub:
|
|
"""Read-only stand-in exposing the accessors the health reconciler uses.
|
|
|
|
``preload`` / ``ensure_background_load`` raise so the tests fail loudly if
|
|
``/debug/warmup`` ever triggers a model load instead of just observing.
|
|
"""
|
|
|
|
def __init__(self, *, backend="onnx", ready=True):
|
|
self.backend = backend
|
|
self.ready = ready
|
|
self.calls: list[str] = []
|
|
|
|
def is_ready(self):
|
|
self.calls.append("is_ready")
|
|
return self.ready
|
|
|
|
def ready_backend(self):
|
|
self.calls.append("ready_backend")
|
|
return self.backend
|
|
|
|
def preload(self):
|
|
raise AssertionError("/debug/warmup must never preload kompress")
|
|
|
|
def ensure_background_load(self):
|
|
raise AssertionError("/debug/warmup must never start a background load")
|
|
|
|
|
|
@contextmanager
|
|
def _deferred_kompress_client(compressor):
|
|
"""Client whose kompress slot still carries the startup ``deferred`` mark.
|
|
|
|
Mirrors the real cold-start shape: ``eager_load_compressors`` reported
|
|
``deferred``, the model then loaded on the request path, and nothing wrote
|
|
the promotion back to the registry.
|
|
"""
|
|
config = ProxyConfig(
|
|
optimize=False,
|
|
cache_enabled=False,
|
|
rate_limit_enabled=False,
|
|
cost_tracking_enabled=False,
|
|
)
|
|
app = create_app(config)
|
|
with TestClient(
|
|
app,
|
|
base_url="http://127.0.0.1",
|
|
client=("127.0.0.1", 12345),
|
|
) as test_client:
|
|
proxy = app.state.proxy
|
|
router = proxy.anthropic_pipeline.transforms[-1]
|
|
router._kompress = compressor
|
|
proxy.warmup.kompress.mark_null()
|
|
proxy.warmup.kompress.info["source_status"] = "deferred"
|
|
yield proxy, test_client
|
|
|
|
|
|
def _clear_kompress_cache(monkeypatch):
|
|
"""Neutralize the process-global ONNX cache the reconciler falls back to."""
|
|
try:
|
|
from headroom.transforms import kompress_compressor
|
|
except ImportError:
|
|
return
|
|
monkeypatch.setattr(kompress_compressor, "_kompress_cache", {}, raising=False)
|
|
|
|
|
|
def test_debug_warmup_promotes_deferred_kompress_after_runtime_load():
|
|
compressor = _KompressStub()
|
|
with _deferred_kompress_client(compressor) as (_proxy, client):
|
|
slot = client.get("/debug/warmup").json()["kompress"]
|
|
|
|
assert slot["status"] == "loaded"
|
|
assert slot["info"]["backend"] == "onnx"
|
|
assert slot["info"]["source_status"] == "runtime"
|
|
|
|
|
|
def test_debug_warmup_keeps_pending_kompress_null(monkeypatch):
|
|
_clear_kompress_cache(monkeypatch)
|
|
compressor = _KompressStub(ready=False)
|
|
with _deferred_kompress_client(compressor) as (_proxy, client):
|
|
slot = client.get("/debug/warmup").json()["kompress"]
|
|
|
|
assert slot["status"] == "null"
|
|
assert slot["info"]["source_status"] == "deferred"
|
|
assert compressor.calls == ["is_ready"]
|
|
|
|
|
|
def test_debug_warmup_never_starts_kompress_loading():
|
|
compressor = _KompressStub()
|
|
with _deferred_kompress_client(compressor) as (_proxy, client):
|
|
client.get("/debug/warmup")
|
|
|
|
# Observation only: no preload(), no ensure_background_load(), no compress().
|
|
assert compressor.calls == ["is_ready", "ready_backend"]
|
|
|
|
|
|
def test_debug_ws_sessions_reports_live_session(app_and_client):
|
|
app, client = app_and_client
|
|
proxy = app.state.proxy
|
|
assert proxy is not None, "create_app must wire app.state.proxy"
|
|
|
|
sid = "sess-debug-http"
|
|
proxy.ws_sessions.register(
|
|
WSSessionHandle(
|
|
session_id=sid,
|
|
request_id="req-debug-http",
|
|
client_addr="127.0.0.1:12345",
|
|
upstream_url="wss://upstream/test",
|
|
)
|
|
)
|
|
try:
|
|
response = client.get("/debug/ws-sessions")
|
|
assert response.status_code == 200
|
|
data = response.json()
|
|
matching = [entry for entry in data if entry["session_id"] == sid]
|
|
assert matching, "expected live session in /debug/ws-sessions output"
|
|
assert matching[0]["request_id"] == "req-debug-http"
|
|
finally:
|
|
proxy.ws_sessions.deregister(sid, cause="response_completed")
|
|
|
|
# After cleanup the session is gone.
|
|
response = client.get("/debug/ws-sessions")
|
|
assert response.status_code == 200
|
|
assert all(entry["session_id"] != sid for entry in response.json())
|
|
|
|
|
|
def test_debug_endpoints_do_not_mutate_state(client):
|
|
# Call each endpoint 100 times and confirm the second read equals
|
|
# the first — no accidental mutation from serialization.
|
|
first_tasks = client.get("/debug/tasks").json()
|
|
first_warmup = client.get("/debug/warmup").json()
|
|
first_ws = client.get("/debug/ws-sessions").json()
|
|
|
|
for _ in range(100):
|
|
client.get("/debug/tasks")
|
|
client.get("/debug/warmup")
|
|
client.get("/debug/ws-sessions")
|
|
|
|
# Warmup and ws-sessions are deterministic (no background work touches
|
|
# them in this test config), so they must be identical.
|
|
assert client.get("/debug/warmup").json() == first_warmup
|
|
assert client.get("/debug/ws-sessions").json() == first_ws
|
|
# Tasks may vary naturally, but the call itself never raises and the
|
|
# shape never changes.
|
|
new_tasks = client.get("/debug/tasks").json()
|
|
assert isinstance(new_tasks, list)
|
|
for entry in new_tasks:
|
|
assert set(entry.keys()) == set(first_tasks[0].keys()) if first_tasks else True
|
|
|
|
|
|
def test_debug_tasks_does_not_leak_coro_locals(client):
|
|
response = client.get("/debug/tasks")
|
|
assert response.status_code == 200
|
|
for entry in response.json():
|
|
# Privacy check: the serializer must not leak coroutine locals,
|
|
# frame state, or request bodies. Only name / qualname / age /
|
|
# depth / done are allowed.
|
|
assert set(entry.keys()) <= {
|
|
"name",
|
|
"coro_qualname",
|
|
"age_seconds",
|
|
"stack_depth",
|
|
"done",
|
|
}
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# HTTP endpoint tests (non-loopback)
|
|
# ---------------------------------------------------------------------------
|
|
|
|
|
|
def test_debug_endpoints_return_404_for_non_loopback_client(app_and_external_client):
|
|
_, client = app_and_external_client
|
|
for path in ("/debug/tasks", "/debug/ws-sessions", "/debug/warmup"):
|
|
response = client.get(path)
|
|
assert response.status_code == 404, path
|
|
# Must be 404, not 403 — invisible to scanners.
|
|
assert response.status_code != 403
|
|
|
|
|
|
def test_debug_endpoints_block_dns_rebinding(app_and_rebinding_client):
|
|
"""Loopback client + ``Host: attacker.com`` must 404 like an external client.
|
|
|
|
Regression for the DNS-rebinding gap: prior to 2026-06 the guard
|
|
only checked ``request.client.host``, which a rebound browser
|
|
passes trivially. Adding a ``Host:`` header allowlist closes that
|
|
gap so a malicious site cannot read /debug/* over the user's
|
|
loopback proxy via the wide-open CORS policy.
|
|
"""
|
|
_, client = app_and_rebinding_client
|
|
for path in ("/debug/tasks", "/debug/ws-sessions", "/debug/warmup"):
|
|
response = client.get(path)
|
|
assert response.status_code == 404, path
|
|
assert response.status_code != 403
|
|
|
|
|
|
def test_existing_health_routes_unchanged(client):
|
|
# Invariant: Unit 5 must not regress the existing health endpoints.
|
|
for path in ("/livez", "/readyz", "/health"):
|
|
response = client.get(path)
|
|
assert response.status_code == 200, path
|