mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
fix(proxy/bedrock): wire PrefixCacheTracker updates into Bedrock backend paths (#2196)
## Description `update_from_response()` was only called from the direct-Anthropic-API branch of `handle_anthropic_messages`. Both Bedrock backend branches (streaming and non-streaming) returned before ever reaching it, so `PrefixCacheTracker` state stayed permanently empty for the life of a session on any `--backend bedrock` deployment: `extract_cache_stable_delta()` always saw no previous turn, and `--mode cache` fell back to full unmodified passthrough on every turn instead of compressing the append-only delta. Closes # ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `headroom/proxy/handlers/anthropic.py`: non-streaming Bedrock branch now mirrors the direct-API branch — builds `next_original_messages`/`next_forwarded_messages` from the response, runs cache-miss attribution, and calls `prefix_tracker.update_from_response()` before returning. - `headroom/proxy/handlers/streaming.py`: `_stream_response_bedrock` gains `prefix_tracker`/`optimized_messages` parameters (previously absent entirely), accumulates raw SSE bytes only when a tracker is present, reconstructs the assistant message via the existing `_parse_sse_to_response` helper in the `finally:` block, then updates the tracker. Mirrors `_finalize_stream_response` and the OpenAI-via-backend sibling (`_stream_openai_via_backend`), which already had this wiring. - `tests/test_bedrock_prefix_tracker_wiring.py` (new): drives real `PrefixCacheTracker` instances (via `session_tracker_store`, not a fake) through both the non-streaming and streaming Bedrock paths using `TestClient`, and asserts the tracker's turn counter and last-forwarded/-original messages actually advance after a Bedrock call. A second non-streaming test drives two turns and asserts turn 2 sees a nonzero `frozen_message_count` once the cached total clears `min_cached_tokens`. Verified these tests fail against the pre-fix `anthropic.py`/`streaming.py` (turn counter stuck at 0) and pass against the fix. - `CHANGELOG.md`: added a `### Fixed` entry under `Unreleased`. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ uv run pytest tests/test_bedrock_prefix_tracker_wiring.py tests/test_backend_nonstreaming_cache_metrics.py tests/test_backend_streaming_cache_metrics.py tests/test_bedrock_streaming_input_tokens.py tests/test_cache/test_prefix_tracker.py tests/test_cache_prefix_overlay.py tests/test_cross_turn_cache_safety.py tests/test_proxy_anthropic_cache_stability.py -q collected 91 items tests/test_bedrock_prefix_tracker_wiring.py ... [ 3%] tests/test_backend_nonstreaming_cache_metrics.py .... [ 7%] tests/test_backend_streaming_cache_metrics.py .... [ 12%] tests/test_bedrock_streaming_input_tokens.py .. [ 14%] tests/test_cache/test_prefix_tracker.py .................................. [ 49%] tests/test_cache_prefix_overlay.py ......... [ 69%] tests/test_cross_turn_cache_safety.py ... [ 72%] tests/test_proxy_anthropic_cache_stability.py ......................... [100%] ======================== 91 passed, 1 warning in 9.15s ========================= $ uv run ruff check headroom/proxy/handlers/anthropic.py headroom/proxy/handlers/streaming.py tests/test_bedrock_prefix_tracker_wiring.py All checks passed! $ uv run mypy headroom/proxy/handlers/anthropic.py headroom/proxy/handlers/streaming.py Success: no issues found in 2 source files ``` ## Real Behavior Proof - Environment: personal fork deployed as a real proxy (macOS launchd service, `headroom install apply`) with `--backend bedrock --mode cache`, fronting a live Claude Code session. - Exact command / steps: ran a two-turn streaming conversation against the running Bedrock-backed proxy, then a third append-only turn, while temporarily adding debug logging around `prefix_tracker.get_frozen_message_count()` / `get_last_original_messages()` (removed before this commit; the automated tests above are the permanent record). - Observed result: before the fix, `prev_orig_len`/`prev_fwd_len` were always 0 on every turn including turn 2+ — the tracker never advanced past its cold-start state. After the fix, turn 2 shows `prev_orig_len`/`prev_fwd_len` populated from turn 1's response, and the append-only turn 3 correctly triggers the delta-compression path (`router:noop` transform, pipeline actually runs) instead of falling to the router-never-called passthrough. In a separate live session captured while validating this fix, one turn showed `cache_write=98242` in the PERF log, and the immediately following turn showed `cache_read=98242 cache_hit_pct=94` — direct proof that the Bedrock path is now feeding real cache-read/write data back into the tracker end-to-end on live traffic, not just synthetic test fixtures. - Not tested: the live full-suite run during development surfaced one pre-existing unrelated failure in `test_provider_model_fallback.py`, confirmed independently failing on the commit prior to this fix (i.e., not introduced by this change, not fixed by it either). ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Screenshots (if applicable) N/A — backend logic change, no UI surface. ## Additional Notes - No linked issue number: found via independent investigation of a personal deployment, not filed as a `headroomlabs-ai/headroom` issue first. - This is the more consequential of two related fixes from the same investigation; the sibling PR (`fix(proxy/savings): append history point on cache-only savings too`) fixes a savings-history reporting gap that this same `--mode cache` + Bedrock deployment surfaced. --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: JerrettDavis <mxjerrett@gmail.com>
This commit is contained in:
parent
0537cbfde4
commit
a352fa0168
4 changed files with 409 additions and 2 deletions
|
|
@ -13,6 +13,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|||
- **install:** `headroom install apply --env KEY=VALUE` (repeatable) passes environment variables into supervised runners (macOS launchd, Linux systemd/cron, Windows services/tasks). These runners previously started with a bare environment and did not inherit the interactive shell's exports — e.g. a custom `HEADROOM_WORKSPACE_DIR` never reached the supervised process, so `headroom install agent run` looked for its manifest in the wrong location and failed outright even though `install apply` itself succeeded. `--env` values are merged into `DeploymentManifest.base_env` last, so they can override auto-derived defaults, and are threaded into the generated `run-headroom.sh`/`ensure-headroom.sh` (and Windows equivalents) as `export`/`$env:` lines before the `exec`.
|
||||
|
||||
### Fixed
|
||||
- **proxy/bedrock:** wire `PrefixCacheTracker` updates into both Bedrock backend paths (`handle_anthropic_messages`'s non-streaming branch in `anthropic.py`, and `_stream_response_bedrock` in `streaming.py`). `update_from_response()` was previously only called from the direct-Anthropic-API branch; both Bedrock branches returned before ever reaching it, so the tracker's state stayed permanently empty for the life of a session on any `--backend bedrock` deployment: `extract_cache_stable_delta()` always saw no previous turn, and `--mode cache` fell back to full unmodified passthrough on every turn instead of freezing the already-cached prefix and compressing only the new suffix.
|
||||
- **install:** `install_supervisor`'s macOS branch did an unconditional `launchctl bootout` followed by a bare `bootstrap` with no retry, unlike `start_supervisor`, which already rides out the ~15s EIO (error 5) window launchd exhibits for several seconds after a bootout. This left `install apply`'s own reinstall path (and anything that re-applies a deployment, e.g. a future `headroom doctor --fix`) exposed to a race that previously required manual recovery (bootout + remove the plist + reapply). Extracted the retry loop already used by `start_supervisor` into a shared `_bootstrap_with_retry()` helper, now used by both call sites.
|
||||
- **proxy/savings:** `SavingsTracker.record_request()` only appended a history point when `tokens_saved > 0` (headroom's own lossy compression). In `--mode cache`, `tokens_saved` is near-always 0 by design, since the frozen prefix is byte-replayed rather than compressed to keep the provider's prompt cache warm. That silently dropped every history point on a cache-mode deployment even when `cache_read_tokens`/`cache_savings_usd` were large, making `headroom-monthly`-style tooling read as a total savings collapse. The guard now fires on `tokens_saved` OR `cache_read_tokens`, and the appended entry carries `cache_read_tokens`/`cache_savings_usd` so downstream consumers can show them; `_normalize_history_entry` defaults both fields to 0/0.0 for legacy entries that predate this change.
|
||||
- **litellm:** vendor-specific top-level fields on `/v1/chat/completions`, including vLLM's `chat_template_kwargs` for per-request Qwen3 thinking-mode toggles, now reach OpenAI-compatible backends through LiteLLM `extra_body` instead of being dropped by the standard-parameter allowlist ([#2128](https://github.com/headroomlabs-ai/headroom/issues/2128)).
|
||||
|
|
|
|||
|
|
@ -2451,6 +2451,8 @@ class AnthropicHandlerMixin:
|
|||
optimization_latency,
|
||||
pipeline_timing=pipeline_timing,
|
||||
original_messages=original_client_messages,
|
||||
prefix_tracker=prefix_tracker,
|
||||
optimized_messages=optimized_messages,
|
||||
)
|
||||
else:
|
||||
async with stage_timer.measure("upstream_connect"):
|
||||
|
|
@ -2536,6 +2538,44 @@ class AnthropicHandlerMixin:
|
|||
0, attempted_input_tokens - cr_tokens - cw_tokens
|
||||
)
|
||||
|
||||
# Update prefix cache tracker for next turn. Mirrors the
|
||||
# direct-Anthropic-API branch below (~line 3011) — without
|
||||
# this, PrefixCacheTracker never sees a turn 2+ update on
|
||||
# the Bedrock path, extract_cache_stable_delta() always
|
||||
# returns None (no previous_original_messages), and cache
|
||||
# mode falls back to full unmodified passthrough every
|
||||
# turn instead of compressing the append-only delta.
|
||||
next_original_messages = copy.deepcopy(original_client_messages)
|
||||
next_forwarded_messages = copy.deepcopy(optimized_messages)
|
||||
assistant_message = self._assistant_message_from_response_json(
|
||||
backend_response.body
|
||||
)
|
||||
if assistant_message is not None:
|
||||
next_original_messages.append(copy.deepcopy(assistant_message))
|
||||
next_forwarded_messages.append(copy.deepcopy(assistant_message))
|
||||
if hasattr(prefix_tracker, "classify_cache_miss"):
|
||||
miss = prefix_tracker.classify_cache_miss(
|
||||
cache_read_tokens=cr_tokens,
|
||||
current_forwarded_messages=optimized_messages,
|
||||
)
|
||||
if miss.is_miss:
|
||||
logger.info(
|
||||
f"[{request_id}] CACHE-MISS-ATTRIBUTION: reason={miss.reason} "
|
||||
f"idle={miss.idle_seconds:.0f}s ttl={miss.cache_ttl_seconds}s "
|
||||
f"expected_cached={miss.expected_cached_tokens:,} "
|
||||
f"prefix_changed={miss.prefix_changed} "
|
||||
f"ttl_exceeded={miss.ttl_exceeded}"
|
||||
)
|
||||
await self.metrics.record_cache_miss_attribution(
|
||||
provider_name, miss.reason
|
||||
)
|
||||
prefix_tracker.update_from_response(
|
||||
cache_read_tokens=cr_tokens,
|
||||
cache_write_tokens=cw_tokens,
|
||||
messages=next_forwarded_messages,
|
||||
original_messages=next_original_messages,
|
||||
)
|
||||
|
||||
await self._record_request_outcome(
|
||||
RequestOutcome(
|
||||
request_id=request_id,
|
||||
|
|
|
|||
|
|
@ -1643,10 +1643,21 @@ class StreamingMixin:
|
|||
optimization_latency: float,
|
||||
pipeline_timing: dict[str, float] | None = None,
|
||||
original_messages: list[dict] | None = None,
|
||||
prefix_tracker: Any | None = None,
|
||||
optimized_messages: list[dict] | None = None,
|
||||
) -> StreamingResponse:
|
||||
"""Stream response from Bedrock backend with metrics tracking.
|
||||
|
||||
Translates Bedrock streaming events to Anthropic SSE format.
|
||||
|
||||
``prefix_tracker``/``optimized_messages`` carry the
|
||||
:class:`PrefixCacheTracker` for the session so cache stats from
|
||||
this turn update the tracker for the next one — mirrors the
|
||||
direct streaming path (``_finalize_stream_response``) and the
|
||||
OpenAI-via-backend sibling (``_stream_openai_via_backend``).
|
||||
Without this, ``extract_cache_stable_delta()`` always sees no
|
||||
previous turn on the Bedrock path and cache mode never compresses
|
||||
anything past the first request in a session.
|
||||
"""
|
||||
from fastapi.responses import StreamingResponse
|
||||
|
||||
|
|
@ -1668,6 +1679,11 @@ class StreamingMixin:
|
|||
"cache_creation_ephemeral_5m_input_tokens": 0,
|
||||
"cache_creation_ephemeral_1h_input_tokens": 0,
|
||||
}
|
||||
# Bytes-level mirror of the SSE stream, used only to reconstruct
|
||||
# the final assistant message for the prefix tracker once the
|
||||
# stream closes (see finally: block below). Not on the hot path
|
||||
# for anything the client sees.
|
||||
full_sse_bytes = bytearray()
|
||||
|
||||
async def generate():
|
||||
try:
|
||||
|
|
@ -1695,10 +1711,13 @@ class StreamingMixin:
|
|||
|
||||
# Format as SSE
|
||||
if event.raw_sse:
|
||||
yield event.raw_sse.encode()
|
||||
chunk_bytes = event.raw_sse.encode()
|
||||
else:
|
||||
sse_line = f"event: {event.event_type}\ndata: {json.dumps(event.data)}\n\n"
|
||||
yield sse_line.encode()
|
||||
chunk_bytes = sse_line.encode()
|
||||
if prefix_tracker is not None:
|
||||
full_sse_bytes.extend(chunk_bytes)
|
||||
yield chunk_bytes
|
||||
|
||||
# Track usage from message_start event
|
||||
if event.event_type == "message_start":
|
||||
|
|
@ -1749,6 +1768,53 @@ class StreamingMixin:
|
|||
_backend_name = (
|
||||
self.anthropic_backend.name if self.anthropic_backend else "anthropic"
|
||||
)
|
||||
|
||||
# Update prefix cache tracker for the next turn — mirrors
|
||||
# _finalize_stream_response (direct-API streaming path)
|
||||
# and _stream_openai_via_backend (OpenAI-via-backend
|
||||
# sibling). Run before the outcome funnel so prefix state
|
||||
# is consistent regardless of metric path.
|
||||
if prefix_tracker is not None:
|
||||
import copy as _copy
|
||||
|
||||
tracker_messages = (
|
||||
optimized_messages
|
||||
if optimized_messages is not None
|
||||
else body.get("messages", [])
|
||||
)
|
||||
next_forwarded = _copy.deepcopy(tracker_messages)
|
||||
next_original = _copy.deepcopy(original_messages or tracker_messages)
|
||||
if full_sse_bytes:
|
||||
parsed = self._parse_sse_to_response(
|
||||
full_sse_bytes.decode("utf-8", errors="replace"), provider
|
||||
)
|
||||
asst_msg = self._assistant_message_from_response_json(parsed)
|
||||
if asst_msg is not None:
|
||||
next_forwarded.append(_copy.deepcopy(asst_msg))
|
||||
next_original.append(_copy.deepcopy(asst_msg))
|
||||
cache_read_tokens = stream_state["cache_read_input_tokens"] or 0
|
||||
cache_write_tokens = stream_state["cache_creation_input_tokens"] or 0
|
||||
if provider == "anthropic" and hasattr(prefix_tracker, "classify_cache_miss"):
|
||||
miss = prefix_tracker.classify_cache_miss(
|
||||
cache_read_tokens=cache_read_tokens,
|
||||
current_forwarded_messages=tracker_messages,
|
||||
)
|
||||
if miss.is_miss:
|
||||
logger.info(
|
||||
f"[{request_id}] CACHE-MISS-ATTRIBUTION: reason={miss.reason} "
|
||||
f"idle={miss.idle_seconds:.0f}s ttl={miss.cache_ttl_seconds}s "
|
||||
f"expected_cached={miss.expected_cached_tokens:,} "
|
||||
f"prefix_changed={miss.prefix_changed} "
|
||||
f"ttl_exceeded={miss.ttl_exceeded}"
|
||||
)
|
||||
await self.metrics.record_cache_miss_attribution(provider, miss.reason)
|
||||
prefix_tracker.update_from_response(
|
||||
cache_read_tokens=cache_read_tokens,
|
||||
cache_write_tokens=cache_write_tokens,
|
||||
messages=next_forwarded,
|
||||
original_messages=next_original,
|
||||
)
|
||||
|
||||
# Active-compression denominator derived inside
|
||||
# ``from_stream`` as ``optimized + saved``. Bedrock
|
||||
# doesn't propagate frozen_message_count either — same
|
||||
|
|
|
|||
300
tests/test_bedrock_prefix_tracker_wiring.py
Normal file
300
tests/test_bedrock_prefix_tracker_wiring.py
Normal file
|
|
@ -0,0 +1,300 @@
|
|||
"""Regression coverage for PrefixCacheTracker wiring on Bedrock backend paths.
|
||||
|
||||
Both Bedrock-routed branches of ``handle_anthropic_messages``
|
||||
(non-streaming in ``anthropic.py``, streaming ``_stream_response_bedrock``
|
||||
in ``streaming.py``) used to return before ever calling
|
||||
``prefix_tracker.update_from_response()``. Only the direct-Anthropic-API
|
||||
branch called it. Practical effect: on any ``--backend bedrock --mode
|
||||
cache`` deployment, ``PrefixCacheTracker`` state stayed permanently at
|
||||
turn 0 for the life of a session — ``get_frozen_message_count()`` always
|
||||
returned 0, ``extract_cache_stable_delta()`` always saw no previous turn,
|
||||
and cache mode fell back to full unmodified passthrough on every single
|
||||
turn instead of freezing the already-cached prefix and compressing only
|
||||
the new suffix.
|
||||
|
||||
These tests drive two turns through the real proxy (with a mocked
|
||||
Bedrock-shaped backend) and inspect the real ``PrefixCacheTracker`` the
|
||||
proxy keeps in ``session_tracker_store`` — not a fake — to pin that the
|
||||
tracker's turn counter and last-forwarded/-original messages actually
|
||||
advance after a Bedrock call, for both the non-streaming and the
|
||||
streaming code path.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from collections.abc import AsyncIterator
|
||||
from typing import Any
|
||||
from unittest.mock import MagicMock, patch
|
||||
|
||||
import pytest
|
||||
|
||||
fastapi = pytest.importorskip("fastapi")
|
||||
httpx = pytest.importorskip("httpx")
|
||||
|
||||
from fastapi.testclient import TestClient # noqa: E402
|
||||
|
||||
from headroom.backends.base import BackendResponse, StreamEvent # noqa: E402
|
||||
from headroom.proxy.server import ProxyConfig, create_app # noqa: E402
|
||||
|
||||
|
||||
def _make_anthropic_backend(body: dict[str, Any]) -> MagicMock:
|
||||
"""Mock backend whose ``send_message`` returns an Anthropic-shaped body."""
|
||||
|
||||
async def fake_send(body_: dict, headers: dict) -> BackendResponse:
|
||||
return BackendResponse(body=body, status_code=200)
|
||||
|
||||
mock = MagicMock()
|
||||
mock.name = "bedrock"
|
||||
mock.send_message = fake_send
|
||||
mock.map_model_id = MagicMock(return_value="claude-3-5-sonnet-20241022")
|
||||
mock.supports_model = MagicMock(return_value=True)
|
||||
return mock
|
||||
|
||||
|
||||
def _make_bedrock_streaming_backend(events: list[StreamEvent]) -> MagicMock:
|
||||
"""Mock backend that yields Anthropic ``StreamEvent`` objects."""
|
||||
|
||||
async def fake_stream(body: dict, headers: dict) -> AsyncIterator[StreamEvent]:
|
||||
for evt in events:
|
||||
yield evt
|
||||
|
||||
mock = MagicMock()
|
||||
mock.name = "bedrock"
|
||||
mock.stream_message = fake_stream
|
||||
mock.map_model_id = MagicMock(return_value="claude-3-5-sonnet-20241022")
|
||||
mock.supports_model = MagicMock(return_value=True)
|
||||
return mock
|
||||
|
||||
|
||||
def _sse_data(event_type: str, data: dict[str, Any]) -> str:
|
||||
return f"event: {event_type}\ndata: {json.dumps(data)}\n\n"
|
||||
|
||||
|
||||
def _cache_config() -> ProxyConfig:
|
||||
return ProxyConfig(
|
||||
optimize=False,
|
||||
cache_enabled=False,
|
||||
rate_limit_enabled=False,
|
||||
backend="anyllm",
|
||||
anyllm_provider="anthropic",
|
||||
mode="cache",
|
||||
)
|
||||
|
||||
|
||||
def _anthropic_body(cache_read: int, cache_write: int) -> dict[str, Any]:
|
||||
return {
|
||||
"id": "msg_1",
|
||||
"type": "message",
|
||||
"role": "assistant",
|
||||
"model": "claude-3-5-sonnet-20241022",
|
||||
"content": [{"type": "text", "text": "hi"}],
|
||||
"stop_reason": "end_turn",
|
||||
"usage": {
|
||||
"input_tokens": 1000,
|
||||
"output_tokens": 50,
|
||||
"cache_read_input_tokens": cache_read,
|
||||
"cache_creation_input_tokens": cache_write,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
# =============================================================================
|
||||
# Non-streaming Bedrock path (anthropic.py)
|
||||
# =============================================================================
|
||||
|
||||
|
||||
def test_bedrock_nonstreaming_advances_prefix_tracker_turn() -> None:
|
||||
"""A non-streaming Bedrock request must call ``update_from_response``.
|
||||
|
||||
Before the fix, the Bedrock non-streaming branch returned its
|
||||
``JSONResponse`` without ever touching ``prefix_tracker`` — the
|
||||
tracker stayed at ``_turn_number == 0`` and ``_last_original_messages
|
||||
== []`` no matter how many turns went through. After the fix, one
|
||||
turn through this path must leave the tracker recording turn 1 and
|
||||
the sent + assistant messages as its "last" snapshot.
|
||||
"""
|
||||
config = _cache_config()
|
||||
backend = _make_anthropic_backend(_anthropic_body(cache_read=500, cache_write=200))
|
||||
|
||||
with patch("headroom.proxy.server.AnyLLMBackend", return_value=backend):
|
||||
app = create_app(config)
|
||||
proxy = app.state.proxy
|
||||
with TestClient(app) as client:
|
||||
resp = client.post(
|
||||
"/v1/messages",
|
||||
json={
|
||||
"model": "claude-3-5-sonnet-20241022",
|
||||
"messages": [{"role": "user", "content": "hi"}],
|
||||
"max_tokens": 64,
|
||||
},
|
||||
headers={
|
||||
"x-api-key": "sk-ant-test",
|
||||
"anthropic-version": "2023-06-01",
|
||||
"x-headroom-session-id": "bedrock-nonstream-session",
|
||||
},
|
||||
)
|
||||
assert resp.status_code == 200, resp.text[:200]
|
||||
|
||||
tracker = proxy.session_tracker_store.get_or_create("bedrock-nonstream-session", "anthropic")
|
||||
assert tracker._turn_number == 1, (
|
||||
"prefix tracker never advanced past turn 0 — update_from_response() "
|
||||
"was not called on the Bedrock non-streaming path"
|
||||
)
|
||||
assert tracker.get_last_original_messages(), (
|
||||
"tracker recorded no 'last turn' messages — the Bedrock non-streaming "
|
||||
"branch is not feeding it the sent + assistant messages"
|
||||
)
|
||||
# cache_read=500 + cache_write=200 = 700 total_cached, above the default
|
||||
# min_cached_tokens=1024 threshold is NOT met here, but the turn/messages
|
||||
# advancing (asserted above) is the actual regression signal — frozen
|
||||
# count only matters once the session crosses the threshold, which is
|
||||
# covered by test_cross_turn_cache_safety.py and test_cache/test_prefix_tracker.py.
|
||||
|
||||
|
||||
def test_bedrock_nonstreaming_second_turn_sees_frozen_prefix() -> None:
|
||||
"""Two Bedrock non-streaming turns: turn 2 must see turn 1 as its frozen prefix.
|
||||
|
||||
This is the concrete consequence of the tracker actually updating:
|
||||
once cache_read+cache_write clears ``min_cached_tokens``, turn 2's
|
||||
``get_frozen_message_count()`` must be nonzero and its
|
||||
``get_last_original_messages()`` must equal turn 1's full message
|
||||
history (user + assistant) — the input the freeze/delta-compression
|
||||
path needs to detect an append-only turn. Before the fix this was
|
||||
always 0 / [] regardless of turn count.
|
||||
"""
|
||||
config = _cache_config()
|
||||
# 1200 total cached tokens clears the default min_cached_tokens=1024.
|
||||
backend = _make_anthropic_backend(_anthropic_body(cache_read=1000, cache_write=200))
|
||||
|
||||
with patch("headroom.proxy.server.AnyLLMBackend", return_value=backend):
|
||||
app = create_app(config)
|
||||
proxy = app.state.proxy
|
||||
with TestClient(app) as client:
|
||||
turn1 = client.post(
|
||||
"/v1/messages",
|
||||
json={
|
||||
"model": "claude-3-5-sonnet-20241022",
|
||||
"messages": [{"role": "user", "content": "hi"}],
|
||||
"max_tokens": 64,
|
||||
},
|
||||
headers={
|
||||
"x-api-key": "sk-ant-test",
|
||||
"anthropic-version": "2023-06-01",
|
||||
"x-headroom-session-id": "bedrock-nonstream-2turn",
|
||||
},
|
||||
)
|
||||
assert turn1.status_code == 200, turn1.text[:200]
|
||||
|
||||
turn2 = client.post(
|
||||
"/v1/messages",
|
||||
json={
|
||||
"model": "claude-3-5-sonnet-20241022",
|
||||
"messages": [
|
||||
{"role": "user", "content": "hi"},
|
||||
{"role": "assistant", "content": [{"type": "text", "text": "hi"}]},
|
||||
{"role": "user", "content": "and again"},
|
||||
],
|
||||
"max_tokens": 64,
|
||||
},
|
||||
headers={
|
||||
"x-api-key": "sk-ant-test",
|
||||
"anthropic-version": "2023-06-01",
|
||||
"x-headroom-session-id": "bedrock-nonstream-2turn",
|
||||
},
|
||||
)
|
||||
assert turn2.status_code == 200, turn2.text[:200]
|
||||
|
||||
tracker = proxy.session_tracker_store.get_or_create("bedrock-nonstream-2turn", "anthropic")
|
||||
assert tracker._turn_number == 2
|
||||
assert tracker.get_frozen_message_count() > 0, (
|
||||
"frozen_message_count stayed 0 on turn 2 despite a cache hit on turn 1 "
|
||||
"— PrefixCacheTracker never saw turn 1's response"
|
||||
)
|
||||
|
||||
|
||||
# =============================================================================
|
||||
# Streaming Bedrock path (streaming.py, _stream_response_bedrock)
|
||||
# =============================================================================
|
||||
|
||||
|
||||
def test_bedrock_streaming_advances_prefix_tracker_turn() -> None:
|
||||
"""A streaming Bedrock request must also call ``update_from_response``.
|
||||
|
||||
Mirrors the non-streaming test above for ``_stream_response_bedrock``.
|
||||
Before the fix, this function had no ``prefix_tracker`` parameter at
|
||||
all — the tracker was never even threaded in, let alone updated.
|
||||
"""
|
||||
config = _cache_config()
|
||||
|
||||
message_start = {
|
||||
"type": "message_start",
|
||||
"message": {
|
||||
"id": "msg_1",
|
||||
"model": "claude-3-5-sonnet-20241022",
|
||||
"role": "assistant",
|
||||
"type": "message",
|
||||
"content": [],
|
||||
"usage": {
|
||||
"input_tokens": 1000,
|
||||
"cache_read_input_tokens": 500,
|
||||
"cache_creation_input_tokens": 200,
|
||||
},
|
||||
},
|
||||
}
|
||||
block_start = {
|
||||
"type": "content_block_start",
|
||||
"index": 0,
|
||||
"content_block": {"type": "text", "text": ""},
|
||||
}
|
||||
block_delta = {
|
||||
"type": "content_block_delta",
|
||||
"index": 0,
|
||||
"delta": {"type": "text_delta", "text": "hi"},
|
||||
}
|
||||
block_stop = {"type": "content_block_stop", "index": 0}
|
||||
message_delta = {
|
||||
"type": "message_delta",
|
||||
"delta": {"stop_reason": "end_turn"},
|
||||
"usage": {"output_tokens": 50},
|
||||
}
|
||||
message_stop = {"type": "message_stop"}
|
||||
|
||||
events = [
|
||||
StreamEvent(event_type=e["type"], data=e, raw_sse=_sse_data(e["type"], e))
|
||||
for e in [message_start, block_start, block_delta, block_stop, message_delta, message_stop]
|
||||
]
|
||||
backend = _make_bedrock_streaming_backend(events)
|
||||
|
||||
with patch("headroom.proxy.server.AnyLLMBackend", return_value=backend):
|
||||
app = create_app(config)
|
||||
proxy = app.state.proxy
|
||||
with TestClient(app) as client:
|
||||
resp = client.post(
|
||||
"/v1/messages",
|
||||
json={
|
||||
"model": "claude-3-5-sonnet-20241022",
|
||||
"messages": [{"role": "user", "content": "hi"}],
|
||||
"max_tokens": 64,
|
||||
"stream": True,
|
||||
},
|
||||
headers={
|
||||
"x-api-key": "sk-ant-test",
|
||||
"anthropic-version": "2023-06-01",
|
||||
"x-headroom-session-id": "bedrock-stream-session",
|
||||
},
|
||||
)
|
||||
assert resp.status_code == 200, resp.text[:200]
|
||||
assert "message_stop" in resp.text
|
||||
|
||||
tracker = proxy.session_tracker_store.get_or_create("bedrock-stream-session", "anthropic")
|
||||
assert tracker._turn_number == 1, (
|
||||
"prefix tracker never advanced past turn 0 on the Bedrock streaming "
|
||||
"path — update_from_response() was not called from "
|
||||
"_stream_response_bedrock"
|
||||
)
|
||||
assert tracker.get_last_original_messages(), (
|
||||
"tracker recorded no 'last turn' messages on the streaming path — "
|
||||
"the reconstructed assistant message from the SSE stream never "
|
||||
"reached the tracker"
|
||||
)
|
||||
Loading…
Add table
Add a link
Reference in a new issue