headroom/tests/test_ccr_row_drop_store_bridge.py
Tejas Chopra f2c48e26c6
feat(compress): reach the lossless provider seam on the general path and default /v1/compress to marker-free output (#2691)
## Description

Two related changes to the compression seams, plus the review fixes for
both. Supersedes #2661 and #2662, which are closed in favour of this
branch — the fixes are inseparable from the code they fix, so reviewing
them together is cheaper than landing two PRs and patching them
afterwards.

**1. A registered lossless provider now competes on the general path.**
The `headroom.transforms.lossless_provider` seam was only ever consulted
from `_lossless_compact_excluded`, gated on `DEFAULT_EXCLUDE_TOOLS`
(`config.py:216` — `Read/Grep/Glob/Write/Edit/WebSearch/WebFetch`).
Gateway traffic carries the caller's own tool names — LiteLLM's
`headroom` guardrail (https://docs.litellm.ai/docs/proxy/headroom) posts
requests containing tools like `search_docs` / `run_ci` / `fetch_rows` —
so a registered provider was structurally unreachable for every
gateway/sidecar deployment. The seam existed; nothing could get to it.

**2. `POST /v1/compress` is marker-free by default.** A CCR marker is
only useful to a caller that also injects the `headroom_retrieve` tool
AND can reach `/v1/retrieve`. Neither holds here: tool injection lives
in the provider request handlers (`handlers/anthropic.py:1894`), never
in `handle_compress`; and every `/v1/retrieve*` route is
`Depends(_require_loopback)` (`server.py:4422, 4470, 4749, 4781`) with
no remote opt-in — `HEADROOM_COMPRESS_ALLOW_REMOTE` drops the loopback
dependency on `/v1/compress` only. So a gateway forwards a `Retrieve
more: hash=…` pointer the model cannot follow, and the proxy pays a CCR
store write nobody reads. `config.mode="ccr"` opts back in.

Closes #

## Type of Change

- [ ] Bug fix (non-breaking change that fixes an issue)
- [x] New feature (non-breaking change that adds functionality)
- [x] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

### Seam — `content_router.py`, `lossless_provider.py`

- `_lossless_first` (STAGE 0, every block on every path) consults
`get_lossless_provider()` and keeps whichever output is smaller.
**Strict no-op when no provider is registered**, which is the default;
and because it is best-of rather than authoritative, a provider can
never do worse than the built-in folds.
- Malformed provider output can no longer escape. Every shape check runs
inside the `try`: result must be `None`, or a 2-element tuple/list of
two `str`. Anything else is ignored at debug level. (Previously the
unpack sat outside the `try`, so a 3-tuple raised `ValueError` up
through `TransformPipeline.apply`, which re-raises.)
- Empty / whitespace-only candidates are rejected rather than silently
replacing block content.
- **Providers are never offered diff content.** Diff folding is
subtractive with no inverse check and a reflowed hunk breaks `git apply`
— the same reason the built-in `diff` fold is restricted at
`content_router.py:2481`.
- The third-party `kind` label is sanitised against `^[a-z0-9_]{1,32}$`
before reaching `transforms_applied` and the per-strategy metric dicts,
so a caller-controlled string cannot explode Prometheus label
cardinality. `fullmatch`, not `match`: `$` also matches before a
trailing newline, which would put a newline in a label.
- `set_lossless_provider(provider, *, verifier=None)` — in lossless-only
mode, where STAGE 0's output is final and there is no marker to recover
from, a registered verifier must confirm the fold or the candidate is
dropped. No verifier registered = today's behaviour. `provider=None`
clears both.
- The provider is invoked once per block, not twice
(`_has_lossless_fold` probes `_lossless_first` and discards the result,
then STAGE 0 recomputes). Bounded memo, wholesale clear on overflow, no
lock — a race costs one redundant fold. The memo keys on the provider
registration generation, so registering or clearing a provider after a
block was already folded takes effect.
- The seam docstring now records that the provider runs on the general
path and inside the parallel compression pool, so it must be thread-safe
as well as deterministic.

### Route — `handlers/openai.py`, `server.py`

- `_derived_compress_pipeline(key, **overrides)` replaces the
copy-pasted pipeline-derivation block; `_no_ccr_pipeline` (the new
default) and `_lossy_inline_pipeline` both use it.
- The default pipeline is **built at startup** and included in
`_eager_preload_transforms`, so a fresh pod does not pay ContentRouter
construction and compressor load on its first request, inside the
compression-executor budget.
- An unrecognised `config.mode` returns 400 naming the valid values
instead of silently falling back to the default.
- Claude-family model names resolve their context limit from the
Anthropic provider. Real divergence:
`bedrock/anthropic.claude-3-5-sonnet` is 200000 there and 128000 on the
OpenAI provider. The tokenizer still comes from the OpenAI pipeline's
provider — a separate, larger change, noted in a comment.
- Documents why the derived router deliberately does **not** share the
base router's compression cache: keys do not encode CCR-marker mode, so
sharing would leak marker-laden entries into the marker-free path.

### Behavior change

A `/v1/compress` caller that relied on default markers now gets none.
The only in-tree caller that can resolve them is the TypeScript SDK
(`sdk/typescript/src/client.ts:398 retrieve`, `:422 handleToolCall`); it
needs `config: {"mode": "ccr"}` to keep today's behaviour, and landing
that SDK default in the same release would leave only gateway callers —
for whom markers were never resolvable — seeing a difference.

No `HEADROOM_COMPRESS_DEFAULT_MODE` compat env deliberately: a flag
nobody sets becomes permanent debt, and the wire-level `mode` already
covers the one caller that needs it.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ python -m pytest tests/test_lossless_first_dispatch.py tests/test_lossless_excluded_compaction.py \
    tests/test_lossless_mode.py tests/test_lossless_then_lossy.py tests/test_lossless_diff_fold_guard.py \
    tests/test_bash_search_lossless_fold.py tests/test_proxy_compress_endpoint.py \
    tests/test_ccr_row_drop_store_bridge.py tests/test_gateway_sidecar_ports.py tests/test_compress_api.py \
    tests/test_platform_stabilization_functional.py tests/test_proxy_eager_preload_bind.py \
    tests/test_proxy_warmup.py tests/test_router_registry_smartcrusher.py -q
189 passed in 32.04s

$ ruff check <all 9 changed files>
All checks passed!

$ ruff format --check <all 9 changed files>
9 files already formatted

$ mypy headroom/transforms/content_router.py headroom/transforms/lossless_provider.py \
      headroom/proxy/handlers/openai.py headroom/proxy/server.py
Success: no issues found in 4 source files
```

New tests cover, one concern each: every malformed provider shape; empty
and whitespace-only results; diff content never reaching a provider
(call-recording); `kind` sanitisation including the trailing-newline
case; the verifier accepting / rejecting / raising; clearing a provider
clearing its verifier; single provider invocation per block; memo
invalidation on re-registration; unknown and valid `mode` values; the
default pipeline existing before any request; and Claude vs OpenAI
context-limit resolution with `token_budget` precedence preserved.

## Real Behavior Proof

- **Environment:** macOS arm64, Python 3.12.6, proxy built from
`_proxy_config_from_env()` with the default `coding` savings profile;
Kompress both disabled and offloaded to a remote `kompress-v2-base`
endpoint; `HEADROOM_COMPRESSION_TIMEOUT_SECONDS=300`.
- **Steps:** `POST /v1/compress` over `TestClient` with OpenAI-shaped
payloads under non-excluded tool names (`run_ci`, `list_files`,
`code_search`, `fetch_rows`) — a CI log with ANSI escapes and repeated
lines, a 160-path listing, a 150-line grep dump, a 150-row JSON array;
plus a second payload with a RAG user blob, a 200-row JSON tool result
and a 300-line log. Ran with and without a provider registered via
`set_lossless_provider`.
- **Observed — seam reachability:**

  | Kompress | no provider registered | provider registered |
  |---|---|---|
  | off | 19,284 → 10,265 tokens (46.8%) | 19,284 → **7,879 (59.1%)** |
| on (remote) | 19,284 → 9,366 tokens (51.4%) | 19,284 → **7,146
(62.9%)** |

Before this change the right-hand column was identical to the left — the
registered provider was never called on this payload.

- **Observed — marker-free default costs nothing:**

  | Config | tokens | saved |
  |---|---|---|
  | markers on (previous default) | 37,791 → 24,415 | 35.4% |
| markers off (new default) | 37,791 → 24,415 | **35.4% — identical** |
  | `mode="lossy_inline"` | 37,791 → 25,129 | 33.5% |
  | `--lossless` | 37,791 → 35,100 | 7.1% |

- **Observed — memo staleness, before the fix:** registering a provider
that folds a grep block to 5 bytes left the block at its 1496-byte
built-in fold, and clearing a provider kept serving the provider's
output. Both correct after keying on the registration generation.
- **Observed — context limit:** `bedrock/anthropic.claude-3-5-sonnet`
resolves 200000 via the Anthropic provider, 128000 via the OpenAI
provider.
- **Not tested:** `tests/test_transforms_content_router.py` was not run
— it does not complete on this machine, wedging on its 5th test while
that test passes in 5.7s alone. Verified pre-existing before this work:
with the diff stashed, the clean tree stalled at the identical test, and
it stalled the same way under `HF_HUB_OFFLINE=1 HEADROOM_OFFLINE=1`. The
machine was also out of disk at the time, which may be the real cause
rather than the suspected native-detector deadlock (#575) — worth a
separate issue either way. CI should be the arbiter here.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md`

## Additional Notes

Two pre-existing test files needed adjusting, both direct consequences
rather than scope creep:

-
`test_platform_stabilization_functional.py::test_v1_compress_success_reports_actual_metrics`
patches `openai_pipeline.apply`, which the default mode no longer routes
through. **This test was already failing on the marker-free-default
commit** before any of the fixes — my original test selection missed it.
It now patches the pipeline the route actually uses.
- `test_proxy_eager_preload_bind.py` substitutes fake pipelines to
control exactly what the preload walks; the eager build injected the
real derived router's statuses into an exact-equality assertion. Its
shared helper now clears the derived cache, preserving each test's
intent without weakening an assertion.

Docs unchecked — follow-ups worth doing in the same release: document
`config.mode` values in `docs/content/docs/litellm.mdx` and
`wiki/proxy.md`; the TS SDK `mode:"ccr"` default; and an operational
note that `COMPRESSION_TIMEOUT_SECONDS` defaults to 30
(`helpers.py:687`) while a remote ML endpoint makes one sequential call
per unit — on a large payload it trips the executor timeout and the
handler fails open, returning `compression_skipped: true` with
`tokens_before: 0`, which reads as "nothing to save" rather than "we
gave up". Those zeroed counters are misleading and worth a separate fix.
2026-07-31 12:31:38 -07:00

491 lines
19 KiB
Python

"""Issue #389: SmartCrusher row-drop CCR hash → Python compression_store bridge.
The row-drop and opaque-blob paths in the Rust SmartCrusher emit
``<<ccr:HASH ...>>`` markers and stash the original payload in the
Rust process-local CCR store. The Python proxy's ``/v1/retrieve``
endpoint queries the Python ``compression_store`` (not the Rust one),
so without a bridge every retrieve call for a Rust-emitted marker
returns 404.
These tests pin the bridge:
1. Unit: lossy crush of a 200-item array populates the Python
compression_store keyed by the same hash that's in the marker, with
the canonical original retrievable via ``store.retrieve(hash)``.
2. Integration: ``/v1/compress`` followed by ``/v1/retrieve/{hash}``
on the same hash returns 200 with the original content. This is
the failing case from the issue's reproducer.
3. Edge: opaque-blob markers (the ``<<ccr:HASH,KIND,SIZE>>`` shape)
also bridge — the document walker emits these for long strings.
If these regress, ``/v1/retrieve`` silently 404s for every
SmartCrusher-emitted marker even though the data is held in the Rust
store. The LLM follows the marker, gets nothing, and the proxy's CCR
contract is broken at the bridge.
"""
from __future__ import annotations
import json
import pytest
def _build_extension() -> None:
try:
from headroom._core import SmartCrusher # noqa: F401
except ImportError:
pytest.skip(
"headroom._core not built — run `bash scripts/build_rust_extension.sh`",
allow_module_level=True,
)
_build_extension()
# Skip if fastapi not available — needed for integration tests but not unit.
try:
import fastapi # noqa: F401
_HAS_FASTAPI = True
except ImportError:
_HAS_FASTAPI = False
# ─── Unit tests: shim populates Python store ─────────────────────────────
def test_lossy_crush_populates_python_compression_store() -> None:
"""The cornerstone unit test: trigger a row-drop via the Python
SmartCrusher shim, then verify the Python compression_store has
an entry keyed by the marker's hash with the canonical original.
"""
from headroom.cache.compression_store import (
get_compression_store,
reset_compression_store,
)
from headroom.config import CCRConfig
from headroom.config import SmartCrusherConfig as PyConfig
from headroom.transforms.smart_crusher import SmartCrusher
reset_compression_store()
try:
crusher = SmartCrusher(PyConfig(), ccr_config=CCRConfig(), with_compaction=False)
# 60 items is well above adaptive_k → lossy path fires.
original = [{"id": i, "status": "ok", "tag": "alpha"} for i in range(60)]
original_json = json.dumps(original)
result = crusher.crush_array_json(original_json)
# Sanity: lossy path actually fired.
assert result["ccr_hash"] is not None, (
f"expected lossy drop, got strategy={result['strategy_info']!r}"
)
ccr_hash = result["ccr_hash"]
# The marker text embeds the same hash.
assert ccr_hash in result["dropped_summary"], (
f"marker {result['dropped_summary']!r} should embed hash {ccr_hash}"
)
# The bridge: Python compression_store now holds an entry keyed
# by the Rust-emitted hash.
store = get_compression_store()
entry = store.retrieve(ccr_hash)
assert entry is not None, (
f"compression_store has no entry for hash {ccr_hash!r}; "
f"the bridge dropped the row-drop hash on the floor"
)
# The entry's original content is the canonical-JSON
# serialization of the original array — same bytes the Rust
# store has under the same hash.
rust_canonical = crusher.ccr_get(ccr_hash)
assert rust_canonical is not None, "Rust store lost the entry"
assert entry.original_content == rust_canonical, (
"Python store's canonical bytes diverged from Rust store"
)
# Round-trip: parse the canonical and compare to input.
retrieved = json.loads(entry.original_content)
assert retrieved == original
finally:
reset_compression_store()
def test_smart_crush_content_populates_python_store() -> None:
"""Same bridge but driven through the runtime path the proxy
actually uses: `_smart_crush_content` (not `crush_array_json`).
This is what `apply()` invokes per-message."""
from headroom.cache.compression_store import (
get_compression_store,
reset_compression_store,
)
from headroom.config import CCRConfig
from headroom.config import SmartCrusherConfig as PyConfig
from headroom.transforms.smart_crusher import SmartCrusher
reset_compression_store()
try:
crusher = SmartCrusher(PyConfig(), ccr_config=CCRConfig(), with_compaction=False)
# Mix of "ok" + occasional "error" → variance the lossy path
# can latch onto. A purely-uniform array gets a `skip:unique_
# entities_no_signal` strategy and never row-drops.
original = [
{
"id": i,
"level": "error" if i % 30 == 0 else "info",
"msg": f"line {i}",
}
for i in range(80)
]
content = json.dumps(original)
crushed, was_modified, info = crusher._smart_crush_content(content)
assert was_modified, f"expected lossy modification, got info={info!r}"
assert "<<ccr:" in crushed, (
f"expected CCR marker in output (info={info!r}): {crushed[:200]!r}"
)
# Pull the hash from the rendered output via JSON parse.
# The output is a JSON array with a `_ccr_dropped` sentinel.
parsed = json.loads(crushed)
assert isinstance(parsed, list), f"expected JSON array, got {type(parsed).__name__}"
sentinel = parsed[-1]
assert isinstance(sentinel, dict) and "_ccr_dropped" in sentinel, (
f"expected _ccr_dropped sentinel as last element, got {sentinel!r}"
)
marker_text = sentinel["_ccr_dropped"]
assert marker_text.startswith("<<ccr:") and "rows_offloaded>>" in marker_text
# Extract hash by structural slice (no regex).
ccr_hash = marker_text[len("<<ccr:") :].split(" ", 1)[0]
assert all(c in "0123456789abcdef" for c in ccr_hash), f"hash should be hex: {ccr_hash!r}"
# The Python store now has the entry under this hash.
store = get_compression_store()
entry = store.retrieve(ccr_hash)
assert entry is not None, (
f"_smart_crush_content didn't bridge hash {ccr_hash!r} to "
f"Python store; /v1/retrieve would 404"
)
# Original is recoverable byte-for-byte from the bridged entry.
retrieved = json.loads(entry.original_content)
assert retrieved == original
finally:
reset_compression_store()
def test_passthrough_does_not_populate_store() -> None:
"""Below adaptive_k → no row drop → no Python store write.
Pins that we don't accidentally store on every crush call."""
from headroom.cache.compression_store import (
get_compression_store,
reset_compression_store,
)
from headroom.config import CCRConfig
from headroom.config import SmartCrusherConfig as PyConfig
from headroom.transforms.smart_crusher import SmartCrusher
reset_compression_store()
try:
crusher = SmartCrusher(PyConfig(), ccr_config=CCRConfig(), with_compaction=False)
# 3 items: well below the threshold; no compression happens.
small = json.dumps([{"id": i} for i in range(3)])
crusher._smart_crush_content(small)
store = get_compression_store()
stats = store.get_stats()
assert stats["entry_count"] == 0, (
f"passthrough crush should not write to compression_store, "
f"got {stats['entry_count']} entries"
)
finally:
reset_compression_store()
def test_marker_disabled_skips_python_store() -> None:
"""`ccr_config.enabled=False` flips off marker emission AND store
writes on the Rust side. The bridge should have nothing to do."""
from headroom.cache.compression_store import (
get_compression_store,
reset_compression_store,
)
from headroom.config import CCRConfig
from headroom.config import SmartCrusherConfig as PyConfig
from headroom.transforms.smart_crusher import SmartCrusher
reset_compression_store()
try:
# Markers disabled → Rust skips both marker emission and store write.
crusher = SmartCrusher(
PyConfig(),
ccr_config=CCRConfig(enabled=False),
with_compaction=False,
)
original = [{"id": i, "status": "ok"} for i in range(60)]
crushed, was_modified, _info = crusher._smart_crush_content(json.dumps(original))
# Compression still happens (rows still drop), but no marker.
assert was_modified
assert "<<ccr:" not in crushed, (
f"expected no marker when ccr_config.enabled=False, got: {crushed[:200]!r}"
)
# And the Python store stays empty — bridge had nothing to mirror.
store = get_compression_store()
assert store.get_stats()["entry_count"] == 0
finally:
reset_compression_store()
def test_distinct_payloads_get_distinct_python_store_entries() -> None:
"""Two unrelated payloads → two row drops → two entries under
distinct hashes in the Python store; both retrievable independently."""
from headroom.cache.compression_store import (
get_compression_store,
reset_compression_store,
)
from headroom.config import CCRConfig
from headroom.config import SmartCrusherConfig as PyConfig
from headroom.transforms.smart_crusher import SmartCrusher
reset_compression_store()
try:
crusher = SmartCrusher(PyConfig(), ccr_config=CCRConfig(), with_compaction=False)
a = [{"id": i, "tag": "alpha"} for i in range(50)]
b = [{"id": i, "tag": "beta"} for i in range(50)]
ra = crusher.crush_array_json(json.dumps(a))
rb = crusher.crush_array_json(json.dumps(b))
ha, hb = ra["ccr_hash"], rb["ccr_hash"]
assert ha and hb and ha != hb
store = get_compression_store()
ea = store.retrieve(ha)
eb = store.retrieve(hb)
assert ea is not None and eb is not None
assert json.loads(ea.original_content) == a
assert json.loads(eb.original_content) == b
finally:
reset_compression_store()
# ─── compression_store: explicit_hash parameter ───────────────────────────
def test_compression_store_explicit_hash_round_trips() -> None:
"""The new `explicit_hash` parameter on `store.store()` keys the
entry by the caller-supplied hash instead of MD5(original)[:24]."""
from headroom.cache.compression_store import (
get_compression_store,
reset_compression_store,
)
reset_compression_store()
try:
store = get_compression_store()
# SmartCrusher emits 12-char SHA-256 hashes — much shorter than
# the default MD5[:24].
explicit = "abc123def456"
returned = store.store(
original='[{"id":1}]',
compressed="<<placeholder>>",
explicit_hash=explicit,
)
assert returned == explicit
entry = store.retrieve(explicit)
assert entry is not None
assert entry.original_content == '[{"id":1}]'
finally:
reset_compression_store()
def test_compression_store_explicit_hash_rejects_non_hex() -> None:
"""Non-hex `explicit_hash` raises ValueError. No silent fallback
to MD5 — that would re-introduce the marker/store mismatch."""
from headroom.cache.compression_store import (
get_compression_store,
reset_compression_store,
)
reset_compression_store()
try:
store = get_compression_store()
with pytest.raises(ValueError, match="hex string"):
store.store(
original="x",
compressed="y",
explicit_hash="NOT_HEX!@#",
)
with pytest.raises(ValueError, match="hex string"):
store.store(
original="x",
compressed="y",
explicit_hash="",
)
finally:
reset_compression_store()
# ─── Integration test: /v1/compress → /v1/retrieve via FastAPI ──────────
@pytest.mark.skipif(not _HAS_FASTAPI, reason="fastapi not installed")
def test_v1_compress_then_v1_retrieve_resolves_marker_hash() -> None:
"""End-to-end issue #389 reproducer:
1. POST a 200-item tool message to /v1/compress.
2. Parse the `<<ccr:HASH N_rows_offloaded>>` marker out of the
compressed messages.
3. GET /v1/retrieve/{hash} — must return 200 with the original.
Before the fix this returns 404 because the Rust crusher's marker
points at a hash the Python compression_store never received.
"""
from fastapi.testclient import TestClient
from headroom.cache.compression_store import reset_compression_store
from headroom.proxy.server import ProxyConfig, create_app
reset_compression_store()
config = ProxyConfig(
optimize=True,
cache_enabled=False,
rate_limit_enabled=False,
cost_tracking_enabled=False,
)
app = create_app(config)
# Build a payload similar to the issue's reproducer — 200 items
# with enough variation to trigger the lossy path. The Rust
# crusher's adaptive_k will keep ~15 and drop the rest.
#
# The blob is unique-per-item and long relative to the key names so
# the lossless Table/CSV path (which wins by stripping repeated keys
# when it saves >= lossless_min_savings_ratio) cannot clear the bar —
# this test exists to exercise the LOSSY row-drop path and its
# Rust -> Python CCR store bridge.
items = [
{
"id": i,
"score": 0.99 if i % 30 == 0 else 0.6,
"msg": f"Result {i:03d}{' error' if i % 30 == 0 else ' ok'}",
"blob": f"payload-{i:04d}-" + "".join(chr(97 + (i * 7 + j) % 26) for j in range(240)),
}
for i in range(200)
]
req = {
"model": "gpt-4o",
# /v1/compress is marker-free by default (no gateway caller can resolve a
# marker); mode="ccr" is the opt-in for callers that run the retrieve loop.
"config": {"mode": "ccr"},
"messages": [
{"role": "user", "content": "Get items"},
{
"role": "assistant",
"content": None,
"tool_calls": [
{
"id": "c1",
"type": "function",
"function": {"name": "get", "arguments": "{}"},
}
],
},
{"role": "tool", "tool_call_id": "c1", "content": json.dumps(items)},
],
}
try:
with TestClient(app, base_url="http://127.0.0.1", client=("127.0.0.1", 12345)) as client:
resp = client.post("/v1/compress", json=req)
assert resp.status_code == 200, resp.text
body = resp.json()
# The compressed messages should embed at least one CCR marker.
messages_blob = json.dumps(body["messages"])
assert "<<ccr:" in messages_blob, (
f"no CCR marker after /v1/compress on a 200-item array; "
f"compression didn't fire as expected: "
f"transforms={body.get('transforms_applied')}"
)
# Pull the hash with a substring scan (no regex).
start = messages_blob.find("<<ccr:") + len("<<ccr:")
end = start
while end < len(messages_blob) and messages_blob[end] in ("0123456789abcdef"):
end += 1
ccr_hash = messages_blob[start:end]
assert ccr_hash, "couldn't extract hash from marker"
# The retrieve stats endpoint should now show ≥ 1 entry.
stats_resp = client.get("/v1/retrieve/stats")
assert stats_resp.status_code == 200
stats = stats_resp.json()
assert stats["store"]["entry_count"] >= 1, (
f"compression_store empty after /v1/compress; stats={stats!r}"
)
# The actual /v1/retrieve call from the issue:
retrieve_resp = client.post("/v1/retrieve", json={"hash": ccr_hash})
assert retrieve_resp.status_code == 200, (
f"/v1/retrieve for marker hash {ccr_hash!r} returned "
f"{retrieve_resp.status_code} ({retrieve_resp.text}). "
f"The Rust→Python store bridge dropped the entry."
)
retrieve_body = retrieve_resp.json()
assert retrieve_body["hash"] == ccr_hash
# The retrieved content should parse back to a JSON array
# of the original items.
retrieved_items = json.loads(retrieve_body["original_content"])
assert isinstance(retrieved_items, list)
# The original was 200 items. The Rust hash is over the
# canonical-JSON form of the parsed input — should round-trip.
assert len(retrieved_items) == 200, (
f"expected 200 items in retrieved content, got {len(retrieved_items)}"
)
# Spot-check the first item.
assert retrieved_items[0]["id"] == 0
# And the GET shape (used by some clients) returns the same.
get_resp = client.get(f"/v1/retrieve/{ccr_hash}")
assert get_resp.status_code == 200
get_body = get_resp.json()
assert get_body["hash"] == ccr_hash
finally:
reset_compression_store()
@pytest.mark.skipif(not _HAS_FASTAPI, reason="fastapi not installed")
def test_v1_retrieve_unknown_hash_still_404() -> None:
"""Sanity: unknown hashes still return 404 (the bridge doesn't
accidentally make the store too permissive)."""
from fastapi.testclient import TestClient
from headroom.cache.compression_store import reset_compression_store
from headroom.proxy.server import ProxyConfig, create_app
reset_compression_store()
config = ProxyConfig(
optimize=False,
cache_enabled=False,
rate_limit_enabled=False,
cost_tracking_enabled=False,
)
app = create_app(config)
try:
with TestClient(app, base_url="http://127.0.0.1", client=("127.0.0.1", 12345)) as client:
resp = client.post("/v1/retrieve", json={"hash": "deadbeef0000"})
assert resp.status_code == 404
finally:
reset_compression_store()