mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
Phase B step 1 of the live-zone-only realignment. Removes ~10K LOC of
"drop messages from history" machinery that became unreachable after
PR-A1 made `/v1/messages` a passthrough on the proxy. Live-zone-only
compression (PR-B2..B7) operates on content blocks within messages;
message-list mutation no longer happens in the pipeline.
Python deletes:
- headroom/transforms/intelligent_context.py (1077 LOC)
- headroom/transforms/rolling_window.py (395 LOC)
- headroom/transforms/progressive_summarizer.py (508 LOC)
- headroom/transforms/scoring.py (459 LOC)
- headroom/transforms/tool_crusher.py (338 LOC)
- 5 corresponding tests/test_transforms/* and tests/test_proxy_intelligent_context.py
Rust deletes:
- crates/headroom-core/src/context/* (manager, config, workspace,
candidate, ccr_drop, strategy/, mod) + safety.rs replaced
- crates/headroom-core/src/scoring/* (mod, score, scorer, traits, weights)
- MessageScorerComparator from crates/headroom-parity (PR #338/#343
becomes deletable; sunk cost stays sunk)
- 13 message_scorer fixtures + record_message_scorer.py
Rust adds (move + rewrite):
- crates/headroom-core/src/transforms/safety.rs — `tool_pair_indices`
preserves the OpenAI/Anthropic tool_use ↔ tool_result pairing rule
the live-zone dispatcher (PR-B2) needs. No IcmConfig dependency.
Surface refactors:
- HeadroomConfig: drop `tool_crusher`, `rolling_window`,
`intelligent_context` fields; hoist `output_buffer_tokens` to top
level (used by client.py).
- ProxyConfig: drop `intelligent_context*` fields.
- `headroom wrap` proxy server: retire IntelligentContextManager
and RollingWindow imports + branch; pipeline is CacheAligner →
ContentRouter (smart_routing) or CacheAligner → SmartCrusher
(legacy).
- CLI: drop `--no-intelligent-context`, `--no-intelligent-scoring`,
`--no-compress-first` flags.
- LangChain memory integration: rename `_apply_rolling_window` →
`_apply_compression`, drop RollingWindowConfig dep. Threshold is
now advisory — B6 will rework the contract.
- TransformPipeline.create_pipeline now takes only cache_aligner_config.
- headroom/__init__.py + headroom/transforms/__init__.py: strip
exports of deleted symbols.
Bug fixes uncovered by full pytest sweep:
- providers/copilot/wrap.py: `environ or os.environ` collapsed
empty-dict to falsy → callers passing `environ={}` accidentally
pulled from os.environ. Use `environ if environ is not None else
os.environ`.
Test correctness fixes:
- _DummyAnthropicHandler._retry_request gains **_kwargs to match
the real handler signature post-A8.
- test_ws_http_fallback extracts JSON from `content=` (post-A3
byte-faithful) rather than the obsolete `json=` kwarg.
- test_ccr_response_handler_extra fixture joins SSE events with
`\n\n` per spec (post-A8 byte-buffer parser requirement).
- test_proxy_responses_phase_preservation: capture via direct
handler attached to the named logger, so the assertion is
order-independent (proxy `_setup_file_logging` flips
`headroom.propagate=False` once any earlier test triggers it).
- conftest.py autouse fixture resets `headroom.propagate=True`
before each test as a defensive measure for the same pollution.
- test_wrap_copilot_translated_backend_still_requires_byok:
monkeypatch.delenv every provider key so the BYOK error
actually fires.
- test_native_installers: skip when system bash < 4.3 (macOS ships 3.2).
- TestGeminiEmbedContent / TestGeminiBatchEmbedContents:
pytest.mark.skip — proxy currently has no :embedContent route;
feature gap, not regression.
Acceptance:
- cargo build --workspace + cargo clippy + cargo fmt --check: green.
- cargo test --workspace --exclude headroom-py: 777 passed.
- pytest: 4892 passed, 240 skipped, 0 failed.
- git grep returns only intentional comments referencing the deletion.
Per-PR-B1 plan: REALIGNMENT/04-phase-B-live-zone.md.
348 lines
13 KiB
Python
348 lines
13 KiB
Python
"""Per-strategy compression observability tests.
|
|
|
|
These guard the forcing function: when any compressor runs in
|
|
production, a `CompressionObserver` notification fires once per real
|
|
compression event, and `PrometheusMetrics` accumulates per-strategy
|
|
counters that the test suite asserts on directly.
|
|
|
|
The TOIN→SmartCrusher silent disconnect (caught three weeks late by
|
|
manual audit) was invisible because no signal distinguished by
|
|
strategy. These tests exist so the next regression of that shape
|
|
fails the suite the day it lands instead of waiting on an audit.
|
|
|
|
The counters live ONLY as in-process state on the metrics instance;
|
|
they are deliberately NOT exported through the Prometheus scrape or
|
|
OTel surface, because the metric→Supabase pipeline treats each
|
|
metric name as a column and we cannot add new columns. CI-level
|
|
observability via these tests is enough to catch silent regressions;
|
|
production export waits on a non-column-adding pipeline.
|
|
|
|
Coverage:
|
|
|
|
1. `ContentRouter.compress(...)` calls observer once per RoutingDecision.
|
|
2. `SmartCrusher.apply(...)` calls observer once per crushed message.
|
|
3. Both transforms tolerate an observer that raises (compression must
|
|
still succeed).
|
|
4. `PrometheusMetrics` correctly satisfies the `CompressionObserver`
|
|
protocol — `record_compression` increments per-strategy counters
|
|
and `tokens_saved_by_strategy` accumulates only positive savings.
|
|
5. The Prometheus scrape output (`export()`) does NOT emit any new
|
|
metric names — the per-strategy state stays internal.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from dataclasses import dataclass, field
|
|
from typing import Any
|
|
|
|
import pytest
|
|
|
|
from headroom.transforms.content_detector import ContentType
|
|
from headroom.transforms.content_router import (
|
|
CompressionStrategy,
|
|
ContentRouter,
|
|
ContentRouterConfig,
|
|
RouterCompressionResult,
|
|
RoutingDecision,
|
|
)
|
|
from headroom.transforms.observability import CompressionObserver
|
|
from headroom.transforms.smart_crusher import SmartCrusher, SmartCrusherConfig
|
|
|
|
# ─── Test doubles ──────────────────────────────────────────────────────
|
|
|
|
|
|
@dataclass
|
|
class SpyObserver:
|
|
"""Captures every `record_compression` call for assertion."""
|
|
|
|
calls: list[tuple[str, int, int]] = field(default_factory=list)
|
|
|
|
def record_compression(
|
|
self,
|
|
strategy: str,
|
|
original_tokens: int,
|
|
compressed_tokens: int,
|
|
) -> None:
|
|
self.calls.append((strategy, original_tokens, compressed_tokens))
|
|
|
|
|
|
@dataclass
|
|
class ExplodingObserver:
|
|
"""Raises on every call. Used to assert observer failures don't
|
|
propagate out and break compression."""
|
|
|
|
raised: int = 0
|
|
|
|
def record_compression(self, *_a: Any, **_kw: Any) -> None:
|
|
self.raised += 1
|
|
raise RuntimeError("simulated observer outage")
|
|
|
|
|
|
# ─── Protocol conformance ──────────────────────────────────────────────
|
|
|
|
|
|
def test_spy_satisfies_observer_protocol():
|
|
spy = SpyObserver()
|
|
# `runtime_checkable` Protocol — isinstance check works.
|
|
assert isinstance(spy, CompressionObserver)
|
|
|
|
|
|
def test_prometheus_metrics_satisfies_observer_protocol():
|
|
from headroom.proxy.prometheus_metrics import PrometheusMetrics
|
|
|
|
m = PrometheusMetrics()
|
|
assert isinstance(m, CompressionObserver)
|
|
|
|
|
|
# ─── ContentRouter wiring ──────────────────────────────────────────────
|
|
|
|
|
|
def test_content_router_records_observer_call_per_routing_decision():
|
|
spy = SpyObserver()
|
|
router = ContentRouter(ContentRouterConfig(), observer=spy)
|
|
|
|
# Forge a routing log directly via the result object — the observer
|
|
# call site walks `result.routing_log`, so we assert the contract
|
|
# without depending on which compressor would actually fire.
|
|
result = RouterCompressionResult(
|
|
compressed="x",
|
|
original="x",
|
|
strategy_used=CompressionStrategy.SMART_CRUSHER,
|
|
routing_log=[
|
|
RoutingDecision(
|
|
content_type=ContentType.JSON_ARRAY,
|
|
strategy=CompressionStrategy.SMART_CRUSHER,
|
|
original_tokens=200,
|
|
compressed_tokens=50,
|
|
),
|
|
RoutingDecision(
|
|
content_type=ContentType.SOURCE_CODE,
|
|
strategy=CompressionStrategy.CODE_AWARE,
|
|
original_tokens=300,
|
|
compressed_tokens=300, # passthrough — still recorded
|
|
),
|
|
],
|
|
)
|
|
router._observe(result)
|
|
|
|
assert spy.calls == [
|
|
("smart_crusher", 200, 50),
|
|
("code_aware", 300, 300),
|
|
]
|
|
|
|
|
|
def test_content_router_with_no_observer_is_silent():
|
|
router = ContentRouter(ContentRouterConfig()) # observer defaults None
|
|
result = RouterCompressionResult(
|
|
compressed="x",
|
|
original="x",
|
|
strategy_used=CompressionStrategy.PASSTHROUGH,
|
|
routing_log=[
|
|
RoutingDecision(
|
|
content_type=ContentType.PLAIN_TEXT,
|
|
strategy=CompressionStrategy.TEXT,
|
|
original_tokens=10,
|
|
compressed_tokens=5,
|
|
)
|
|
],
|
|
)
|
|
# Should not raise.
|
|
router._observe(result)
|
|
|
|
|
|
def test_content_router_swallows_observer_failures():
|
|
boom = ExplodingObserver()
|
|
router = ContentRouter(ContentRouterConfig(), observer=boom)
|
|
result = RouterCompressionResult(
|
|
compressed="x",
|
|
original="x",
|
|
strategy_used=CompressionStrategy.TEXT,
|
|
routing_log=[
|
|
RoutingDecision(
|
|
content_type=ContentType.PLAIN_TEXT,
|
|
strategy=CompressionStrategy.TEXT,
|
|
original_tokens=10,
|
|
compressed_tokens=5,
|
|
)
|
|
],
|
|
)
|
|
# Must not raise — observability failures are not compression failures.
|
|
router._observe(result)
|
|
assert boom.raised == 1
|
|
|
|
|
|
# ─── SmartCrusher wiring (legacy direct-pipeline path) ─────────────────
|
|
|
|
|
|
def _bigger_array(n: int = 60) -> str:
|
|
import json as _json
|
|
|
|
items = [{"status": "ok", "tag": "x", "n": i} for i in range(n)]
|
|
return _json.dumps(items)
|
|
|
|
|
|
@pytest.fixture
|
|
def isolated_toin(tmp_path, monkeypatch):
|
|
"""Point TOIN at a tempdir for the duration of the test.
|
|
|
|
SmartCrusher.apply() feeds the global TOIN learning store via
|
|
`record_compression`. Its default storage path is
|
|
`~/.headroom/toin.json`, which persists across pytest invocations.
|
|
On Python 3.11 CI runs the suite twice (regular + coverage); a
|
|
pattern written in run #1 changes which rows the lossy sampler
|
|
keeps in run #2 and breaks `test_first_last_items_always_preserved`
|
|
in `test_evals.py`.
|
|
|
|
Isolating the TOIN file per test contains the side effect.
|
|
"""
|
|
from pathlib import Path
|
|
|
|
from headroom.telemetry.toin import TOIN_PATH_ENV_VAR, reset_toin
|
|
|
|
storage = str(Path(tmp_path) / "toin.json")
|
|
monkeypatch.setenv(TOIN_PATH_ENV_VAR, storage)
|
|
reset_toin()
|
|
yield
|
|
reset_toin()
|
|
|
|
|
|
def test_smart_crusher_apply_records_observer_per_crushed_message(isolated_toin):
|
|
"""End-to-end: SmartCrusher.apply() walks messages, crushes the
|
|
big tool_result, fires the observer with strategy='smart_crusher'."""
|
|
from headroom.providers.openai import OpenAITokenCounter
|
|
from headroom.tokenizer import Tokenizer
|
|
|
|
spy = SpyObserver()
|
|
crusher = SmartCrusher(SmartCrusherConfig(), observer=spy)
|
|
tok = Tokenizer(OpenAITokenCounter("gpt-4o-mini"), model="gpt-4o-mini")
|
|
|
|
messages = [
|
|
{"role": "user", "content": "what's in the data?"},
|
|
{"role": "tool", "content": _bigger_array(60)},
|
|
]
|
|
result = crusher.apply(messages, tok)
|
|
# If the analyzer chose passthrough this run, the observer wasn't
|
|
# fired; that's fine for the wiring test — we only assert it WAS
|
|
# fired in the case it crushed.
|
|
if "smart_crush:" in ",".join(result.transforms_applied):
|
|
assert spy.calls, "smart_crusher crushed but observer wasn't notified"
|
|
for strategy, original, compressed in spy.calls:
|
|
assert strategy == "smart_crusher"
|
|
assert original > 0
|
|
assert compressed >= 0
|
|
|
|
|
|
def test_smart_crusher_apply_swallows_observer_failures(isolated_toin):
|
|
"""Observer raises → compression still completes, returns valid
|
|
TransformResult, count of raises matches the crushed_count."""
|
|
from headroom.providers.openai import OpenAITokenCounter
|
|
from headroom.tokenizer import Tokenizer
|
|
|
|
boom = ExplodingObserver()
|
|
crusher = SmartCrusher(SmartCrusherConfig(), observer=boom)
|
|
tok = Tokenizer(OpenAITokenCounter("gpt-4o-mini"), model="gpt-4o-mini")
|
|
messages = [{"role": "tool", "content": _bigger_array(60)}]
|
|
result = crusher.apply(messages, tok)
|
|
# Either the analyzer didn't crush (boom.raised == 0) or it did
|
|
# (boom.raised >= 1) — but in both cases compression returned a
|
|
# valid TransformResult. No exception escaped.
|
|
assert result.messages is not None
|
|
|
|
|
|
# ─── PrometheusMetrics implementation ──────────────────────────────────
|
|
|
|
|
|
def test_prometheus_metrics_accumulates_per_strategy_counters():
|
|
from headroom.proxy.prometheus_metrics import PrometheusMetrics
|
|
|
|
m = PrometheusMetrics()
|
|
|
|
m.record_compression("smart_crusher", original_tokens=200, compressed_tokens=50)
|
|
m.record_compression("smart_crusher", original_tokens=100, compressed_tokens=40)
|
|
m.record_compression("diff", original_tokens=80, compressed_tokens=80) # no savings
|
|
m.record_compression("code_aware", original_tokens=50, compressed_tokens=70) # negative savings
|
|
|
|
assert m.compressions_by_strategy == {
|
|
"smart_crusher": 2,
|
|
"diff": 1,
|
|
"code_aware": 1,
|
|
}
|
|
# Tokens saved is `max(0, original - compressed)` per strategy.
|
|
# smart_crusher: 150 + 60 = 210; diff: 0 (no savings, dict entry omitted);
|
|
# code_aware: 0 (negative).
|
|
assert m.tokens_saved_by_strategy == {"smart_crusher": 210}
|
|
|
|
|
|
def test_prometheus_export_does_not_leak_per_strategy_metrics():
|
|
"""Per-strategy state is tracked in-process only. The Prometheus
|
|
scrape output deliberately must NOT emit new metric names — the
|
|
metric→Supabase pipeline treats each metric name as a column, and
|
|
we cannot add new columns. This test guards that constraint: if a
|
|
future change adds the metric to the scrape, this fails and forces
|
|
a conscious decision."""
|
|
import asyncio
|
|
|
|
from headroom.proxy.prometheus_metrics import PrometheusMetrics
|
|
|
|
m = PrometheusMetrics()
|
|
m.record_compression("smart_crusher", original_tokens=200, compressed_tokens=50)
|
|
m.record_compression("diff", original_tokens=120, compressed_tokens=70)
|
|
|
|
output = asyncio.run(m.export())
|
|
|
|
assert "headroom_compressions_total" not in output
|
|
assert "headroom_tokens_saved_by_strategy_total" not in output
|
|
|
|
|
|
# ─── End-to-end smoke (router + metrics together) ──────────────────────
|
|
|
|
|
|
def test_router_with_prometheus_observer_increments_counters():
|
|
"""Plumbing test: a router wired to a real PrometheusMetrics
|
|
instance lights up the per-strategy counters as routing decisions
|
|
accumulate. This is the production wiring shape from
|
|
`headroom/proxy/server.py`."""
|
|
from headroom.proxy.prometheus_metrics import PrometheusMetrics
|
|
|
|
m = PrometheusMetrics()
|
|
router = ContentRouter(ContentRouterConfig(), observer=m)
|
|
|
|
fake_result = RouterCompressionResult(
|
|
compressed="x",
|
|
original="x",
|
|
strategy_used=CompressionStrategy.MIXED,
|
|
routing_log=[
|
|
RoutingDecision(
|
|
content_type=ContentType.JSON_ARRAY,
|
|
strategy=CompressionStrategy.SMART_CRUSHER,
|
|
original_tokens=300,
|
|
compressed_tokens=80,
|
|
),
|
|
RoutingDecision(
|
|
content_type=ContentType.SOURCE_CODE,
|
|
strategy=CompressionStrategy.CODE_AWARE,
|
|
original_tokens=200,
|
|
compressed_tokens=120,
|
|
),
|
|
RoutingDecision(
|
|
content_type=ContentType.JSON_ARRAY,
|
|
strategy=CompressionStrategy.SMART_CRUSHER,
|
|
original_tokens=100,
|
|
compressed_tokens=40,
|
|
),
|
|
],
|
|
)
|
|
router._observe(fake_result)
|
|
|
|
assert m.compressions_by_strategy == {"smart_crusher": 2, "code_aware": 1}
|
|
assert m.tokens_saved_by_strategy == {
|
|
"smart_crusher": (300 - 80) + (100 - 40), # 280
|
|
"code_aware": (200 - 120), # 80
|
|
}
|
|
|
|
|
|
# IntelligentContextManager observability tests retired with PR-B1 —
|
|
# the manager itself was deleted along with the message-dropping
|
|
# strategy. Inner-router observability is now exercised solely
|
|
# through ContentRouter, covered by
|
|
# `test_content_router_records_observer_call_per_routing_decision`.
|