headroom/RUST_DEV.md
chopratejas ea78cf6252 fix(proxy): cache concurrency lock, multi-worker docs, bounded compression executor
Three audit follow-ups from issue #327's deep-dive review.

C1 — CompressionCache concurrency lock
======================================

`CompressionCache` instances are shared per `session_id` and accessed from
async-dispatched threadpool workers. Pre-fix, concurrent requests for the
same session raced on `_cache`, `_stable_hashes`, `_first_seen`, and
`_total_tokens_saved` with no synchronization. Observable failures:

* Lost-update on `_total_tokens_saved` (read-modify-write).
* `RuntimeError: OrderedDict mutated during iteration` from `apply_cached`
  when a concurrent `store_compressed` evicts during the walk.
* Lost stable-hash records — next-turn compute_frozen_count reads
  inconsistent state.

May also explain part of SvenMeyer's `_cache: 0 entries / 1003 misses`
observation: the cache was being clobbered concurrently.

Added `threading.RLock` guarding all mutating methods. `RLock` (not `Lock`)
so future code can call locked methods from inside another locked method
without self-deadlock. Also locked `HeadroomProxy._compression_caches`
dict-of-caches access via a separate `_compression_caches_lock` so two
concurrent calls for the same session_id can't each create distinct
CompressionCache objects (which would split the cache state between them).
The `/stats` endpoint snapshots the cache list under the dict lock before
iterating to avoid eviction-during-iteration.

C2 — Multi-worker CCR fragmentation: documented + startup warning
=================================================================

The in-memory `InMemoryCcrStore` (Rust), `_compression_caches` (Python),
`session_tracker_store` (Python), and TOIN learner state are ALL
per-process. Multi-worker uvicorn round-robins requests across workers,
so a session whose turn-1 lands on worker A may have turn-2 land on
worker B. Worker B has zero knowledge of A's CCR markers, replay cache,
or prefix-cache state. Result: `Retrieve original: hash=X` markers stay
in-context as opaque directives, every fresh tool_result is recompressed
from scratch, and `frozen_message_count=0` causes Anthropic prefix-cache
busts on every cross-worker turn.

Added a "Multi-worker deployment — CCR fragmentation" section in
`RUST_DEV.md` documenting the failure modes, the supported configuration
(`--workers 1`), and the sticky-session workaround for horizontal scale.
The proxy emits a `WARNING`-level log line on startup if `workers > 1` is
detected, pointing at the doc section.

C3 — Bounded compression executor with cancel-aware metrics
===========================================================

`asyncio.wait_for(asyncio.to_thread(pipeline.apply), timeout=...)`
cancellation does NOT propagate into the threadpool worker that's running
Rust code. Once the worker has picked up the task,
`concurrent.futures.Future.cancel()` returns False and the thread runs to
completion. Stuck threads accumulated invisibly on asyncio's default
executor, contending with unrelated `to_thread` callers (file IO, etc.).

Replaced all 7 `asyncio.to_thread` call sites for `pipeline.apply()`
across `proxy/handlers/anthropic.py` (3) and `proxy/handlers/openai.py` (4)
with a new `HeadroomProxy._run_compression_in_executor(fn, *, timeout)`
helper that:

  1. Submits to a dedicated bounded `ThreadPoolExecutor` named
     `headroom-compress` (configurable via
     `ProxyConfig.compression_max_workers`; defaults to
     `min(32, (cpu_count or 1) * 4)`).
  2. Increments `_compression_in_flight` (gauge) when work starts and
     decrements when work completes; tracks `_compression_in_flight_max`
     as a high-water mark.
  3. Detects "leaked threads" by comparing wall-clock elapsed against the
     timeout in the worker's `finally` block. Increments
     `_compression_leaked_threads` when a worker finishes after its
     asyncio future was cancelled. Operators can see the leaked-thread
     rate climbing in `/stats runtime.compression_executor` BEFORE the
     pool fills up.

Tests
=====

* `TestCompressionCacheConcurrency` (3 tests) — many threads
  store_compressed / apply_cached / update_from_result on a single
  CompressionCache; assert no exceptions, no lost updates, no partial
  state.
* `test_get_compression_cache_returns_same_instance_under_contention` —
  32 concurrent `_get_compression_cache(same_id)` calls return the
  identical instance (would split pre-lock).
* `test_proxy_compression_executor.py` (8 tests) — pool size respects
  config, in-flight gauge tracks running compressions, high-water mark
  is monotonic, timeout propagates to awaiter, leaked-thread counter
  increments on post-deadline completion, `/stats` surfaces all three
  gauges.

Verification
============

* All 123 targeted regression tests pass.
* `make ci-precheck` clean.
* No `Co-Authored-By` trailer; conventional `fix:` prefix; no
  `--no-verify`.
2026-05-01 15:25:18 -07:00

17 KiB

Headroom Rust Rewrite — Developer Guide

This document covers the Rust port of Headroom. It is the only new top-level doc created in Phase 0; longer-form design/plan writeups live elsewhere and are not versioned in this repo.

Workspace layout

Cargo.toml                       # workspace root
rust-toolchain.toml              # pins stable rustc with rustfmt+clippy
crates/
  headroom-core/                 # library: shared types + transform trait surface
  headroom-proxy/                # binary: axum /healthz (Phase 2 grows this)
  headroom-py/                   # PyO3 cdylib exposing `headroom._core`
  headroom-parity/               # lib + `parity-run` CLI for Python parity tests
tests/parity/
  fixtures/<transform>/*.json    # recorded Python outputs (Phase 1 ports match)
  recorder.py                    # Python-side fixture recorder
scripts/record_fixtures.py       # entry point for running the recorder

cargo build --workspace builds every crate. default-members drops headroom-py from cargo run/bare-cargo test flows so that cargo test --workspace does not try to execute the PyO3 cdylib standalone (it can't find libpython without a Python interpreter hosting it).

Common commands

just is not installed on dev boxes here; a Makefile at the repo root exposes the same targets:

Target What it does
make test cargo test --workspace
make test-parity Builds headroom-py via maturin, runs parity-run run
make bench cargo bench --workspace
make build-proxy Release-builds headroom-proxy, strips, prints size
make build-wheel maturin build --release -m crates/headroom-py/pyproject.toml
make fmt cargo fmt --all
make lint cargo fmt --check + cargo clippy --workspace -- -D warnings

Running the proxy

headroom-proxy is a transparent reverse proxy. Phase 1 forwards HTTP/1.1, HTTP/2, SSE, and WebSocket traffic verbatim to a configured upstream — no provider logic yet. The intent is that operators run the existing Python proxy on a private port and put headroom-proxy on the public port pointed at it; end users notice nothing.

# Build
make build-proxy
./target/release/headroom-proxy --help

# Run against a local upstream
./target/release/headroom-proxy \
    --listen 0.0.0.0:8787 \
    --upstream http://127.0.0.1:8788

# Health checks
curl -s http://127.0.0.1:8787/healthz            # => {"ok":true,...}
curl -s http://127.0.0.1:8787/healthz/upstream   # => 200 if upstream reachable

Operator runbook (Phase 1 cutover)

# 1. Move the Python proxy to a private port (e.g. 8788)
HEADROOM_BIND=127.0.0.1:8788 python -m headroom.proxy &     # or your existing launcher

# 2. Run the Rust proxy on the previously-public port (8787) pointing at it
./target/release/headroom-proxy --listen 0.0.0.0:8787 --upstream http://127.0.0.1:8788 &

# 3. End users keep hitting :8787 unchanged.
# 4. Confirm passthrough:
curl -si http://127.0.0.1:8787/v1/models
# 5. Rollback = stop the Rust proxy and rebind Python back to 8787.

Configuration flags

Flag Env var Default Notes
--listen HEADROOM_PROXY_LISTEN 0.0.0.0:8787 bind address
--upstream HEADROOM_PROXY_UPSTREAM (required) base URL the proxy forwards to
--upstream-timeout 600s end-to-end request timeout (long for streams)
--upstream-connect-timeout 10s TCP/TLS connect timeout
--max-body-bytes 100MB for buffered cases; streams bypass
--log-level info RUST_LOG-style filter
--rewrite-host / --no-rewrite-host rewrite rewrite Host to upstream (default)
--graceful-shutdown-timeout 30s wait for in-flight on SIGTERM/SIGINT

Picking the next port: invocation telemetry

Before porting another Python compressor to Rust, check what's actually running. The Python proxy already exposes per-transform telemetry on /stats (headroom.proxy.prometheus_metrics):

# Top compressors by invocation count (last process lifetime)
curl -s http://127.0.0.1:8788/stats | jq '.compressions_by_strategy'
# {
#   "intelligent_context": 12453,
#   "smart_crusher": 487,
#   "search":         312,
#   "diff":            28,
#   "code":             0,        # ← never fires; safe to defer porting
#   ...
# }

# Per-transform timing (avg/max/count by transform name)
curl -s http://127.0.0.1:8788/stats | jq '.pipeline_timing'

# Token savings attributable to each strategy
curl -s http://127.0.0.1:8788/stats | jq '.tokens_saved_by_strategy'

This is the data the audit-cleanup PR (2026-04-30) recommended for prioritizing the next Python → Rust port. Strategies with zero or near-zero invocations are deferral candidates; strategies on the hot path are porting candidates regardless of LOC count.

Reserved paths

/healthz and /healthz/upstream are intercepted by the Rust proxy and not forwarded. Operators must not name a real upstream route either of these. Everything else is a catch-all forward.

Maturin + Python wiring

headroom-py is a PyO3 cdylib that exposes headroom._core in Python. The extension-module feature is opt-in so plain cargo build --workspace does not try to link against libpython on systems that don't have it.

python3.11 -m venv /tmp/hr-rust-venv
source /tmp/hr-rust-venv/bin/activate
pip install maturin
cd crates/headroom-py
maturin develop           # editable dev build, installs headroom._core
cd /tmp                   # IMPORTANT: step out of the repo root first
python -c "from headroom._core import hello; print(hello())"
# => headroom-core

Why cd /tmp? The repo root also contains the Python headroom/ package. Running the smoke import from the repo root makes Python resolve headroom to ./headroom/__init__.py (the full SDK, which pulls in heavy deps) instead of the lightweight namespace package installed by maturin. Tests should either run outside the repo root, or ensure headroom is installed into the same venv (then the maturin-installed _core.so lands alongside it and both imports resolve).

Release wheels

make build-wheel
# wheels land under target/wheels/

CI (.github/workflows/rust.yml) builds linux-x86_64, macos-arm64, and macos-x86_64 wheels via PyO3/maturin-action and uploads them as artifacts.

Parity harness

crates/headroom-parity owns the Rust-vs-Python oracle:

  • JSON fixtures under tests/parity/fixtures/<transform>/ (schema: { transform, input, config, output, recorded_at, input_sha256 }).
  • TransformComparator trait — one impl per transform. Phase 0 stubs return Err(...); the harness flags those as Skipped, not panics.
  • parity-run CLI: cargo run -p headroom-parity -- run [--only TRANSFORM].
  • Unit tests in crates/headroom-parity/src/lib.rs include a negative test (harness_reports_diff_for_divergent_comparator) proving the harness detects mismatched output before any real port lands.

Recording fresh fixtures

source .venv/bin/activate           # the main Python SDK venv
python scripts/record_fixtures.py   # uses tests/parity/recorder.py
ls tests/parity/fixtures/*/ | sort | uniq -c

The recorder monkey-patches the in-process transform classes (see record_all() in tests/parity/recorder.py). It does not modify any file under headroom/.

Known regressions in retired-Python components

The Stage 3b/3c.1b retirements deleted Python source for DiffCompressor and SmartCrusher and replaced them with PyO3-delegating shims. The 2026-04-28 audit found that the retirements shipped with subsystems silently disconnected. This section tracks each gap and its disposition so they don't regress further or get forgotten.

SmartCrusher

Subsystem State Tracked by
TOIN learning loop Re-attached 2026-04-28. Shim's crush() and _smart_crush_content() now call toin.record_compression() after a real compression. Filtered on strategy != "passthrough" to ignore JSON re-canonicalization. Best-effort: TOIN failures are logged at debug level and don't break compression. tests/test_smart_crusher_toin_attachment.py
CCR marker emission knob Honored end-to-end 2026-04-29. New enable_ccr_marker: bool field on Rust SmartCrusherConfig; crush_array checks it before emitting the <<ccr:HASH>> marker text and the CCR store write. Python shim flips it from ccr_config.enabled and ccr_config.inject_retrieval_marker — both flags collapse to the same Rust gate, since storing payloads under either off-switch makes no sense. Scope: gates only the row-drop sentinel path; Stage-3c.2 opaque-string CCR substitutions still emit always (no Python equivalent, no production caller asks for suppression). tests/test_smart_crusher_toin_attachment.py + crates/headroom-core/.../crusher.rs::tests::enable_ccr_marker_*
Custom relevance scorer Closed (fail-loud) 2026-04-29. relevance_config and scorer constructor args remain in the signature for source compat, but the shim raises NotImplementedError when either is non-None — silently dropping a user-supplied scorer is a textbook silent-fallback bug. Full plumbing waits on Stage-3c.2's relevance-crate Python bridge. tests/test_smart_crusher_toin_attachment.py::test_custom_*_arg_raises_not_implemented
Per-tool TOIN learning hook Re-attached partially. _smart_crush_content accepts tool_name and now threads it into the TOIN record. The hook is best-effort — it improves query_context aggregation but doesn't drive per-tool overrides yet. tests/test_smart_crusher_toin_attachment.py::test_smart_crush_content_records_to_toin

DiffCompressor

Subsystem State
Adaptive context windows Honored byte-for-byte (parity fixture-locked).
TOIN integration Never had one — DiffCompressor records via _record_to_toin in ContentRouter, which already runs for non-SmartCrusher strategies. No regression.

Phase 3e.1 — signals/ trait module + KeywordDetector (2026-04-29)

The Python error_detection.py regex registry was retired and reborn as a trait + tier system in crates/headroom-core/src/signals/. See signals/README.md for the full architecture; the highlights:

  • Per-granularity traits. LineImportanceDetector ships today; future ContentTypeDetector and ItemImportanceDetector<I> will follow as their consumers get touched.
  • Tiered<T> combinator. Composition, not inheritance. Future ML detectors slot in as new tiers without changes to KeywordDetector or any caller.
  • One concrete impl. KeywordDetector (aho-corasick) is the only tier registered today. No NoOp/stub impls — per project no-silent-fallbacks rule, future tiers land with their real implementations.
  • Bug fixes baked in. ERROR_KEYWORDS regex now includes timeout|abort|denied|rejected (previously drifted from the keyword set); token dropped from SECURITY_KEYWORDS (false-positived on every LLM metric reference). Both fixed in the Python regex too via the shim that recompiles patterns from the Rust-exposed keyword tables.
  • Companion canonical extension path. signals/README.md documents the BGE classifier head — a 384-dim → 4-class softmax on top of the already-loaded bge-small-en-v1.5 embedder — as the natural ML tier. Two alternatives kept open: distilled tinyBERT in ONNX, logistic regression on lexical features.

Phase 3g (queued) — Compression Pipeline Formalization (issue #315)

Strategic decision 2026-04-29: after Phase 3e (compressor ports) and Phase 3f (Rust MCP scaffold) wrap, formalize the lossless-then-lossy- then-CCR ordering as a cross-cutting CompressionPipeline orchestrator

  • LosslessTransform / LossyTransform traits in crates/headroom-core/src/pipeline/. Existing compressors get refactored as compositions of pluggable transforms. The crucial design choice — parsers for structure, models at the prose/structure boundary — is captured in issue #315 and memory/project_lossless_first_pipeline.md. Do NOT start coding before 3e/3f finish.

Watch list (potential regressions, not yet audited)

  • CCRConfig.enabled=False end-to-end — closed 2026-04-29. Both enabled=False and inject_retrieval_marker=False collapse to the same Rust enable_ccr_marker=False gate (no marker, no store write). See the SmartCrusher table above.
  • SmartCrusherConfig.use_feedback_hints=False — config field is forwarded to Rust but its honoring inside the Rust crusher hasn't been verified against a parity fixture for the disabled path.

When any item above changes, update both this section and the test file. The shim's docstring also references this section — keep them aligned.

Phase 0 Blockers

These are known limitations for Phase 0. They are tracked here so Phase 1 doesn't rediscover them.

  • cache_aligner fixtures: CacheAligner.apply() takes (messages, tokenizer, **kwargs) — a Tokenizer is provider-specific and its cheapest NoopTokenCounter / TiktokenTokenCounter construction still requires pulling headroom.providers.* which imports the full observability stack (opentelemetry, etc). The recorder records cache_aligner only if a usable tokenizer is cheaply available; otherwise it logs a blocker and skips. See recorder.py::_build_cache_aligner_tokenizer.
  • ccr is not a single class: The repo has CCRToolInjector, CCRResponseHandler, CCRToolCall, CCRToolResult etc. rather than a single CCR class. The recorder targets the encoder-style entry point most analogous to the Rust port (CCRToolInjector.inject_tool and CCRResponseHandler.parse_response). If Phase 1 wants a different split it should update recorder.py::record_all accordingly.
  • Pre-commit hook noise: scripts/sync-plugin-versions.py mutates .claude-plugin/marketplace.json, .github/plugin/marketplace.json, and plugins/headroom-agent-hooks/**/plugin.json on every commit. Those changes are harmless but each commit in Phase 0 picks them up. Phase 1 does not need to do anything special — just let the hook run.
  • rust-toolchain.toml pins channel = "stable" rather than a specific version so CI picks up the same toolchain the local box uses. Tighten to a pinned version (e.g. 1.78) once the port stabilizes.

Multi-worker deployment — CCR fragmentation

Recommendation: run the proxy with --workers 1 (the default). The in-memory CCR store is per-process. Multi-worker uvicorn deployments produce silent retrieval failures — and Compress-Cache-Retrieve is the lossless half of the pipeline.

What goes wrong with --workers N > 1

Each uvicorn worker is a separate Python process. Each process holds its own copies of:

  1. InMemoryCcrStore (crates/headroom-core/src/ccr.rs:78) — the sharded DashMap mapping hash → original_content for content the compressor replaced with Retrieve original: hash=X markers.
  2. HeadroomProxy._compression_caches (headroom/proxy/server.py:367) — the per-session CompressionCache dict.
  3. HeadroomProxy.session_tracker_store — per-session prefix-tracker state derived from Anthropic's cache_read_input_tokens responses.
  4. TOIN learner state — pattern statistics used to bias the compressor.

When uvicorn round-robins requests across workers, a session whose turn-1 landed on worker A may have turn-2 land on worker B. Worker B has zero knowledge of what worker A did:

  • The CCR marker Retrieve original: hash=X is in the conversation, but worker B's InMemoryCcrStore returns None for X. The marker stays in-context as an opaque directive the model can't act on. Tokens spent, no retrieval value.
  • The CompressionCache on worker B has no replay entries → every fresh tool_result is recompressed from scratch, even content that worker A already compressed once. CPU wasted; observable as compression latency doubling on round-robin sessions.
  • The prefix_tracker on worker B starts at frozen_message_count = 0 → worker B compresses positions that Anthropic has already cached. Cache bust + write premium paid on every cross-worker session turn.

What works today

--workers 1 is the only fully-supported configuration. The proxy is async and a single worker handles thousands of concurrent requests via the event loop; CPU-bound Rust work releases the GIL via py.allow_threads, so vertical scaling on one process is the intended path.

What we'd need for multi-worker

A backend implementation of the CcrStore trait (crates/headroom-core/src/ccr.rs) backed by a shared store — Redis, Memcached, or a sticky-session reverse proxy. The _compression_caches and session_tracker_store would also need shared backing. None of this is implemented yet. If you need horizontal scale today, run multiple single-worker proxy processes behind a sticky-session load balancer (hash on session_id) — that pins each session to one worker and avoids the fragmentation entirely.

Detecting it in the wild

The proxy emits a WARNING-level log line on startup if it detects WEB_CONCURRENCY or uvicorn --workers set to anything > 1, pointing operators at this section.