Three audit follow-ups from issue #327's deep-dive review.
C1 — CompressionCache concurrency lock
======================================
`CompressionCache` instances are shared per `session_id` and accessed from
async-dispatched threadpool workers. Pre-fix, concurrent requests for the
same session raced on `_cache`, `_stable_hashes`, `_first_seen`, and
`_total_tokens_saved` with no synchronization. Observable failures:
* Lost-update on `_total_tokens_saved` (read-modify-write).
* `RuntimeError: OrderedDict mutated during iteration` from `apply_cached`
when a concurrent `store_compressed` evicts during the walk.
* Lost stable-hash records — next-turn compute_frozen_count reads
inconsistent state.
May also explain part of SvenMeyer's `_cache: 0 entries / 1003 misses`
observation: the cache was being clobbered concurrently.
Added `threading.RLock` guarding all mutating methods. `RLock` (not `Lock`)
so future code can call locked methods from inside another locked method
without self-deadlock. Also locked `HeadroomProxy._compression_caches`
dict-of-caches access via a separate `_compression_caches_lock` so two
concurrent calls for the same session_id can't each create distinct
CompressionCache objects (which would split the cache state between them).
The `/stats` endpoint snapshots the cache list under the dict lock before
iterating to avoid eviction-during-iteration.
C2 — Multi-worker CCR fragmentation: documented + startup warning
=================================================================
The in-memory `InMemoryCcrStore` (Rust), `_compression_caches` (Python),
`session_tracker_store` (Python), and TOIN learner state are ALL
per-process. Multi-worker uvicorn round-robins requests across workers,
so a session whose turn-1 lands on worker A may have turn-2 land on
worker B. Worker B has zero knowledge of A's CCR markers, replay cache,
or prefix-cache state. Result: `Retrieve original: hash=X` markers stay
in-context as opaque directives, every fresh tool_result is recompressed
from scratch, and `frozen_message_count=0` causes Anthropic prefix-cache
busts on every cross-worker turn.
Added a "Multi-worker deployment — CCR fragmentation" section in
`RUST_DEV.md` documenting the failure modes, the supported configuration
(`--workers 1`), and the sticky-session workaround for horizontal scale.
The proxy emits a `WARNING`-level log line on startup if `workers > 1` is
detected, pointing at the doc section.
C3 — Bounded compression executor with cancel-aware metrics
===========================================================
`asyncio.wait_for(asyncio.to_thread(pipeline.apply), timeout=...)`
cancellation does NOT propagate into the threadpool worker that's running
Rust code. Once the worker has picked up the task,
`concurrent.futures.Future.cancel()` returns False and the thread runs to
completion. Stuck threads accumulated invisibly on asyncio's default
executor, contending with unrelated `to_thread` callers (file IO, etc.).
Replaced all 7 `asyncio.to_thread` call sites for `pipeline.apply()`
across `proxy/handlers/anthropic.py` (3) and `proxy/handlers/openai.py` (4)
with a new `HeadroomProxy._run_compression_in_executor(fn, *, timeout)`
helper that:
1. Submits to a dedicated bounded `ThreadPoolExecutor` named
`headroom-compress` (configurable via
`ProxyConfig.compression_max_workers`; defaults to
`min(32, (cpu_count or 1) * 4)`).
2. Increments `_compression_in_flight` (gauge) when work starts and
decrements when work completes; tracks `_compression_in_flight_max`
as a high-water mark.
3. Detects "leaked threads" by comparing wall-clock elapsed against the
timeout in the worker's `finally` block. Increments
`_compression_leaked_threads` when a worker finishes after its
asyncio future was cancelled. Operators can see the leaked-thread
rate climbing in `/stats runtime.compression_executor` BEFORE the
pool fills up.
Tests
=====
* `TestCompressionCacheConcurrency` (3 tests) — many threads
store_compressed / apply_cached / update_from_result on a single
CompressionCache; assert no exceptions, no lost updates, no partial
state.
* `test_get_compression_cache_returns_same_instance_under_contention` —
32 concurrent `_get_compression_cache(same_id)` calls return the
identical instance (would split pre-lock).
* `test_proxy_compression_executor.py` (8 tests) — pool size respects
config, in-flight gauge tracks running compressions, high-water mark
is monotonic, timeout propagates to awaiter, leaked-thread counter
increments on post-deadline completion, `/stats` surfaces all three
gauges.
Verification
============
* All 123 targeted regression tests pass.
* `make ci-precheck` clean.
* No `Co-Authored-By` trailer; conventional `fix:` prefix; no
`--no-verify`.