mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
3 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7bfb1d7f38
|
fix(cache): stable session identity and per-conversation prefix trackers under agentic clients (#2193)
## Description Running headroom as the proxy for Claude Code destroys Anthropic prompt-cache reuse (#2085: ~4.4x cache-creation inflation, 2.5–3x net cost). Tracing live Claude Code traffic through the proxy shows **two independent session-identity defects**, both of which orphan or thrash the frozen-prefix state; this PR fixes both. ### Defect 1: `<system-reminder>` turns rotate the fallback session id mid-conversation Claude Code interleaves reminder turns into the history as actual `role:"system"` messages (hook output, skills lists, file-truncation notices). `compute_session_id` hashed **every** system message, so the id rotated each time a reminder landed. Live trace (subagent reading two 80KB files; sid changes exactly when the truncation reminder appears, and the tracker restarts at turn 0): ``` REQ#2 sid=68d4ee666990 nmsg=3 [0]SYSTEM<<top-level system>> [1]user [2]SYSTEM<<skills reminder>> REQ#3 sid=6944948c9fb2 nmsg=6 ... [5]SYSTEM<<Truncated: PARTIAL view ...>> <- id rotated ``` Everything keyed on the session id is orphaned at that moment: the prefix tracker (freeze never survives past a reminder-bearing turn), beta-header stickiness, the CCR and memory-tool registries, and the compression cache. **Fix:** hash only the **leading run** of system messages (everything before the first non-system turn) — the top-level system prompt on the Anthropic path (folded in as the synthetic first message), the conventional leading system message(s) on the OpenAI path. Stable for the life of a conversation; mid-history system turns are content, not identity. ### Defect 2: conversations sharing a (now stable) id thrash one tracker With ids stable, the fallback tuple `model + system prompt` is identical across every same-type parallel subagent (and any sessions reusing one system prompt) — all of them collapse onto one `PrefixCacheTracker`, and their interleaved histories cross-contaminate the freeze state: the forwarded prefix is byte-unstable on nearly every turn and the provider cache is re-written instead of read. Reproduced against the real code paths (script below): ``` 1) fallback session ids: A=3dc639aaf4f48fa1 B=3dc639aaf4f48fa1 -> COLLIDE=True 2) single conversation, legacy : stable prefix on 4/4 later turns, trackers=1 2) interleaved (subagents), legacy : stable prefix on 0/9 later turns, trackers=1 2) interleaved, lineage resolution : stable prefix on 8/8 later turns, trackers=2 ``` **Fix:** `SessionTrackerStore.resolve_tracker` — within a session id, reuse the tracker whose previous request messages are a prefix of the incoming history (client histories are append-only, so a conversation's next request always extends its previous one); a diverging or rewritten history (client-side compaction) starts a fresh lineage. Matching uses the repo's existing canonical cross-turn equivalence (`_canonicalize_for_prefix_compare`, the same one the cache-stable delta path uses) on the **original client bytes**, so moved cache breakpoints, string<->block sugar, transport annotations, or a tail-mutating `pre_compress` hook never read as a rewrite. Byte-identical histories (templated fan-outs before they diverge) intentionally share a tracker — their provider cache line is identical too. ### Both fixes together, on live Claude Code traffic (sonnet, 2 parallel Explore agents) ``` main conversation: sid=5b7e245a... one tracker, turns 0->4, id stable across reminders agents (collide): sid=2bdffc9e... -> lineage bare (alpha) turns 0->1->2 -> lineage "~1" (beta) turns 0->1->2 ``` Before: the agents' ids rotated per reminder (every tracker stuck at turn 0), and whenever they did share an id they thrashed one tracker (`0/9` stable prefixes in the repro). ### Why not key the session id on conversation content? Draft #1912 folds the first user turn into the fallback id; this change composes with it, but identity-level keying alone can't close #2085: identical first turns (templated fan-outs) still collide, and everything keyed on the session id rotates with it when the client rewrites history. The "session" (client/workspace grouping) and the "conversation" (positional cache lineage) are different identities; only the tracker holds positional per-turn state that thrashes under collision — beta stickiness is a monotone union and the compression cache is content-addressed — so lineage resolution lives one level below the session id and leaves the id semantics (and every other consumer) untouched. ## Changes Made - `headroom/cache/prefix_tracker.py`: - `compute_session_id`: harvest only the leading system run (defect 1). - `SessionTrackerStore.resolve_tracker`: conversation-lineage resolution (defect 2). First lineage lives under the bare session id — single-conversation sessions behave byte-identically to before; degrades to `get_or_create` when messages are absent or prefix freeze is disabled. - Lineages are capped per session id (`PrefixFreezeConfig.max_lineages_per_session`, default 32). **Over-cap conversations share one overflow tracker instead of evicting an established lineage** — any eviction policy degrades every conversation once the working set exceeds the cap (under round-robin the victim is always the conversation about to arrive), while overflow sharing degrades only the over-cap tail, to exactly the pre-lineage shared behavior; `0` disables lineage splitting. Chains are stored as structural snapshots that normalize `NaN` (`json.loads` accepts bare NaN, and `NaN != NaN` would read a byte-identical resend as a rewrite). Synthetic lineage keys use a `\x00` separator, which cannot appear in an HTTP header value, so they can never collide with a client-supplied `x-headroom-session-id`. - `headroom/proxy/handlers/anthropic.py`, `openai.py`: the session id and the lineage both derive from the **same original client bytes** (a turn-dependent hook rewrite can no longer rotate one without the other); anthropic folds in its synthetic system message so explicit-header clients with different system prompts stay separate. Plus a docstring correction in `streaming.py` that falsely claimed its coarse mid-turn key "mirrors" `compute_session_id`. - `tests/test_cache/test_prefix_tracker.py`: 24 new test cases — reminder-rotation regression; interleaved isolation + per-conversation turn state; identical-first-turn share-then-split; cache_control movement (3 cases); representation churn (string<->block sugar / streaming `index` / Bedrock cachePoint); rewritten history → fresh lineage (compacted / middle-edited / truncated); legacy no-messages / freeze-disabled / empty-canonical fallbacks; NaN-in-tool-payload stability; overflow sharing, established-lineages-survive-cap, and a cap+1 round-robin no-cliff guard; TTL cleanup; session-id-not-rotated-by-lineage guard. One existing test renamed (`uses_all_system_messages` → `distinguishes_leading_system_run`) to match the new contract. - Three SimpleNamespace stub stores in existing tests gained a `resolve_tracker` field (handlers call it unconditionally — a silent `hasattr` fallback would degrade to the pre-fix behavior with no signal). One of them is the cold-start fast-pass suite (#2073), which landed while this branch was in review. - `CHANGELOG.md` entry. ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) ## Testing - [x] Unit tests pass (`pytest`) — 11 failed, 8652 passed, 528 skipped in 4:37 (the 11 are pre-existing on unmodified `main` — verified by rerunning the same node ids on a clean checkout: gh-CLI/onnx/PID-reuse/deadline flakes and order-dependent cases, none touching session/cache/proxy paths) - [x] Linting passes (`ruff check .`) — All checks passed (ruff 0.15.17, CI-pinned; `ruff format --check .` clean) - [x] Type checking passes (`mypy headroom`) — Success: no issues found in 471 source files - [x] New tests added for new functionality — 24 test cases; the rotation/isolation/no-cliff ones fail on `main` - [x] Manual testing performed — live Claude Code end-to-end, below ### Test Output ```text $ python -m pytest tests/ -q 11 failed, 8652 passed, 528 skipped, 5857 warnings in 276.68s (0:04:36) # same 11 fail on unmodified main (env/order-dependent: test_wrap_claude_base_url pid-reuse, # copilot_auth gh-cli fallback, image_compression onnx, content_router deadline, rtk/output-shaper/dedup order flakes) $ python -m pytest tests/test_cache/test_prefix_tracker.py -q 63 passed $ uvx ruff@0.15.17 check . && uvx ruff@0.15.17 format --check . All checks passed! / 1208 files already formatted $ mypy headroom Success: no issues found in 471 source files $ python repro_2085.py 1) fallback session ids: A=3dc639aaf4f48fa1 B=3dc639aaf4f48fa1 -> COLLIDE=True 2) single conversation, legacy : stable prefix on 4/4 later turns, trackers=1 2) interleaved (subagents), legacy : stable prefix on 0/9 later turns, trackers=1 2) interleaved, lineage resolution : stable prefix on 8/8 later turns, trackers=2 ``` ## Real Behavior Proof - Environment: macOS arm64, Python 3.13, `uv sync --extra dev --extra proxy`; real Claude Code CLI pointed at the proxy via `ANTHROPIC_BASE_URL=http://127.0.0.1:8790`, real Anthropic backend. - Exact command / steps: ran Claude Code sessions that launch 2–3 parallel Explore subagents (each reading multi-KB JSON files, several tool-loop turns each), with an observability wrapper printing each request's resolved session id, tracker identity, and turn counter inside the proxy. - Observed result: on `main`, subagent session ids rotate on reminder-bearing turns (trackers permanently stuck at turn 0); when conversations do share an id they share one tracker whose turn counter interleaves all of them. On this branch: ids stable for the life of each conversation; colliding subagents resolve to separate lineages (`bare`, `~1`) with clean per-conversation turn progressions (trace above). Unit-level repro shows forwarded-prefix stability going 0/9 → 8/8 for the interleaved shape. - Not tested: reporter-scale cache-economics (his 4.4x needs his long-session workload against a paid backend); happy to coordinate with @RomanAlexanderW on a before/after — the number to watch is the cache-read ratio in Claude Code transcripts recovering toward ~96%. <details> <summary>repro_2085.py</summary> ```python """Repro for #2085: concurrent conversations sharing a fallback session id (same model + system prompt — e.g. a Claude Code session and its parallel subagents) collapse onto one PrefixCacheTracker and thrash its frozen-prefix state -> byte-unstable forwarded prefixes -> the provider prompt cache is re-written on nearly every call. Uses headroom's real code paths. Run from the repo root: python ../repro_2085.py """ from headroom.cache.prefix_tracker import PrefixFreezeConfig, SessionTrackerStore MODEL = "claude-sonnet-5" # Claude Code system prompt: long, static, identical across the main session # and every parallel subagent of the same type. SYSTEM = ("You are Claude Code, Anthropic's official CLI for Claude. " * 40)[:2000] def convo(name: str, turns: int) -> list[dict]: msgs = [{"role": "system", "content": SYSTEM}] for t in range(turns): msgs.append({"role": "user", "content": f"[{name}] user turn {t}: " + ("x" * 800)}) msgs.append( {"role": "assistant", "content": f"[{name}] tool_result {t}: " + ('{"data": 1}' * 200)} ) return msgs class _Req: # request stub: no x-headroom-session-id header headers: dict = {} # --- Part 1: identity collision (real derivation) ---------------------------- store = SessionTrackerStore(PrefixFreezeConfig()) id_a = store.compute_session_id(_Req(), MODEL, convo("A", 3)) id_b = store.compute_session_id(_Req(), MODEL, convo("B", 5)) print(f"1) fallback session ids: A={id_a} B={id_b} -> COLLIDE={id_a == id_b}") # --- Part 2: interleaved conversations thrash the freeze state --------------- def run(interleave: bool, lineage_resolution: bool) -> tuple[int, int, int]: store = SessionTrackerStore(PrefixFreezeConfig()) stable_turns = 0 later_turns = 0 seq = [] for t in range(1, 6): seq.append(("A", convo("A", t))) if interleave: seq.append(("B", convo("B", t))) for _name, msgs in seq: sid = store.compute_session_id(_Req(), MODEL, msgs) if lineage_resolution: tracker = store.resolve_tracker(sid, "anthropic", messages=msgs) else: tracker = store.get_or_create(sid, "anthropic") if tracker._turn_number > 0: later_turns += 1 if tracker._forwarded_prefix_stable(msgs): stable_turns += 1 tracker.update_from_response( cache_read_tokens=5000 * len(msgs), cache_write_tokens=2000, messages=msgs, ) return stable_turns, later_turns, store.active_sessions for label, interleave, fixed in ( ("single conversation, legacy ", False, False), ("interleaved (subagents), legacy ", True, False), ("interleaved, lineage resolution ", True, True), ): stable, later, sessions = run(interleave, fixed) print(f"2) {label}: stable prefix on {stable}/{later} later turns, trackers={sessions}") ``` </details> ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation (CHANGELOG only — no docs describe the tracker store) - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Additional Notes - Addresses the session-identity mechanisms of #2085; intentionally does not `Closes` it — the reporter should confirm the cache-read ratio recovers on live traffic first. - Composes with draft #1912 (first-user-turn fallback id). - Known bounded tradeoffs (all strictly milder than the per-turn thrash this fixes): a fork-style branch that resends a parent's full history adopts the parent's lineage, costing the parent one cold restart at its next turn; a request that aborts before the response and is retried with different bytes starts a fresh lineage; history truncation/tail-edit starts a fresh lineage even though the shorter provider prefix may still be warm. - Hot-path cost, measured on a 199-message/2.1MB agentic history: canonical projection 0.21ms + structural snapshot 0.92ms + match loop 0.06ms with 32 candidate lineages (2.27ms absolute worst case) ≈ **1.3ms per request** — same order as the handler's existing request deepcopy (0.80ms) and below one `json.dumps` of the body (2.9ms). Chain memory is structure-only (~180-330KB per lineage; message strings are shared with state the tracker already retains). - Known semantic shift to flag: hashing only the leading system run means conversations distinguished ONLY by mid-list system messages (e.g. clients injecting a per-conversation system context late in the list) now share a fallback id. The tracker is protected by lineage resolution; the residual sharing concentrates in the CCR sticky-tool registry and the monotone beta union — the same pre-existing class as same-system-prompt conversations today. Happy to file the CCR-stickiness scoping as a follow-up. - Out of scope, observed while tracing: `SessionCcrTracker.has_done_ccr` mildly cross-contaminates conversations sharing an id (monotone, no thrash) — can file separately if useful. --------- Co-authored-by: JerrettDavis <mxjerrett@gmail.com> |
||
|
|
8522fcbc40
|
fix(code): quarantine Perl parser from code-aware compression (#2204)
## Description `headroom proxy` can wedge when code-aware compression enters the tree-sitter Perl external scanner and the native scan keeps the GIL indefinitely. Current main still has two routes into that scanner: explicit `perl` or `pl` hints flow through `CodeAwareCompressor`, and `detect_language()` can nominate Perl on non-Perl code because its prefilter matches generic sigils such as decorators, JSDoc tags, and shell variables before phase 2 parses every surviving candidate grammar. This change quarantines Perl at the code-aware compression funnel without widening scope. Perl remains recognized at the input boundary, but live proxy compression no longer requests a Perl parser. Non-Perl code keeps its existing code-aware behavior. Real Perl falls back through the existing safe Kompress or passthrough contract instead of entering tree-sitter. The diff stays inside `headroom/transforms/code_compressor.py` plus focused parser-safety regressions. Refs #2185 ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - Added a Perl quarantine list in `headroom/transforms/code_compressor.py` and used it to stop Perl candidate parsing during language detection. - Added a hard `_get_parser()` guard so no live code-aware path can construct a Perl parser. - Routed resolved explicit Perl hints through the existing safe fallback or passthrough contract before AST compression. - Added focused parser-safety regressions for non-Perl candidate bleed, explicit `perl` and `pl`, inferred real Perl, mixed fenced Perl content, fallback-disabled passthrough, and non-Perl negative space. ## Testing - [x] Unit tests pass (`uv run pytest tests/test_perl_scanner_safety.py -q`) - [x] Linting passes (`uv run ruff check headroom/transforms/code_compressor.py tests/test_perl_scanner_safety.py`) - [ ] Type checking passes (`uv run mypy headroom`) - [x] New tests added for new functionality when applicable - [ ] Manual testing performed ### Test Output ```text uv run pytest tests/test_perl_scanner_safety.py -q 8 passed in 1.93s uv run ruff check headroom/transforms/code_compressor.py tests/test_perl_scanner_safety.py All checks passed! uv run ruff format headroom/transforms/code_compressor.py tests/test_perl_scanner_safety.py --check 2 files already formatted ``` ## Real Behavior Proof - Environment: Windows, Python from the synced `uv` environment, `dev` and `code` extras installed, no provider call - Exact command / steps: Run `uv run pytest tests/test_perl_scanner_safety.py -q`. - Observed result: `8 passed in 1.93s`; the suite proves explicit `perl` and `pl`, inferred real Perl, mixed fenced Perl, and non-Perl negative-space routes all avoid Perl parser entry. - Not tested: the reporter's macOS payload and long-running concurrent workload ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [ ] I have commented my code, particularly in hard-to-understand areas - [ ] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Additional Notes - Use `Refs #2185`, not `Closes #2185`. The wedge surface is this PR's scope, but #2185 also carries a separate orphaned `headroom mcp serve` report that this slice does not address. - This is a reachability fix. It does not repair the upstream Perl scanner and it does not harden other grammars against the same class of native wedge. - `CHANGELOG.md` remains unchanged because Headroom generates release notes from conventional commits. - Merged PR https://github.com/headroomlabs-ai/headroom/pull/2114 addresses a different cooperative compression stall and stays separate from this native parser-entry slice. --------- Co-authored-by: JerrettDavis <mxjerrett@gmail.com> |
||
|
|
fd9ddaa238
|
fix(proxy): cold-start fast pass — defer only Kompress, not the whole pipeline (#2073)
## Description Since #1850, the freeze path forwards a session's provider-cached prefix byte-identical — so a session is permanently locked to whatever form its cold start put in the provider cache. That fix is correct (it stopped token-mode cache busting measured at +41% cost), but it interacts badly with off-path background compression (#1171): when `HEADROOM_BACKGROUND_COMPRESSION=1` defers a cold-start-large request (frozen=0, ≥50k tokens), the ENTIRE pipeline is deferred and the raw transcript is forwarded, cached, and frozen. The background job's results can never be applied afterward (doing so would rewrite the frozen prefix), so the session forfeits its compression savings for its lifetime. Field data (same day, same session, A/B across a version boundary): ~15k tokens/turn saved when the cold start compressed synchronously vs 0/turn forever when it deferred. Notably, the recurring savings came from `read_lifecycle` stale-read drops completing in ~300ms — deferral throws away sub-second lossless wins to avoid a 30s Kompress pass. Only the Kompress ML stage can blow the request budget (the #1171 cascade). This PR splits the two: - The deferral branch now runs the pipeline synchronously with a new `skip_kompress=True` per-call kwarg — everything except the ML stage — under a bounded budget, and forwards the pruned form. The provider caches (and #1850 freezes) the *compressed* transcript, so the cheap savings persist for the session's lifetime. - The full pipeline (Kompress included) still goes to the background job, unchanged, keyed against the original messages so its content-hash results remain reusable at future cache-miss boundaries. - Fail-open: on fast-pass timeout or error, the request forwards uncompressed exactly as before this change. ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - `headroom/transforms/content_router.py`: new per-call `skip_kompress` runtime kwarg (follows the existing `_runtime_force_kompress` pattern). Gates only the Kompress deep-path call site; units routed there take the identical fallback used when the model isn't ready. Wins over `force_kompress`. - `headroom/proxy/helpers.py`: `COLD_START_FAST_PASS_TIMEOUT_SECONDS` (env `HEADROOM_COLD_START_FAST_PASS_TIMEOUT_SECONDS`, default 10s), documented next to `COMPRESSION_TIMEOUT_SECONDS`. - `headroom/proxy/handlers/anthropic.py`: the background-deferral branch runs the fast pass synchronously, stores its result in the session `CompressionCache`, forwards the pruned messages, and tags `deferred:kompress_background` (or `deferred:dropped` when the enqueue was dropped). On failure it constructs the same `_DeferredCompressionResult` as before. The Anthropic handler is the only deferral site (OpenAI/Gemini handlers don't defer). - `tests/test_transforms/test_content_router.py`: `skip_kompress` never invokes the ML stage and wins over `force_kompress` (mirrors the existing `force_kompress` test). - `tests/test_cold_start_fast_pass.py`: handler-level tests — exactly one synchronous `skip_kompress=True` pass, the background job runs the full pipeline, the forwarded body carries the fast-pass form, fast-pass results land in the compression cache; and the fail-open path (executor timeout → original messages forwarded, background job still queued). - `CHANGELOG.md`: Bug Fixes entry. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check .`) - [x] Type checking passes (`mypy headroom`) - [x] New tests added for new functionality - [ ] Manual testing performed ### Test Output ```text $ uv run --frozen --extra dev pytest tests/test_cold_start_fast_pass.py -v tests/test_cold_start_fast_pass.py::test_cold_start_runs_fast_pass_and_defers_only_kompress PASSED tests/test_cold_start_fast_pass.py::test_fast_pass_failure_falls_back_to_full_deferral PASSED ============================== 2 passed in 0.28s =============================== $ uv run --frozen --extra dev pytest tests/test_proxy/test_background_compression.py \ tests/test_proxy/test_phase3_byte_identity.py tests/test_anthropic_stage_timings.py \ tests/test_transforms/test_content_router.py tests/test_anthropic_pre_upstream_backpressure.py ======================== 90 passed, 1 warning in 10.47s ======================== $ uv run --frozen --extra dev mypy headroom/proxy/handlers/anthropic.py headroom/proxy/helpers.py headroom/transforms/content_router.py Success: no issues found in 3 source files $ ruff check <changed files> && ruff format --check <changed files> All checks passed! / 5 files already formatted ``` ## Real Behavior Proof - Environment: macOS 15 (darwin 24.6.0), Python 3.10 via `uv run --frozen --extra dev`; field logs from a production desktop deployment (Python 3.12, `HEADROOM_MODE=token`, `HEADROOM_BACKGROUND_COMPRESSION=1`, subscription auth policy). - Exact command / steps: compared per-request PERF log lines for the same Claude Code session served by 0.30.0-lineage (sync cold start) vs 0.31.0-lineage (deferred cold start) on the same day. - Observed result: deferred-cold-start sessions log `tok_saved=0` on every subsequent turn with `Pipeline: freezing first 281/284 messages`; sync-cold-start sessions log `tok_saved=15526-18791` per turn with `read_lifecycle:stale` transforms at `opt_ms≈300`. - Not tested: this patch has not run against a live proxy yet (behavior verified at the handler-test level); `ruff`/`mypy` scoped to changed files. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md if applicable ## Screenshots (if applicable) N/A — proxy pipeline change, no UI. ## Additional Notes - Companion to #2057 (nested tool_result image token counting) and #2058 (new-content-relative savings rate) — all three came out of the same investigation into near-zero reported savings on long 1M-context Claude Code sessions. - Deliberate scope cuts: the OpenAI/Gemini handlers don't have a deferral branch, so nothing to change there; the background job is left keyed to original messages (not the fast-pass output) so its cached results match client-resent bytes at future cache-miss boundaries. - Timeout leak caveat is documented in code: a fast-pass timeout briefly leaks an executor worker, but without the ML stage the pass is bounded by routing + statistical crushers (observed 5-8s worst case on multi-M-token counted transcripts). 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Tejas Chopra <chopratejas@gmail.com> Co-authored-by: JerrettDavis <mxjerrett@gmail.com> |