mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
20 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ddc6f6ceb0 |
fix: C1 — byte-level SSE parser + state machines
Foundation of Phase C. Delivers:
* Byte-level SSE framing (bytes::Bytes / BytesMut) with UTF-8
decoded only at \n\n event boundaries — no per-chunk decode,
no errors=ignore data loss across TCP reads.
* Three provider state machines:
- Anthropic: blocks keyed by index, all delta types
(text/thinking/input_json/citations/signature) preserved
byte-equal.
- OpenAI Chat: ToolCallState concatenation, refusal field,
include_usage final chunk handling.
- OpenAI Responses: items keyed by id (not position) for
out-of-order completion; full event coverage.
* State machine runs in parallel with byte-passthrough via a
tokio::spawn task fed by a bounded mpsc — clients see raw
bytes immediately; telemetry populates without blocking.
Retires P1-8, P1-9, P1-14, P1-15, P1-17, P4-48 in the Rust path
(Python A8 hotfix preserved as fallback until Phase H).
Per-PR-C1 plan: REALIGNMENT/05-phase-C-rust-proxy.md.
|
||
|
|
00902b8fea |
fix: B7 — CCR hardening: persistent backends + always-on tool
P2-25, P2-26: CCR (Compress-Cache-Retrieve) used an in-memory store
that fragmented across uvicorn workers and was wiped on restart, and
the `headroom_retrieve` tool was registered/unregistered per-request
based on whether the latest body happened to contain compression
markers — every flip busted the prompt cache. Both are sticky
side-channels: once a session has done CCR, the tool list bytes and
the retrieval store must stay stable. This PR fixes both.
Rust:
* Split `ccr.rs` into `ccr/` with `backends/` submodule
(`in_memory.rs`, `sqlite.rs`, `redis.rs` cfg-gated).
* `SqliteCcrStore` (production default): WAL mode, prepared upsert,
lazy TTL purge on read, persistent across worker restarts and
shareable across workers on the same host via SQLite file locking.
* `RedisCcrStore` (cfg-gated behind `feature = "redis"`): SETEX with
startup PING smoke-test, no key-prefix collision risk, no sticky
session required at the LB.
* `CcrBackendConfig::{InMemory, Sqlite, Redis}` + `from_config(...)`
factory — every init failure surfaces (no silent fallback per
`feedback_no_silent_fallbacks.md`).
* `ccr::compute_key` (BLAKE3 → first 24 hex chars) and
`ccr::marker_for("HASH") -> "<<ccr:HASH>>"` centralize the hash +
marker format; one definition for the live-zone dispatcher and the
Python regex (`headroom/ccr/tool_injection.py:211`).
* `compress_anthropic_live_zone_with_ccr` accepts
`Option<&dyn CcrStore>`. When wired, every accepted compression
puts the original bytes into the backend and appends `<<ccr:HASH>>`
to the compressed string. The token-validation gate runs on the
marker-augmented string so the `compressed_tokens >=
original_tokens` rejection stays honest.
Python:
* `SessionCcrTracker` + `apply_session_sticky_ccr_tool` mirror the
PR-A7 `SessionToolTracker` / `apply_session_sticky_memory_tools`
pattern: once a session has done CCR, every subsequent request
injects the recorded golden tool-definition bytes. Tool list bytes
are byte-stable across turns (snapshot test pins them).
* `headroom/ccr/tool_injection.py::inject_tool_definition` accepts a
new `session_has_done_ccr` kwarg per the PR-B7 spec change at line
302-328. The legacy per-request path stays intact for callers that
don't yet thread a session id (e.g. Google handler).
* Anthropic + OpenAI handlers route their CCR tool-list updates
through `apply_session_sticky_ccr_tool`, keyed off the existing
`session_tracker_store.compute_session_id(...)` plumbing.
Backend selection model: `CcrBackendConfig::Sqlite { path }` is the
production default — single host, persistent, multi-worker safe with
sticky session. `CcrBackendConfig::Redis { url }` is the multi-host
scale-out option — no stickiness needed. `InMemory` is for tests
and single-worker dev only. RUST_DEV.md "Multi-worker deployment —
CCR fragmentation" rewritten around this matrix.
Tests:
* `crates/headroom-core/tests/ccr_backends.rs` — 7 tests covering
SQLite round-trip, TTL purge, proxy-restart survival, cross-backend
byte-equal keys, `from_config` paths, and the no-redis-feature
loud-failure check (+ 2 redis tests gated behind the feature).
* `crates/headroom-core/tests/live_zone_ccr.rs` — confirms
`<<ccr:HASH>>` marker injection, store population, and
no-marker-when-no-store invariants end-to-end.
* `tests/test_ccr_tool_always_on.py` — 12 tests pinning the
always-on behaviour, session/provider isolation, LRU bound, no-
session-id fallback, and (per-acceptance-criterion) the byte-stable
tool-definition snapshot.
Per-PR-B7 plan: REALIGNMENT/04-phase-B-live-zone.md.
|
||
|
|
a974bb153a |
fix(rust): PR-A1 — make /v1/messages compression a passthrough
Stop calling IntelligentContextManager from the Rust proxy on
/v1/messages. The proxy is now a byte-faithful passthrough on this
endpoint. Eliminates the C1+C2+C3+C4 cache-killer cluster (P0-3,
P0-4, P0-5, P1-13) by not running ICM with `frozen_message_count: 0`
hardcoded — Phase B PR-B2 brings live-zone-only compression back.
Per REALIGNMENT/03-phase-A-lockdown.md.
Changes:
- Add `--compression-mode {off,live_zone}` flag and
`HEADROOM_PROXY_COMPRESSION_MODE` env var. Default `off`. Both
modes passthrough in PR-A1; `live_zone` warns loudly because
Phase B isn't implemented yet (no silent fallback).
- Replace `compress_anthropic_request` body with a passthrough
stub that emits a structured `tracing::info!` decision log line
(request_id, path, method, compression_mode, decision,
reason="phase_a_lockdown", body_bytes) and returns
`Outcome::NoCompression`. Function signature preserved so
Phase B PR-B2 is a pure body swap.
- Delete `compression/icm.rs` (per the realignment plan: ICM
modules in headroom-core are deleted in PR-B1).
- Drop the `Arc<IntelligentContextManager>` field from `AppState`
— no longer used.
- Add request-entry `tracing::debug!` with auth_mode_placeholder
("unknown" until Phase F PR-F1 wires the auth-mode classifier).
- Add `debug_assert!` on the NoCompression branch that the
buffered bytes length is stable, locking in Phase A's
cache-safety invariant at the call site.
- Tighten existing tests from `len()` equality to SHA-256 byte
equality. Rename `compression_on_oversized_body_trims_messages`
→ `compression_on_long_body_passes_through_in_phase_a` and
flip the assertion to byte-equal.
- Add new tests: passthrough_mode_off_byte_equal_sha256,
passthrough_mode_live_zone_currently_passthrough_byte_equal_sha256,
passthrough_preserves_numeric_precision (literal-byte body so
serde_json's f64 quantization can't mask a regression),
passthrough_preserves_cache_control_markers,
passthrough_preserves_thinking_signature,
passthrough_preserves_redacted_thinking_data,
passthrough_recorded_fixture_byte_equal_sha256,
tracing_capture::compression_decision_logged.
- Add fixture
`crates/headroom-proxy/tests/fixtures/anthropic_messages_request_real.json`
with system block list + cache_control markers, tools with
nested JSON Schema, messages containing text + thinking +
signature + tool_use + tool_result + image, non-ASCII content,
large numbers. Used as the canonical SHA-256 round-trip gate.
Constraints honored: configurable (compression_mode is the only
new knob), no hardcoded thresholds, no regex usage, no silent
fallbacks (live_zone-not-implemented warns), structured tracing
on every cache-affecting decision, comprehensive tests.
Acceptance criteria from PR-A1 spec:
- `cargo build --workspace` clean
- `cargo test --workspace` green (886 tests pass)
- `cargo clippy --workspace -- -D warnings` clean
- `cargo fmt --all --check` clean
- `make ci-precheck` green
- New SHA-256 byte-equality tests pass against the recorded fixture
- `tracing::info!` decision-log line is observable
- `--compression-mode` CLI + env var work
- No regex import added
|
||
|
|
378d8a0f05 |
fix(rust): audit cleanup — DiffCompressor CCR leak, CCR TOCTOU race, clippy debt, dep dedup
Closes findings from the post-Phase-3g audit. Five surgical fixes plus telemetry-discoverability docs. PyO3 0.22 → 0.24 security upgrade is its own PR (issue #335). 1. DiffCompressor cache_key persistence (production bug) --------------------------------------------------------- Pre-fix: `RustDiffCompressor.compress()` minted a `cache_key`, embedded `[... hash=abc123]` in the wire marker, and returned without storing the original anywhere. Python ContentRouter then returned the compressed text with a dangling marker — every retrieval tool call from the LLM 404'd. Sibling compressors (LogCompressor, SearchCompressor) already had the right pattern: Rust mints the key, Python's `_persist_to_python_ccr` writes the original to the production `CompressionStore`. DiffCompressor was the asymmetric one. Fix: - Rust: add `DiffCompressor::compress_with_store(content, context, Option<&dyn CcrStore>)` mirroring siblings. Calls `store.put` when a key is minted; legacy `compress()` and `compress_with_stats()` delegate with `None` for parity. - Python: add `_persist_to_python_ccr` helper to `headroom/transforms/diff_compressor.py.compress()` mirroring `log_compressor.py` and `search_compressor.py`. - Pipeline `DiffOffload`: switch to `compress_with_store(Some(store))` and drop the post-hoc double-store hack that papered over this bug at the orchestrator boundary. 2. CCR store TOCTOU race in `get()` ----------------------------------- `InMemoryCcrStore::get()` checked TTL under a read lock, dropped the lock, then called `remove()`. Between drop and remove a concurrent `put()` of the same hash with fresh data could land — and our `remove` would then wipe that fresh entry. Under multi-worker proxy load this manifested as "I just stored it; why is it gone?" Fix: use `DashMap::remove_if`. Predicate runs under the shard write lock so check-and-remove is atomic. New regression test exercises a tight contention loop between writer and reader on the same key. 3. Pre-existing clippy debt in smart_crusher -------------------------------------------- - 3× `field_reassign_with_default` in `crusher.rs` test setup — switch to struct-update syntax `Config { field: x, ..Default }`. - `hash_array_for_ccr` was `#[cfg(test)]` but unused; deleted with a comment so a future test can reintroduce it as a one-liner. `cargo clippy --workspace --all-targets -- -D warnings` is now clean across the whole workspace; previous CI patches that allowed these warnings can be removed in a follow-up. 4. Tokenizers dependency dedup ------------------------------ `tokenizers 0.21` (direct dep) + `tokenizers 0.22` (transitive via fastembed) compiled twice into the binary. Bumped direct dep to `0.22` to align; API is compatible (verified by full tokenizer test suite). Saves compile time + binary bloat. 5. Telemetry-discoverability doc (no new code) ---------------------------------------------- The audit recommended a per-transform invocation counter to inform the next Python → Rust port. Discovered the infrastructure already exists at `/stats`: - `compressions_by_strategy` — invocation count per strategy - `pipeline_timing` — count + avg/max ms per transform name - `tokens_saved_by_strategy` — savings attribution Added a section to `RUST_DEV.md` showing the `curl + jq` recipes to read this data, with example output highlighting how to spot zero-invocation deferral candidates (e.g. `code_compressor`). Verification: workspace tests 734 + 14 + 5 + 4 + 6 + 5 + 2 + 2 + 3 + 4 + 2 + 2 + 1 = all green; cargo fmt clean; cargo clippy --all-targets clean; Python tests 185 pass; commitlint clean. |
||
|
|
01a423a316 |
fix(rust): reformat/offload pipeline + log templates + diff noise (Phase 3g rework)
Replaces PR1's lossless/lossy split with ReformatTransform (pack denser, no info lost) and OffloadTransform (drop bytes, CCR-stash original via required cache_key). With CCR every transform is information-preserving end-to-end, so the lossless/lossy distinction misnamed the architecture. OffloadTransform carries a cheap, structural estimate_bloat() method scoped to its domain — generic byte-redundancy heuristics miss domain semantics. The orchestrator runs reformat phase + per-offload bloat estimation in parallel via rayon::join + par_iter, then runs offload iff bloat clears threshold OR reformat underwhelmed. Transforms shipped: REFORMATS (lossless): - JsonMinifier: serde_json round-trip whitespace stripping. - LogTemplate: Drain-inspired order-preserving template miner. Collapses consecutive runs of same-template lines into [Template Tn: ...] (Nx) + variant table. Win comes from emitting the constant-token prefix once instead of N times. Lossless: every original line reconstructible from template + variants. OFFLOADS (drop bytes, stash original via CCR): - LogOffload: wraps existing LogCompressor; bloat = repetition x uniqueness_weight + dilution x priority_dilution_weight. - DiffOffload: wraps existing DiffCompressor; bloat = context-to- change ratio. Bug-fix-on-port — persists original under the cache_key the parity-bound DiffCompressor mints (closes a leak). - DiffNoise: drops lockfile hunks (Cargo.lock, package-lock.json, yarn.lock, etc., suffix list configurable in TOML) and whitespace-only hunks. Stashes original via CCR for retrieval. Search offload exists but is not in default re-exports — modern agents (Claude Code, Codex) use scoped rg/grep, the marginal value didn't justify default registration. Reach via the explicit module path if opting in. JSON Offload is intentionally absent from this PR — already lives at SmartCrusher; Phase 3g PR3 wraps it in the OffloadTransform contract. Thresholds and weights live in config/pipeline.toml, embedded via include_str!; PipelineConfig::from_toml_str loads runtime overrides. 98 new pipeline tests; full headroom-core suite (714) and workspace tests green; cargo fmt clean. No regex, per project convention. |
||
|
|
12c2665531 |
feat(rust): signals trait module + KeywordDetector (Phase 3e.1)
Establish `crates/headroom-core/src/signals/` as a top-level module holding cross-cutting detection traits. Phase 3e.1 ports `error_detection.py` to a `LineImportanceDetector` trait + a `Tiered<T>` combinator + a single concrete `KeywordDetector` impl backed by aho-corasick. Three traits at three granularities are sketched (line / blob / item); only line-importance is implemented today. Two bug fixes from the Python source bake into both the Rust impl and the Python regex shim: 1. `ERROR_KEYWORDS` listed `timeout|abort|denied|rejected` but `ERROR_PATTERN` regex omitted them. Lines like `"Connection timeout"` were silently neutral despite the keyword being canonical. Both surfaces now flag them. 2. `SECURITY_KEYWORDS` carried `token`, which false-positived on every reference to LLM tokens (`input_tokens`, `tokens_saved`, ...) in our own product. Dropped from the security set. The Python `error_detection.py` shim now reflects keyword data out of Rust via `keyword_registry_snapshot()` and recompiles the legacy `re.Pattern` objects on the fly. Existing callers (text_compressor, search_compressor, intelligent_context) continue to import the same names with no source changes; caller migration to the trait API happens in their own port PRs. The trait architecture is the seam where a future ML detector slots in without touching `KeywordDetector` or any caller. The canonical extension is documented in `signals/README.md` as a classifier head on the existing `bge-small-en-v1.5` embedder loaded by `relevance::EmbeddingScorer` -- 384-dim -> 4-class softmax, ~1.5 KB head, ~1 ms inference, no extra model file. Two alternatives (distilled tinyBERT in ONNX, logistic regression on lexical features) are kept open in case BGE-head underfits. Per the no-silent-fallbacks rule: only `KeywordDetector` lands as a concrete impl. No NoOp, no MockDetector, no stub-ML -- those will arrive with their real implementations. Phase 3g (Compression Pipeline Formalization, issue #315) is queued as the cross-cutting follow-up that will make lossless-then-lossy- then-CCR ordering an explicit, observable architecture rather than implicit per-compressor logic. Trait shapes there will reuse the signals primitive landed in this PR. |
||
|
|
19fc49ac39 |
chore(rust): port unidiff Tier 2 diff detector (Stage 3d PR4)
Adds the second tier of the Stage-3d ContentRouter detection arch.
Magika (PR3) is a probabilistic ML classifier — short, prose-prefixed,
or "looks like code because the lines are code" diffs can slip past
it into PlainText. PR4 catches those by running the [`unidiff`]
parser as a deterministic oracle: anything that parses to ≥1
PatchedFile with ≥1 hunk is a diff.
What lands:
- `crates/headroom-core/src/transforms/unidiff_detector.rs`:
- `is_diff(content) -> bool`: predicate.
- `detect_diff(content) -> Option<ContentType>`: typed wrapper for
the router (PR5) to chain after Magika.
- Empty input shortcuts to false without invoking the parser.
- "Found zero hunks" is treated as **not** a diff — `unidiff::PatchSet
::parse` returns Ok(()) on plain text (just finds zero files);
we explicitly require non-empty patch + non-empty hunk to avoid
silently routing prose through the diff compressor.
- 14 unit tests: standard git diff, naked hunk without git header,
multi-file, added/removed-only files, JSON/HTML/YAML/source/prose
negatives, "almost looks like a diff" prose with @@/--- in passing,
truncated-diff canary.
Known gaps (deliberately punted to PR5+):
- Combined-merge headers (`@@@ ... @@@`) — `unidiff`'s hunk regex is
for plain `@@`. Rare in proxy traffic; PR5 router can fall back
to the regex content_detector if needed.
- Pathological CRLF-stripped inputs — `input.lines()` strips `\r`
only when paired with `\n`. Acceptable.
What does NOT land here (per PR scope):
- No PyO3 surface — module-only.
- No router rewiring — the existing regex `content_detector` still
drives `ContentRouter`. PR5 chains magika → unidiff → PlainText.
The `unidiff` crate brings `regex` (already in tree) and `encoding_rs`
(default features) — small dep impact.
`make ci-precheck` green.
|
||
|
|
d34658b22e |
chore(rust): port Magika detection (Stage 3d PR3 — Tier 1)
Adds Google's `magika` ONNX-backed content classifier as the first
tier of the new Stage-3d ContentRouter detection arch (`magika` →
`unidiff-rs` → `PlainText` fall-through; no regex tier on the Rust
side).
What lands:
- New module `crates/headroom-core/src/transforms/magika_detector.rs`:
- `magika_detect(content: &str) -> Result<ContentType, _>`
- `OnceLock<Mutex<Result<Session, _>>>` singleton: model loads
once per process; init failure is recorded once and cheaply
replayed (no retry — rust-side `feedback_no_silent_fallbacks`).
- `map_magika_label(&str) -> ContentType`: explicit match arms
against magika's 200+ labels, mapped onto Headroom's existing
`ContentType` enum so the dispatch (PR5) stays enum-stable.
Unmapped labels passthrough to `PlainText` rather than misroute.
- 16 unit tests: empty fast-path, JSON / Python / Rust / JS /
diff / markdown / plain prose / HTML / YAML / shell / SQL,
singleton-reuse smoke, default-passthrough for unmapped labels,
pure-table-lookup sanity.
What does NOT land here (per PR scope):
- No PyO3 surface yet — PR3 is detector-only.
- No router rewiring — the existing regex `content_detector` still
drives `ContentRouter` until PR5 flips the dispatch.
- No `unidiff-rs` Tier-2 — that's PR4.
The `magika` crate brings `ndarray` + `ort` (already in our dep
tree via `fastembed`); adding it shares the ONNX Runtime singleton
rather than pulling a second ML stack.
`make ci-precheck` green.
|
||
|
|
29aadb1054 |
perf(rust): tier-1 multi-worker wins — GIL release, sharded CCR store, single-serialize CCR write
Three orthogonal hot-path fixes targeting concurrent-request throughput.
Each is independently bench-measured below; the proxy hot path benefits
from all three at once.
== 1. PyO3 GIL release on heavy compute ==
PyO3 methods (crush, smart_crush_content, crush_array_json,
compact_document_json, compress, compress_with_stats) used to hold the
GIL across the entire Rust call. Result: a 100ms compress() blocked
EVERY other Python thread for 100ms — multi-worker uvicorn deployments
serialized through SmartCrusher.
Wrap each compute call in `py.allow_threads(|| ...)`. Inputs (`&str`
from Python) are copied to owned `String` first because PyO3 ties them
to the GIL hold. PyDict construction stays on the GIL side.
Measured: 4 Python threads each running 20 crushes:
before (GIL held): ~3.3s wall (serialized — equivalent to 4×0.83s)
after (allow_threads): 826ms wall (4.01x speedup, perfect parallel)
== 2. CcrStore: Mutex<HashMap> -> DashMap-backed sharded ==
Single Mutex was the dominant bottleneck under multi-worker load — every
put/get serialized through one lock. Replace with DashMap (sharded
concurrent map, lock-free reads within a shard) plus a separate
small Mutex<VecDeque> for FIFO insertion-order eviction. Reads of
distinct keys never contend; writes only contend during the brief
order-queue push or capacity-sweep.
A/B bench (200 mixed put/get ops × N threads, in benches/ccr_store.rs):
Threads | DashMap Legacy Mutex Speedup
-------------------------------------------
1 | 63 µs 71 µs 1.13x
2 | 98 µs 194 µs 2.0x
4 | 178 µs 707 µs 4.0x
8 | 342 µs 1267 µs 3.7x
Legacy degrades ~linearly with thread count; DashMap stays near-flat
per-thread. Real multi-worker scaling.
== 3. Single-serialize the lossy CCR payload ==
The lossy `crush_array` path used to serialize the full array TWICE:
once in `hash_array_for_ccr` (allocates `Value::Array(items.to_vec())`,
deep-clones every Value subtree, then serializes), and a second time
in the store-write site. For a 50-item dict array that's ~MB of
allocator pressure per crushed array.
Introduce `canonical_array_json` (serializes `&[Value]` directly — same
bytes as `Value::Array(items.to_vec())` but no wrapper allocation +
no tree clone), call it ONCE per lossy path, then both hash and store
from those same bytes. Hash-format stable — all 17 parity fixtures
match byte-for-byte.
== Tests ==
- 8 ccr.rs unit tests including a new concurrent-stress test (8 threads
× 200 puts/gets, every key readable afterwards)
- 14 ccr_roundtrip integration tests stay green
- parity-run smart_crusher: 17/17 fixtures match
- 479 lib + 14 integration + 185 Python tests all pass
- New benches/ccr_store.rs runs the A/B and is committed for regression
visibility
== Dependencies added ==
- dashmap v6 (mature, widely-used in tokio/linkerd ecosystem)
|
||
|
|
22c8fec4c1 |
chore(rust): SmartCrusher CCR storage layer + roundtrip verification
CcrStore trait + InMemoryCcrStore (1000 entries, 5-min TTL, FIFO eviction, idempotent re-store) live at the crate root. SmartCrusher's lossy crush_array path now actually stashes the full original [items] canonical-JSON into the configured store keyed by the same ccr_hash it embeds in the prompt marker -- closing the no-data-loss contract that was previously hash-only. PyO3 surface: - crusher.crush_array_json(items_json) -> dict with ccr_hash + kept items - crusher.ccr_get(hash) -> Optional[str] for retrieval - crusher.ccr_len() -> int for telemetry Python shim passes both through. Default constructors enable the store (matches Python's CCR-enabled default); without_compaction() also gets it because CCR is a contract, not an opt-in extra. Tests proving compress -> store -> retrieve -> reconstruct: - 7 unit tests in ccr.rs (put/get/eviction/expiry) - 9 Rust integration tests (crates/headroom-core/tests/ccr_roundtrip.rs) - 10 Python tests including 4 explicit before/after element-equality assertions through both the native PyO3 surface and the Python shim Plugin manifest versions auto-bumped by the sync-plugin-versions pre-commit hook (unrelated to CCR but co-resident in the working tree). |
||
|
|
1945e5f55b |
feat(rust): real fastembed-rs EmbeddingScorer (BAAI/bge-small-en-v1.5)
Replace the embedding scorer stub with a real fastembed-rs implementation. Same library + same model as the Python side will use after the next commit, giving byte-equal embeddings on identical inputs. Cargo.toml: fastembed = "5". Default features pull in `ort` (ONNX Runtime) with auto-download of the runtime binary at build time (~21s additional first-build); model weights (BAAI/bge-small-en-v1.5, ~30 MB int8-quantized ONNX) auto-download from HuggingFace Hub on first use. embedding.rs: - EmbeddingScorer wraps Option<Mutex<TextEmbedding>>. Mutex required because TextEmbedding::embed needs &mut self (single-threaded ONNX session); concurrent callers serialize on the lock, fine for the SmartCrusher hot path where inference dominates lock contention. - EmbeddingScorer::try_new() — explicit construction with HF Hub download. Returns Result; surface errors to callers. - EmbeddingScorer::try_new_with_model(EmbeddingModel) — bring your own model from fastembed's catalog. - EmbeddingScorer::default() — STUB only (model=None, is_available()=false). Mirrors Python's "sentence-transformers not installed" branch byte-for-byte. To get a real scorer, call try_new() and pass via HybridScorer::with_scorers(). Why default() is a stub: with auto-load Default, model availability would depend on whether HF Hub cache has the file — non-deterministic in tests. Explicit try_new() keeps Default cheap and predictable. cosine_similarity: - f32 vec inputs (fastembed returns Vec<Vec<f32>>). - Clamped to [0, 1] (mirrors Python _cosine_similarity — only positive similarity matters for relevance). - Defensive: zero vectors / mismatched dims → 0.0. score / score_batch: - Empty input / unavailable model → empty score with explanatory reason. - Batch encodes items + context in one model call (Python parity: amortizes model dispatch). - Inference failures degrade gracefully with empty scores rather than panicking. Tests: - 5 cosine-similarity unit tests (offline). - 3 unavailable-scorer tests (model=None path). - 3 model-backed integration tests gated on RUN_FASTEMBED_TESTS=1 (semantic-match-outranks-unrelated, batch-shape, model-loads). - All 388 headroom-core tests pass without RUN_FASTEMBED_TESTS; with it set, the gated 3 also pass. Net: 388 unit tests, clippy clean. HybridScorer's BM25-fallback path remains correct (default embedding scorer reports unavailable). Stage 3c.1 next: switch Python's relevance/embedding.py to fastembed PyPI package + record parity fixtures with real embeddings on both sides. |
||
|
|
9d515fb78e |
feat(rust): smart_crusher universal crushers — string, number, object
Three crushers from headroom/transforms/smart_crusher.py ported. Each takes a SmartCrusherConfig + bias and returns (crushed_items, strategy_string). All schema-preserving — output is items/values from the original; no generated text. What's in: 1. compute_k_split (smart_crusher.py:2693) Wraps adaptive_sizer::compute_optimal_k. Splits k_total into first/last/importance via config.first_fraction / last_fraction. Uses f64::round_ties_even() (Rust 1.77+) to match Python's banker's-rounding round() — important for off-by-one parity on .5-edged k computations. 2. crush_string_array (smart_crusher.py:2727) Adaptive K via Kneedle. Mandatory-keep: error-keyword strings + length-anomaly strings (>variance_threshold σ from mean length). Boundary-keep: first K_first + last K_last. Stride-based diverse fill with content-dedup. Output preserves original array order (BTreeSet iteration). Strategy includes dedup= and errors= counts when nonzero. 3. crush_number_array (smart_crusher.py:2810) — CARRIES BUG #1 Statistics-driven (mean/median/stdev/p25/p75). Outliers flagged at variance_threshold σ. Change-points via window-mean comparison (config.preserve_change_points + n>10 gates). Strategy string embeds full stats summary via format_g (Python's :.4g approximation). BUG #1 — percentile off-by-one — ported AS-IS: sorted_finite[len/4] / sorted_finite[3*len/4]. Cosmetic (strategy-string only). Test bug1_percentile_off_by_one_documented pins the buggy index choice; commit 7 fixes both languages and regenerates fixtures. 4. crush_object (smart_crusher.py:3015) Token-budget gate (config.min_tokens_to_crush=200). Three passthrough exits: n<=8, total tokens too low, k_total>=n. Always keeps: error-keyword values + small values (<=12 tokens via len/4 + len/4 + 2 heuristic). Boundary keys + stride fill with Python's recompute-each-iter cap (mirrored faithfully — slower but parity-true). Output preserves key insertion order via serde_json/preserve_order's IndexMap. Supporting helpers in stats_math.rs: - median(values) — Python statistics.median (mean-of-middles for even, total_cmp sort for NaN determinism). - format_g(x) — approximate Python f"{x:.4g}" (4 sig figs, scientific outside [-4, 4) exponent range, trailing-zero strip, explicit-sign 2-digit exponent). Pinned by 5 fixed-output tests. Field iteration order: key/object iteration uses BTreeMap-sorted (in analyzer) and IndexMap-insertion-order (in serde_json::Map for crush_object). The Python sorted-key fix scheduled for commit 7 also covers crush_object's iteration paths. Net: 266 unit tests passing in headroom-core, clippy clean (MSRV 1.80), parity harness intact (4/4 diff_compressor). Next commit: planning + execution layer (_create_plan, _execute_plan, plan-builder methods) with BUG #4 fix (k-split overshoot). |
||
|
|
a64716d5d1 |
fix(rust): smart_crusher scaffold review findings — hash truncation, int parse, python-repr matcher
Code review (`/code-review` on commit `
|
||
|
|
d219beecab |
feat(rust): scaffold smart_crusher module + foundational helpers
Stage 3c.1 — like-for-like Rust port of `headroom/transforms/smart_crusher.py`. This commit lays the foundation: module layout, configuration, foundational data types, and the simpler helpers (classification, hashing, anchors, basic statistics). Subsequent commits add the analyzer, crushers, plan execution, and the orchestrator. # What's in this commit `crates/headroom-core/src/transforms/smart_crusher/`: - `mod.rs` — module entry, public re-exports, port narrative. - `classifier.rs` — `classify_array` / `ArrayType` (dict/string/number/ bool/nested/mixed/empty). Direct port of `_classify_array`. - `config.rs` — `SmartCrusherConfig` with defaults pinned to Python byte-for-byte. - `hashing.rs` — `hash_field_name` (SHA-256 truncated to 16 hex chars), matches `hashlib.sha256(name.encode()).hexdigest()[:16]` exactly. - `statistics.rs` — `is_uuid_format`, `calculate_string_entropy`, `detect_sequential_pattern` (with **BUG #2 fix** — see below). - `anchors.rs` — `extract_query_anchors`, `item_matches_anchors`. Five regex patterns ported via `std::sync::LazyLock`. - `types.rs` — `CompressionStrategy`, `FieldStats`, `CrushabilityAnalysis`, `ArrayAnalysis`, `CompressionPlan`, `CrushResult`. Field-by-field mirror of the Python @dataclasses so the PyO3 bridge in 3c.1b can reconstruct them without manual translators. # Bug #2 fixed in this commit (Python fix lands later in same PR) `smart_crusher.py:444-448` — `_detect_sequential_pattern` calls `int(string_value)` and silently strips zero-padding, so padded string IDs like `["001", "002", ..., "100"]` get misclassified as a sequential numeric pattern. Fix: track whether each parsed numeric value originated as a string. If EVERY parsed value was a string, refuse to flag as sequential. Mixed numeric+string fields still detect correctly because the unambiguous numerics dominate. Test: `bug2_zero_padded_strings_no_longer_misclassified`. # What's NOT in this commit (subsequent commits) - `SmartAnalyzer` — `analyze_array`, `_analyze_field`, `_detect_change_points`, `_detect_pattern`, `_detect_temporal_field`, `analyze_crushability`, `_select_strategy`, `_estimate_reduction`. - The five array crushers (`_crush_array`, `_crush_string_array`, `_crush_number_array`, `_crush_mixed_array`, `_crush_object`). - Planning (`_compute_k_split`, `_create_plan`, `_plan_*` family). - Orchestration (`_prioritize_indices`, `_deduplicate_indices_by_content`, `_fill_remaining_slots`). - `SmartCrusher` orchestrator class itself. - Parity harness fixtures. - The remaining 3 Python bug fixes (#1, #3, #4) — landed alongside the code paths they affect. # Build / test - `cargo build -p headroom-core` — clean. - `cargo clippy -p headroom-core -- -D warnings` — clean. - 55 new unit tests across the 6 new files, all passing. Architectural improvements (lossless-first, unified saliency score, structured CCR markers) are deferred to Stage 3c.2 — see design doc at `~/Desktop/SmartCrusher-Architecture-Improvements.md`. |
||
|
|
5c3c9c49f2 |
feat(rust): diff_compressor port — byte-equal parity + sidecar stats
Stage 3a: first real transform port. Faithful Rust port of
`headroom.transforms.diff_compressor` with byte-equal parity against all
20 recorded fixtures.
# Algorithm (matching Python)
1. Hand-rolled unified-diff parser (state machine over `diff --git`,
`index`, `--- a/`, `+++ b/`, `@@`, mode/binary/rename markers, +/- /
space lines, "other" lines like `\ No newline at end of file`).
2. File cap (`max_files=20`): when fired, sort by total changes (most
first) and keep top N.
3. Per-file hunk cap (`max_hunks_per_file=10`): keep first + last + top
relevance-scored middle, then resort by hunk-header start line to
restore appearance order.
4. Relevance scoring: change-density base + user-query word overlap
+ priority patterns (ERROR / IMPORTANCE / SECURITY regexes —
matches `error_detection.PRIORITY_PATTERNS_DIFF`).
5. Per-hunk context trim: keep `max_context_lines=2` lines either side
of each `+`/`-` line.
6. CCR cache_key: `md5(original)[:24]` (matches
`compression_store.CompressionStore.store`). Emitted only when
compression saved >20% of lines.
Parity result: `[diff_compressor ] total=20 matched=20 skipped=0 diffed=0`.
# Information preservation hardening
Three pass-through paths inherited from Python that we keep deliberate
(would lose info if we changed them):
- Below `min_lines_for_ccr` (50): return input unchanged.
- No diff sections parsed: return input unchanged.
- Below 20% compression savings: emit compressed output but no CCR
marker (the original is the cheaper representation anyway).
Plus a parity-bound subtlety: `compressed_line_count` is captured BEFORE
the CCR retrieval marker is appended, both for the marker text
(`compressed to N`) and the result field. The output string therefore
ends up with one more line than the field reports — by design, matching
Python exactly. An off-by-one bug from recounting after appending the
CCR marker was caught and pinned by a synthetic 8-file diff test.
# Observability — the Rust escape hatch
Python's `DiffCompressionResult` has thin observability: input/output
line counts, additions/deletions, hunks_kept/removed, files_affected,
cache_key. The Rust port adds a sidecar `DiffCompressorStats` struct
with metrics Python doesn't emit:
- `files_dropped: Vec<String>` — names (old → new path) of files
silently discarded by the `max_files` cap. Python loses these.
- `hunks_dropped_per_file: BTreeMap<String, usize>` — per-file hunk
drops, stable iteration via `BTreeMap`.
- `context_lines_input` / `context_lines_kept` / `context_lines_trimmed`
— directly proxies info loss from the context trim.
- `largest_hunk_kept_lines` / `largest_hunk_dropped_lines` — outlier
detection (a single huge dropped hunk is much worse than many small).
- `parse_warnings: Vec<String>` — surfaces malformed input rather than
dropping silently.
- `processing_duration_us` — latency budget.
- `cache_key_emitted` + `ccr_skipped_reason: Option<String>` — explicit
signal for "we chose not to emit CCR and this is why".
A `tracing::info!(target: "diff_compressor", ...)` event is emitted on
every call, carrying these fields for OTel scraping in prod. The
sidecar struct is returned alongside via `compress_with_stats`; the
parity-only `compress` API discards it.
# Module layout
- `crates/headroom-core/src/transforms/mod.rs` — namespace, doc comment
with the guiding principle ("information preservation > aggressive
compression") so future ports inherit the philosophy.
- `crates/headroom-core/src/transforms/diff_compressor.rs` — full port
(parser, scorer, hunk selector, context trimmer, formatter, CCR layer,
stats, tracing).
# Dependencies added to headroom-core
- `md-5 = "0.10"` — for the CCR cache_key (matches Python MD5[:24]).
- `regex = "1"` — was a transitive dep via tokenizers; now a direct
dependency for the hunk-header parser and priority patterns.
# Tests
6 unit tests covering pass-through paths, MD5 hex truncation, the
Python `split("\n")` line-count semantics, sidecar stats emission,
and a synthetic 8-file diff that locks the byte-equal behavior found
in the parity fixtures.
|
||
|
|
a23ee8e70b |
feat(rust): HfTokenizer::from_pretrained — HuggingFace Hub auto-download
Stage 2.1: closes the loop on the HuggingFace tokenizer story. Stage 2
shipped `HfTokenizer::from_bytes`/`from_file`, which required callers to
manage their own tokenizer.json files. This adds the third constructor:
let t = HfTokenizer::from_pretrained("CohereForAI/c4ai-command-r-v01")?;
register_hf("command-", t);
`from_pretrained` is a thin wrapper around the `hf-hub` crate's blocking
`ureq` API. First call downloads `tokenizer.json` to `~/.cache/huggingface/
hub` (or `$HF_HOME` if set); subsequent calls reuse the on-disk cache. Uses
the `main` revision; gated repos (Llama, Mistral) require `HF_TOKEN` in env
or `~/.cache/huggingface/token`.
Also adds `try_register_hf(prefix, repo)` as the obvious one-liner for
proxy startup code:
let _ = try_register_hf("command-", "CohereForAI/c4ai-command-r-v01");
let _ = try_register_hf("mistral-", "mistralai/Mistral-7B-v0.1");
Each call is independent — a download failure for one model (e.g. gated
without a token) does not affect others.
`HfTokenizerError` gains a new `Hub` variant so callers can distinguish
"couldn't fetch" from "fetched but malformed" — relevant when deciding
whether to retry, surface to the user, or fall back to the estimator.
Why blocking, not async: `from_pretrained` is called once at startup. A
sync API works from `main()`, from a `OnceLock` initializer, or from
`tokio::task::spawn_blocking` if a tokio caller needs it later. The async
hf-hub backend would force callers to await at startup, which doesn't fit
the `register_hf` registry pattern.
Why rustls, not native-tls: keeps the binary statically linkable for AWS
deploys (no system OpenSSL dependency).
Tests: a network-dependent integration test (`#[ignore]`d in CI; hits HF
for `gpt2`, ~1.4 MB) verifies the real download + load + count path. A
non-network negative test verifies that an invalid repo name surfaces as
`HfTokenizerError::Hub`, not a panic. 44 unit tests + 5 proptests +
1 doctest pass; parity stays 40/40 byte-equal.
|
||
|
|
9ce1c01b87 |
feat(rust): tokenizer crate with tiktoken-rs + HuggingFace + estimator
Stage 2 of the Rust port: a `headroom_core::tokenizer` module mirroring the Python `headroom.tokenizers` surface, with three backends behind a single `Tokenizer` trait. Backends, in dispatch order: 1. HuggingFace (`HfTokenizer`) — pure-Rust `tokenizers` crate loading any public `tokenizer.json`. Covers the gap between OpenAI (tiktoken) and the Anthropic/Gemini estimator: Cohere `command-*`, Llama-3.x, Mistral, Qwen, BERT, T5, etc. Construct from bytes or a file path; register against a model-name prefix via `register_hf` for automatic dispatch. No `hf-hub` auto-download yet — keeps networking, auth, and `~/.cache/huggingface` out of core. Longest-prefix wins; lookups are RwLock-protected. 2. Tiktoken (`TiktokenCounter`) — `tiktoken-rs` 0.11 BPE for OpenAI / o-series families. Byte-identical to Python `tiktoken` for ordinary text. Lazy shared `Arc<CoreBPE>` per encoding (o200k_base, cl100k_base, p50k_base, r50k_base). 3. Estimation (`EstimatingCounter`) — `chars / cpt` last-resort fallback. Matches Python's `max(1, int(len(text) / cpt + 0.5))` round-half-up formula (a self-review caught and fixed an earlier `ceil`-based version that diverged in the middle of the range, e.g. 5 chars at 4.0 cpt). Tests: 43 unit tests + 5 proptests; parity 40/40 byte-equal. Bench: criterion baseline on small/medium/large inputs. Workspace MSRV bumped 1.78 → 1.80 for `LazyLock`/`OnceLock`. No proxy wiring. Library-only; production behavior unchanged. |
||
|
|
c2749c0fb6 |
docs(rust): lockfile + RUST_DEV.md for proxy CLI
Cargo.lock: pick up tokio-util added in the WS half-close fix. RUST_DEV.md: document how to run headroom-proxy in passthrough mode (listen + upstream flags, e2e test gate, env vars). |
||
|
|
128a910ebb |
feat(rust): axum reverse proxy skeleton + http catch-all (phase-1)
Builds out crates/headroom-proxy from a /healthz stub into a transparent reverse proxy: catch-all router that forwards every method/path/query to --upstream verbatim, streaming both request and response bodies through reqwest without buffering. Adds clap-based config (CLI + env), thiserror error type with sane upstream-status mapping, JSON tracing-subscriber logging, and graceful shutdown. The library surface (build_app, AppState, Config) is reused by the integration tests. |
||
|
|
0414cb70e4 |
feat(rust): scaffold workspace + parity harness (phase-0)
Bootstrap the Rust port of Headroom. Additive only — no existing Python code modified. Ships the workshop, not the widgets. Layout Cargo.toml (workspace) + rust-toolchain.toml crates/headroom-core — transform library, stub only crates/headroom-proxy — axum binary, /healthz only crates/headroom-py — PyO3 cdylib, exposes headroom._core.hello() crates/headroom-parity — Rust-vs-Python oracle harness + parity-run CLI Tooling Makefile: test, test-parity, bench, build-proxy, build-wheel, fmt, lint .github/workflows/rust.yml: test, wheels (linux/mac), audit, parity-nightly deny.toml for cargo-deny Parity corpus tests/parity/recorder.py + scripts/record_fixtures.py 125 recorded fixtures across 5 leaf transforms (ccr, tokenizer, log_compressor, diff_compressor, cache_aligner) Docs RUST_DEV.md — developer setup and workspace reference docs/spec/022-rust-migration.md — migration plan and stage breakdown .gitignore: whitelist scripts/record_fixtures.py; ignore target/ Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |