mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-10 14:27:00 -04:00
## Description `TextCrusher` (the native extractive prose compressor added in #1171) only handled ASCII: `split_segments` split on `.!?`+whitespace and `tokens` split on whitespace/alphanumeric runs. CJK (Chinese/Japanese/Korean) has neither spaces nor ASCII terminators, so a whole CJK paragraph collapsed into **one segment / one token** — it passed through at ~0% compression, and BM25 relevance + salience scored zero terms. This makes `TextCrusher` CJK-aware. CJK-bearing content takes an ICU (`icu_segmenter`, UAX#29 sentence + dictionary word) segmentation path, with a length fallback for terminator-sparse runs, a local BM25 over the ICU word tokens, and ICU-token salience. Dispatch is on **content only**, so pure-ASCII text is byte-identical to before — the shared `BM25Scorer` and the ASCII path are untouched. It also adds a committed, reproducible answer-retention eval (`benchmarks/i18n_compression_eval.py`) with a deterministic zh/ja/ko CI regression gate, so the improvement below is permanently verifiable rather than a one-off measurement. Extends #1171. ## Type of Change - [x] Bug fix (CJK passed through near-uncompressed) - [x] New feature (CJK segmentation / relevance support) - [x] Performance improvement (CJK now compresses; ICU segmenters cached, not rebuilt per call) ## Changes Made - `is_cjk` predicate gates a CJK path (ideographs, kana, Hangul, CJK punctuation, full/half-width forms). - `split_segments` → ICU `SentenceSegmenter` for CJK + a mandatory length fallback (whitespace / CJK punctuation / hard cap) for terminator-sparse runs; ASCII path unchanged. - `tokens` → ICU `WordSegmenter` (dictionary) for CJK; ASCII path unchanged. - `relevance_cjk`: a local BM25 over ICU word tokens — the shared ASCII `BM25Scorer` scores zero terms for CJK and is parity-locked, so this is an intentional separate scorer (documented in code). - CJK salience uses ICU tokens (whitespace-split gave one giant "word" → zero salience). - `count_tokens`: CJK-aware so `compression_ratio` isn't nonsense for space-free text. - ICU segmenters resolved once in `static LazyLock` (compiled_data is static) instead of rebuilt per call. - New dep `icu_segmenter` 2.2, `compiled_data` only (see Dependency below). - `benchmarks/i18n_compression_eval.py` + `tests/test_transforms/test_text_crusher_cjk_eval.py`: a zh/ja/ko answer-retention eval — a deterministic needle CI gate (always-runs, no external data), real-transcript fidelity with CJK-aware salient, and optional `multi-wiki-qa` natural-data retention (loaded via the `[evals]` `datasets` extra, skipped if absent; data never vendored — CC-BY-NC-SA). ## Testing - [x] Unit tests pass (`pytest` + `cargo test`) - [x] Linting passes (`ruff check`/`format` on the new eval + test — clean) - [ ] Type checking passes (`mypy headroom`) — N/A, the only Python added is a benchmark + test, not `headroom/` source - [x] New tests added for new functionality - [x] Manual testing performed (see Real Behavior Proof) ### Test Output ```text $ cargo test -p headroom-core --lib text_crusher running 12 tests test result: ok. 12 passed; 0 failed; 0 ignored; 0 measured; 841 filtered out $ .venv/bin/python -m pytest tests/test_transforms/test_text_crusher*.py 15 passed $ .venv/bin/python -m pytest tests/test_transforms/test_text_crusher_cjk_eval.py 6 passed # deterministic zh/ja/ko needle CI gate $ cargo clippy -p headroom-core && ruff check benchmarks/i18n_compression_eval.py # both clean ``` ## Real Behavior Proof - Environment: macOS (Darwin 25.3.0), Python in a uv venv, `headroom-core` built via `uv pip install -e .` (maturin), branch `feat/cjk-text-compression`. - Exact command / steps: built `_core`, then ran a mixed Chinese+Japanese doc (no spaces, `。` terminators) through `TextCrusher().compress(doc, "认证令牌缓存策略", 0.3)`; separately evaluated answer-retention on the public CMRC2018 Chinese QA dev set (bury the gold-answer paragraph among 25 distractors, query = the question, compress to 30%, check the gold answer survives), and end-to-end through `ContentRouter`. - Observed result: a mixed Chinese+Japanese doc compressed 189 → 78 tokens (ratio 0.41, kept 3/8 segments) with the query-relevant sentence surviving — before this change the same doc was a single segment → 100% passthrough. On the public CMRC2018 Chinese QA dev set, answer-retention under 30% compression rose 34% → 93% (multiple seeds). End-to-end through `ContentRouter` on real CJK content, aggregate savings rose 16% → 40%. Pure-ASCII (English) output stayed byte-identical (the English parity fixtures did not move). Demo terminal output: ```text ORIGINAL tokens= 189 chars=189 COMPRESS tokens= 78 ratio=0.41 segments kept 3/8 QUERY-RELEVANT sentence survived: True --- compressed output (verbatim kept CJK sentences) --- 认证令牌的缓存策略采用最近最少使用淘汰算法来管理过期条目。 请求重试使用指数退避并设置最大次数上限。 数据备份每天凌晨执行并保留最近三十天的快照。 ``` The committed eval now demonstrates this across all three CJK languages. The deterministic needle gate (in CI via `tests/test_transforms/test_text_crusher_cjk_eval.py`, 6 passed) has TextCrusher keep the query-relevant needle while truncate/random drop it in zh, ja, and ko. On real `multi-wiki-qa` natural data (n=80/lang), query-aware answer-retention is **zh 74% / ja 70% / ko 50%** vs **25–41%** for the truncate/random baselines: ```text === Part A: multi-wiki-qa answer-retention (n=80/lang, target_ratio=0.3) === lang text_crusher truncate random zh-cn 74% 25% 38% ja 70% 31% 39% ko 50% 26% 41% ``` Korean is measurably weaker (ICU has no Korean dictionary and falls back to UAX#29 word-breaking) — still well above baselines, and scoped as a follow-up. - Not tested: the live proxy HTTP path (validated at the `ContentRouter` / `TextCrusher` layer, not via a running proxy); no-space Korean (standard Korean is space-delimited and is covered); non-CJK SE-Asian scripts (out of scope). ## Dependency (per CONTRIBUTING supply-chain policy) `icu_segmenter` 2.2 (ICU4X), `features = ["compiled_data"]`: - **Why this package (vs. ourselves / existing deps):** CJK needs dictionary/UAX#29 segmentation. A hand-rolled char-bigram scored slightly worse on real data (CMRC2018 answer-retention: 92.5% ICU vs 91% bigram, 4 seeds); jieba/lindera are ZH-only or 13–207 MB dicts. ICU4X covers zh/ja/ko in one crate. The existing `unicode-segmentation` does UAX#29 only (no CJK dictionary), so it can't word-segment space-free CJK. - **Who maintains it:** the official `unicode-org` ICU4X project; active release cadence (2.2 in 2025); no known CVEs. - **Install surface:** ~13 new pure-Rust crates, no build scripts, no native code, no build/runtime network. `compiled_data` bundles locale data at compile time (hermetic). `auto`/`lstm` deliberately NOT enabled — LSTM covers SE-Asian scripts (Thai/Lao), not CJK, and would pull in `libm` for nothing. - **Why this version:** 2.x is the stabilized ICU4X API (1.x used a different data-provider model); floored at 2.2 (Cargo.lock pins the patch) since segmenter boundaries are observable in output and bumps should be deliberate. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation (CHANGELOG) - [x] My changes generate no new warnings (clippy + fmt clean) - [x] I have added tests that prove my feature works - [x] New and existing unit tests pass locally with my changes - [x] I have updated the CHANGELOG.md ## Additional Notes - **Parity:** the shared `BM25Scorer` (byte-exact parity-locked with `headroom/relevance/bm25.py`) is untouched. `relevance_cjk` is a separate local scorer because the shared one's tokenizer is ASCII-only. The whole CJK path lives in Rust (`text_crusher.py` is a thin wrapper over `_core`), so there is no Python mirror to keep in sync; the parity fixtures stay green (only the CJK `unicode` fixture was re-recorded, intentionally; English fixtures unchanged). - **Known by-design gap (not a bug):** CJK content + a pure-ASCII query yields no token overlap, so relevance falls back to recency + salience (cross-script query matching is unsupported). - The Python added is a benchmark (`benchmarks/i18n_compression_eval.py`) plus its test, not `headroom/` runtime source — both are `ruff`-clean; `mypy headroom` is unaffected. - **License:** the optional Part A pulls `alexandrainst/multi-wiki-qa` (CC-BY-NC-SA-4.0) at run time via the `[evals]` extra and is skipped if absent — the dataset is never vendored into the repo, and the always-run CI gate (Part C) uses only our own deterministic data. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
282 lines
10 KiB
Python
282 lines
10 KiB
Python
#!/usr/bin/env python3
|
|
"""i18n compression-quality eval (zh/ja/ko): does extractive compression keep
|
|
the answer-bearing content in CJK? No LLM/API calls -- fully local.
|
|
|
|
Part C -- our own DETERMINISTIC needle answer-retention (zh/ja/ko): the always-
|
|
runs regression gate. A distinctive needle sentence is buried (in the middle) in
|
|
language-matched distractor sentences; compress query-aware; assert the needle
|
|
survives. No external data. TextCrusher (query-aware) vs truncate (keep-recent)
|
|
vs random baselines.
|
|
|
|
Part B -- real-transcript fidelity with CJK-aware salient: optional, anonymized.
|
|
|
|
Part A -- natural-data answer-retention on alexandrainst/multi-wiki-qa
|
|
(zh-cn/ja/ko): optional, via the [evals] datasets extra, skipped if absent.
|
|
|
|
Usage: python benchmarks/i18n_compression_eval.py [transcript.jsonl]
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import glob
|
|
import os
|
|
import random
|
|
import re
|
|
import sys
|
|
import time
|
|
|
|
from headroom.transforms.text_crusher import TextCrusher
|
|
|
|
_REDACT = [
|
|
(re.compile(r"/Users/[^/\s]+"), "/Users/USER"),
|
|
(re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.-]+\b"), "EMAIL"),
|
|
(re.compile(r"\b(?:sk|pk|ghp|gho|xox[baprs])-[A-Za-z0-9_-]{10,}\b"), "TOKEN"),
|
|
(re.compile(r"\b[A-Fa-f0-9]{40,}\b"), "HEX"),
|
|
]
|
|
# Split on ASCII and full-width CJK terminators so baselines segment CJK too.
|
|
_SEG = re.compile(r"(?<=[.!?。!?])\s*|\n+")
|
|
_CJK_RUN = re.compile(r"[㐀-鿿-ヿ가-]+")
|
|
|
|
|
|
def anon(t: str) -> str:
|
|
for rx, rep in _REDACT:
|
|
t = rx.sub(rep, t)
|
|
return t
|
|
|
|
|
|
def norm(s: str) -> str:
|
|
# CJK has no spaces; drop all whitespace so substring match is robust.
|
|
return re.sub(r"\s+", "", s.lower())
|
|
|
|
|
|
def _segs(text: str) -> list[str]:
|
|
return [s for s in _SEG.split(text) if s.strip()]
|
|
|
|
|
|
def truncate_keep_last(text: str, ratio: float) -> str:
|
|
segs = _segs(text)
|
|
budget = int(sum(len(s) for s in segs) * ratio)
|
|
kept: list[str] = []
|
|
c = 0
|
|
for s in reversed(segs):
|
|
if c >= budget:
|
|
break
|
|
kept.append(s)
|
|
c += len(s)
|
|
return "".join(reversed(kept))
|
|
|
|
|
|
def random_keep(text: str, ratio: float, seed: int) -> str:
|
|
segs = _segs(text)
|
|
idx = list(range(len(segs)))
|
|
random.Random(seed).shuffle(idx)
|
|
budget = int(sum(len(s) for s in segs) * ratio)
|
|
kept: set[int] = set()
|
|
c = 0
|
|
for i in idx:
|
|
if c >= budget:
|
|
break
|
|
kept.add(i)
|
|
c += len(segs[i])
|
|
return "".join(segs[i] for i in sorted(kept))
|
|
|
|
|
|
# --- Part C: deterministic needle retention (zh / ja / ko) ---------------------
|
|
|
|
# Each needle carries a distinctive verbatim KEY that must survive. Distractors
|
|
# are generated (deterministic, distinct, topic-unrelated to the query) so the
|
|
# haystack is large enough to FORCE real compression -- the needle only survives
|
|
# under TextCrusher because it is query-relevant, not because of passthrough.
|
|
_NEEDLES = {
|
|
"zh": {
|
|
"query": "认证令牌缓存淘汰策略",
|
|
"key": "最近最少使用淘汰",
|
|
"needle": "认证令牌的缓存采用最近最少使用淘汰算法来管理过期条目。",
|
|
"distractor": lambda i: f"第{i}号监控服务器的日志显示子系统{i}今天运行平稳没有出现异常。",
|
|
},
|
|
"ja": {
|
|
"query": "認証トークン キャッシュ 破棄 アルゴリズム",
|
|
"key": "最長未使用",
|
|
"needle": "認証トークンのキャッシュは最長未使用アルゴリズムで管理される。",
|
|
"distractor": lambda i: (
|
|
f"{i}番目の監視サーバーのログには{i}番のサブシステムが本日も正常に稼働したと記録されている。"
|
|
),
|
|
},
|
|
"ko": {
|
|
"query": "인증 토큰 캐시 제거 알고리즘",
|
|
"key": "최근 최소 사용",
|
|
"needle": "인증 토큰 캐시는 최근 최소 사용 알고리즘으로 관리된다.",
|
|
"distractor": lambda i: (
|
|
f"{i}번 모니터링 서버의 로그에는 {i}번 하위 시스템이 오늘도 정상 작동했다고 기록되어 있다."
|
|
),
|
|
},
|
|
}
|
|
|
|
|
|
def _haystack(spec: dict, n_distract: int = 24) -> str:
|
|
half = n_distract // 2
|
|
before = [spec["distractor"](i) for i in range(half)]
|
|
after = [spec["distractor"](i) for i in range(half, n_distract)]
|
|
# needle in the MIDDLE so keep-recent (truncate) reliably misses it.
|
|
return "".join(before + [spec["needle"]] + after)
|
|
|
|
|
|
def retention_synthetic(lang: str, ratio: float = 0.3, seed: int = 0) -> dict[str, bool]:
|
|
spec = _NEEDLES[lang]
|
|
hay = _haystack(spec)
|
|
key = norm(spec["key"])
|
|
tc = TextCrusher()
|
|
out_tc = tc.compress(hay, spec["query"], ratio).compressed
|
|
return {
|
|
"text_crusher": key in norm(out_tc),
|
|
"truncate": key in norm(truncate_keep_last(hay, ratio)),
|
|
"random": key in norm(random_keep(hay, ratio, seed)),
|
|
}
|
|
|
|
|
|
def eval_synthetic(ratio: float = 0.3) -> None:
|
|
print(f"\n=== Part C: synthetic needle retention (zh/ja/ko, target_ratio={ratio}) ===")
|
|
print(f" {'lang':5} {'text_crusher':>13} {'truncate':>9} {'random':>7}")
|
|
for lang in ("zh", "ja", "ko"):
|
|
r = retention_synthetic(lang, ratio)
|
|
print(
|
|
f" {lang:5} {str(r['text_crusher']):>13} {str(r['truncate']):>9} {str(r['random']):>7}"
|
|
)
|
|
print(" (needle must survive under TextCrusher; baselines are the contrast)")
|
|
|
|
|
|
# --- Part B: real CJK transcript fidelity (CJK-aware salient) ------------------
|
|
|
|
# ASCII salient (identifiers/numbers/errors) STILL matters in CJK coding context.
|
|
_SALIENT_ASCII = re.compile(
|
|
r"\b(?:error|exception|fail(?:ed|ure)?|warning|traceback|assert|todo|fixme)\b"
|
|
r"|\b[A-Z]{2,}\b|\b[A-Za-z_][A-Za-z0-9_]*\.[A-Za-z_][A-Za-z0-9_]*\b|\b\d+\b"
|
|
)
|
|
|
|
|
|
def _cjk_hapax(text: str) -> set[str]:
|
|
# distinctive CJK content = char-bigrams occurring exactly once (rare = must-keep)
|
|
grams: dict[str, int] = {}
|
|
for run in _CJK_RUN.findall(text):
|
|
for i in range(len(run) - 1):
|
|
g = run[i : i + 2]
|
|
grams[g] = grams.get(g, 0) + 1
|
|
return {g for g, c in grams.items() if c == 1}
|
|
|
|
|
|
def salient_set(text: str) -> set[str]:
|
|
return set(_SALIENT_ASCII.findall(text)) | _cjk_hapax(text)
|
|
|
|
|
|
def _block_texts(jsonl_path: str, min_chars: int, limit: int) -> list[str]:
|
|
import json
|
|
|
|
out: list[str] = []
|
|
with open(jsonl_path, encoding="utf-8") as fh:
|
|
for line in fh:
|
|
try:
|
|
o = json.loads(line)
|
|
except json.JSONDecodeError:
|
|
continue
|
|
c = (o.get("message") or {}).get("content")
|
|
parts = (
|
|
[c]
|
|
if isinstance(c, str)
|
|
else [
|
|
p["text"] for p in c if isinstance(p, dict) and isinstance(p.get("text"), str)
|
|
]
|
|
if isinstance(c, list)
|
|
else []
|
|
)
|
|
for t in parts:
|
|
if len(t) >= min_chars and _CJK_RUN.search(t): # CJK-bearing only
|
|
out.append(anon(t))
|
|
if len(out) >= limit:
|
|
break
|
|
return out[:limit]
|
|
|
|
|
|
def eval_transcript(
|
|
jsonl_path: str, ratio: float = 0.4, min_chars: int = 600, limit: int = 40
|
|
) -> None:
|
|
blocks = _block_texts(jsonl_path, min_chars, limit)
|
|
if not blocks:
|
|
print(
|
|
f"\n=== Part B: no CJK blocks >= {min_chars} chars in {os.path.basename(jsonl_path)} ==="
|
|
)
|
|
return
|
|
tc = TextCrusher()
|
|
ratios: list[float] = []
|
|
times: list[float] = []
|
|
retentions: list[float] = []
|
|
for b in blocks:
|
|
sal_before = salient_set(b)
|
|
t0 = time.perf_counter()
|
|
out = tc.compress(b, "", ratio).compressed
|
|
times.append((time.perf_counter() - t0) * 1000)
|
|
retentions.append(len(sal_before & salient_set(out)) / max(1, len(sal_before)))
|
|
ratios.append(len(out) / max(1, len(b)))
|
|
n = len(blocks)
|
|
print(
|
|
f"\n=== Part B: real CJK transcript fidelity (n={n}, anonymized, target_ratio={ratio}) ==="
|
|
)
|
|
print(f" mean char-ratio kept: {sum(ratios) / n:.2f}")
|
|
print(f" mean speed: {sum(times) / n:.1f} ms/block")
|
|
print(f" CJK-aware salient retention: {sum(retentions) / n:.1%}")
|
|
|
|
|
|
# --- Part A: optional natural-data retention (multi-wiki-qa zh/ja/ko) ----------
|
|
# Schema verified: row = {id, title, context, question, answers:{text:[...]}}.
|
|
# Answers are guaranteed verbatim substrings of the (long) context; CC-BY-NC-SA.
|
|
|
|
|
|
def eval_multiwiki(
|
|
langs=("zh-cn", "ja", "ko"), n: int = 80, ratio: float = 0.3, seed: int = 0
|
|
) -> None:
|
|
try:
|
|
from datasets import load_dataset
|
|
except ImportError:
|
|
print(
|
|
"\n=== Part A: `datasets` not installed; skipping (pip install headroom-ai[evals]) ==="
|
|
)
|
|
return
|
|
tc = TextCrusher()
|
|
print(f"\n=== Part A: multi-wiki-qa answer-retention (n={n}/lang, target_ratio={ratio}) ===")
|
|
print(f" {'lang':6} {'text_crusher':>13} {'truncate':>9} {'random':>7}")
|
|
for lang in langs:
|
|
try:
|
|
ds = load_dataset("alexandrainst/multi-wiki-qa", lang, split=f"train[:{n * 2}]")
|
|
except Exception as e: # noqa: BLE001 -- optional path, fail-open
|
|
print(f" {lang}: load failed ({e}); skipping")
|
|
continue
|
|
ex = []
|
|
for r in ds:
|
|
ans = r.get("answers")
|
|
a = ans["text"][0] if isinstance(ans, dict) and ans.get("text") else None
|
|
if r.get("context") and r.get("question") and a:
|
|
ex.append((r["context"], r["question"], a))
|
|
random.Random(seed).shuffle(ex)
|
|
ex = ex[:n]
|
|
hit = {"text_crusher": 0, "truncate": 0, "random": 0}
|
|
for ctx, q, ans in ex:
|
|
a = norm(ans)
|
|
hit["text_crusher"] += a in norm(tc.compress(ctx, q, ratio).compressed)
|
|
hit["truncate"] += a in norm(truncate_keep_last(ctx, ratio))
|
|
hit["random"] += a in norm(random_keep(ctx, ratio, seed))
|
|
m = max(1, len(ex))
|
|
print(
|
|
f" {lang:6} {hit['text_crusher'] / m:>12.0%} {hit['truncate'] / m:>9.0%} {hit['random'] / m:>7.0%}"
|
|
)
|
|
|
|
|
|
if __name__ == "__main__":
|
|
eval_synthetic()
|
|
tx = sys.argv[1] if len(sys.argv) > 1 else None
|
|
if tx is None:
|
|
found = glob.glob(os.path.expanduser("~/.claude/projects/*headroom*/*.jsonl"))
|
|
tx = max(found, key=os.path.getsize) if found else None
|
|
if tx and os.path.exists(tx):
|
|
eval_transcript(tx)
|
|
else:
|
|
print("\nno transcript jsonl found; skipping Part B")
|
|
eval_multiwiki()
|