headroom/docs/rtk-architecture.md

Ignoring revisions in .git-blame-ignore-revs. Click here to bypass and see the normal blame view.

123 lines
4.8 KiB
Markdown
Raw Normal View History

fix(observability): wire Phase G PR-G3 RTK + proxy metrics (H-blocker) Phase H ("retire the Python proxy") needs cache-hit-rate parity between the Rust and Python proxies during canary. This PR lands the per-invocation RTK metrics and the proxy-side observability surface that the canary gate depends on. Rust observability: - `proxy_cache_hit_rate_per_session{provider}` — histogram, emitted per session at SSE state-machine close (Anthropic message_delta, OpenAI Chat final usage chunk, OpenAI Responses response.completed). The Phase H canary gate metric. - `proxy_compression_ratio_by_strategy{strategy, content_type}` — histogram; one sample per shrunk block. - `proxy_compression_rejected_by_token_check_total{strategy}` — counter for tokenizer-validated rejections. - `proxy_passthrough_bytes_modified_total{path}` — counter (must stay 0 outside compression hot path; alarmable via PromQL rate). - `proxy_rate_limit_remaining_{requests,tokens,input_tokens,output_tokens}{provider}` — gauges populated from anthropic-ratelimit-* / x-ratelimit-* headers. - `proxy_service_tier_count_total{tier}` and `proxy_response_status_count_total{status}` — counters for Responses-API outcome telemetry. - `proxy_image_generation_call_log_redacted_total` — counter. - `wrap_rtk_invocations_total{tool}` and `wrap_rtk_tokens_saved_per_session` — RTK metrics exposed via the proxy's /metrics scrape so wrap-side tail can increment through one observability surface. All metric names and label keys live in a single `observability/metric_names.rs` constants module per realignment build-constraint "configurable". Bounded label vocabularies (service_tier, response_status, provider) are defined alongside. Python (P4-45): - `headroom/proxy/request_logger.py` — base64-image payloads in request/response logs over 1024 bytes are replaced with `<image:base64-redacted bytes=N>` placeholders. Walks Anthropic source.data and OpenAI data URLs. No regexes — substring + density heuristic. Tests: - `crates/headroom-proxy/tests/integration_metrics.rs` — 6 tests covering cache-hit-rate, compression-ratio, passthrough-bytes, service-tier, response-status, and rate-limit-snapshot. - `tests/test_image_log_redaction.py` — 13 tests for the Python redaction helper. - Existing tests: 1100+ Rust + 76 Python regression checks green. Docs: - `docs/observability.md` — metric catalogue + PromQL queries. - `docs/rtk-architecture.md` — locks the wrap-CLI-only decision so future contributors don't relitigate proxy-side RTK. No silent fallbacks: zero-denominator cache-hit-rate logs and skips rather than synthesising 0.0. Unparseable rate-limit headers stay None rather than coerced to 0. Missing upstream JSON fields log + skip emit rather than fabricating data.
2026-05-22 13:18:42 -07:00
# RTK architecture — why wrap-CLI only
**Status:** decided. Locked at Phase G PR-G3 (2026-05).
**Owner:** Headroom realignment.
## TL;DR
**RTK is a wrap-CLI hook, not a proxy-side compressor.** The Headroom
proxy does NOT invoke RTK on tool-result content. Future contributors
who consider moving RTK into the proxy hot path: read this doc first.
## Background
RTK (Realtime Token Kompress) rewrites shell **commands** at exec
time so that a `git diff` or `grep` invocation emits a more
compressed output before the agent ever ingests it. RTK runs in the
wrap-CLI tail — `headroom wrap claude`, `headroom wrap codex`, etc.
— where it installs a `~/.rtk/bin/rtk` shim ahead of the agent CLI
and intercepts shelled-out subprocesses.
It surfaces value in two places:
1. **Tokens saved per invocation** — measured by `rtk gain --format json`.
2. **Tokens saved per session** — aggregated at wrap-session end.
Both signals feed `wrap_rtk_invocations_total` and
`wrap_rtk_tokens_saved_per_session` (registered by the Rust proxy's
observability surface so a single `/metrics` scrape exposes the full
picture).
## Proxy-side RTK was considered and rejected
At Phase G scoping, three reviewers floated the idea of invoking
RTK on the **proxy** side: when a `tool_result` block flows
upstream, dispatch it through RTK to shrink the content before it
hits the model.
**Decision: rejected.** Three load-bearing reasons.
### 1. Cache hot zone risk
The proxy's Phase B cache-safety contract pins `tool_result`
content as part of the cache hot zone. Compression there bursts
the prompt cache because the rewritten bytes diverge from the
canonical wire bytes the upstream cached. Phase B PR-B2 → PR-B7
spent ~3000 LOC carving the live-zone-only surface specifically
to prevent this class of cache-invalidation. Inserting RTK
proxy-side would re-introduce it.
### 2. Parallel implementation with `log_compressor.rs`
The Rust proxy already has a `crates/headroom-core/src/transforms/log_compressor.rs`
that compresses **tool output text** in the live zone. It uses the
same heuristics RTK uses (whitespace de-dup, line de-dup,
file-listing collapse) but invoked at the proxy's per-block
dispatcher rather than at the shell exec boundary. Adding RTK
proxy-side would mean two implementations of the same compression
in the same hot path; "no silent fallbacks, no parallel impls" is
explicit project policy.
### 3. Command-rewrite vs output-rewrite — different value propositions
RTK rewrites **commands** before they execute. The
`git log --oneline` you typed becomes `git log --oneline -n 50`
because RTK has learned that the first 50 commits are usually
enough context. That's a fundamentally different mechanism from
compressing the **output** of an unmodified command. A proxy-side
invocation would skip the command-rewrite half — the half that
generates the largest savings on heavy shell workloads — and only
catch the output side, which is already covered by
`log_compressor` and `code_compressor`.
## What the proxy does provide
Per Phase G PR-G3, the proxy exposes RTK-derived metrics via its
registry:
- `wrap_rtk_invocations_total{tool}` — driven by the wrap-CLI
polling `rtk gain --format json` and incrementing the registered
counter by the delta since last poll.
- `wrap_rtk_tokens_saved_per_session` — emitted at wrap-session
close.
This keeps the operator dashboard single-pane-of-glass without
re-implementing RTK inside the proxy.
## What the wrap CLI does
Every `headroom wrap <agent>` subcommand:
1. Ensures the RTK binary is installed via `_ensure_rtk_binary()`.
2. Injects the `<!-- headroom:rtk-instructions -->` block into the
agent's instruction file (e.g. `AGENTS.md`, `.cursorrules`).
3. Spawns the proxy and the agent CLI side-by-side.
4. Polls `rtk gain --format json` on a 5-second memoization window
and feeds the delta into the proxy's metric registry.
See `headroom/cli/wrap/` for the per-agent shims.
## Re-litigation policy
A change to this architecture should:
1. Quote the live-zone-only contract from
`REALIGNMENT/04-phase-B-live-zone.md` and explain why the
cache-burst risk is acceptable.
2. Show measurements (not estimates) that proxy-side RTK adds value
beyond `log_compressor.rs` on real production traffic.
3. Have an exit ramp: a CLI flag to disable proxy-side RTK without
reverting the wrap-CLI integration.
Without all three, treat the proposal as a regression and link this
doc.
## References
- `REALIGNMENT/09-phase-G-rtk-observability.md` — Phase G plan.
- `REALIGNMENT/04-phase-B-live-zone.md` — cache hot-zone contract.
- `headroom/cli/wrap/` — wrap-CLI implementation.
- `crates/headroom-core/src/transforms/log_compressor.rs` — the
proxy-side log compressor RTK would parallel.
- 2026-05-01 user direction message archived in
`project_compression_realignment_2026_05` memory note.