mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-27 14:17:10 -04:00
Phase H ("retire the Python proxy") needs cache-hit-rate parity
between the Rust and Python proxies during canary. This PR lands
the per-invocation RTK metrics and the proxy-side observability
surface that the canary gate depends on.
Rust observability:
- `proxy_cache_hit_rate_per_session{provider}` — histogram, emitted
per session at SSE state-machine close (Anthropic message_delta,
OpenAI Chat final usage chunk, OpenAI Responses response.completed).
The Phase H canary gate metric.
- `proxy_compression_ratio_by_strategy{strategy, content_type}` —
histogram; one sample per shrunk block.
- `proxy_compression_rejected_by_token_check_total{strategy}` —
counter for tokenizer-validated rejections.
- `proxy_passthrough_bytes_modified_total{path}` — counter (must
stay 0 outside compression hot path; alarmable via PromQL rate).
- `proxy_rate_limit_remaining_{requests,tokens,input_tokens,output_tokens}{provider}` —
gauges populated from anthropic-ratelimit-* / x-ratelimit-* headers.
- `proxy_service_tier_count_total{tier}` and
`proxy_response_status_count_total{status}` — counters for
Responses-API outcome telemetry.
- `proxy_image_generation_call_log_redacted_total` — counter.
- `wrap_rtk_invocations_total{tool}` and
`wrap_rtk_tokens_saved_per_session` — RTK metrics exposed via
the proxy's /metrics scrape so wrap-side tail can increment
through one observability surface.
All metric names and label keys live in a single
`observability/metric_names.rs` constants module per realignment
build-constraint "configurable". Bounded label vocabularies
(service_tier, response_status, provider) are defined alongside.
Python (P4-45):
- `headroom/proxy/request_logger.py` — base64-image payloads in
request/response logs over 1024 bytes are replaced with
`<image:base64-redacted bytes=N>` placeholders. Walks Anthropic
source.data and OpenAI data URLs. No regexes — substring +
density heuristic.
Tests:
- `crates/headroom-proxy/tests/integration_metrics.rs` — 6 tests
covering cache-hit-rate, compression-ratio, passthrough-bytes,
service-tier, response-status, and rate-limit-snapshot.
- `tests/test_image_log_redaction.py` — 13 tests for the Python
redaction helper.
- Existing tests: 1100+ Rust + 76 Python regression checks green.
Docs:
- `docs/observability.md` — metric catalogue + PromQL queries.
- `docs/rtk-architecture.md` — locks the wrap-CLI-only decision so
future contributors don't relitigate proxy-side RTK.
No silent fallbacks: zero-denominator cache-hit-rate logs and
skips rather than synthesising 0.0. Unparseable rate-limit headers
stay None rather than coerced to 0. Missing upstream JSON fields
log + skip emit rather than fabricating data.
122 lines
4.8 KiB
Markdown
122 lines
4.8 KiB
Markdown
# RTK architecture — why wrap-CLI only
|
|
|
|
**Status:** decided. Locked at Phase G PR-G3 (2026-05).
|
|
**Owner:** Headroom realignment.
|
|
|
|
## TL;DR
|
|
|
|
**RTK is a wrap-CLI hook, not a proxy-side compressor.** The Headroom
|
|
proxy does NOT invoke RTK on tool-result content. Future contributors
|
|
who consider moving RTK into the proxy hot path: read this doc first.
|
|
|
|
## Background
|
|
|
|
RTK (Realtime Token Kompress) rewrites shell **commands** at exec
|
|
time so that a `git diff` or `grep` invocation emits a more
|
|
compressed output before the agent ever ingests it. RTK runs in the
|
|
wrap-CLI tail — `headroom wrap claude`, `headroom wrap codex`, etc.
|
|
— where it installs a `~/.rtk/bin/rtk` shim ahead of the agent CLI
|
|
and intercepts shelled-out subprocesses.
|
|
|
|
It surfaces value in two places:
|
|
1. **Tokens saved per invocation** — measured by `rtk gain --format json`.
|
|
2. **Tokens saved per session** — aggregated at wrap-session end.
|
|
|
|
Both signals feed `wrap_rtk_invocations_total` and
|
|
`wrap_rtk_tokens_saved_per_session` (registered by the Rust proxy's
|
|
observability surface so a single `/metrics` scrape exposes the full
|
|
picture).
|
|
|
|
## Proxy-side RTK was considered and rejected
|
|
|
|
At Phase G scoping, three reviewers floated the idea of invoking
|
|
RTK on the **proxy** side: when a `tool_result` block flows
|
|
upstream, dispatch it through RTK to shrink the content before it
|
|
hits the model.
|
|
|
|
**Decision: rejected.** Three load-bearing reasons.
|
|
|
|
### 1. Cache hot zone risk
|
|
|
|
The proxy's Phase B cache-safety contract pins `tool_result`
|
|
content as part of the cache hot zone. Compression there bursts
|
|
the prompt cache because the rewritten bytes diverge from the
|
|
canonical wire bytes the upstream cached. Phase B PR-B2 → PR-B7
|
|
spent ~3000 LOC carving the live-zone-only surface specifically
|
|
to prevent this class of cache-invalidation. Inserting RTK
|
|
proxy-side would re-introduce it.
|
|
|
|
### 2. Parallel implementation with `log_compressor.rs`
|
|
|
|
The Rust proxy already has a `crates/headroom-core/src/transforms/log_compressor.rs`
|
|
that compresses **tool output text** in the live zone. It uses the
|
|
same heuristics RTK uses (whitespace de-dup, line de-dup,
|
|
file-listing collapse) but invoked at the proxy's per-block
|
|
dispatcher rather than at the shell exec boundary. Adding RTK
|
|
proxy-side would mean two implementations of the same compression
|
|
in the same hot path; "no silent fallbacks, no parallel impls" is
|
|
explicit project policy.
|
|
|
|
### 3. Command-rewrite vs output-rewrite — different value propositions
|
|
|
|
RTK rewrites **commands** before they execute. The
|
|
`git log --oneline` you typed becomes `git log --oneline -n 50`
|
|
because RTK has learned that the first 50 commits are usually
|
|
enough context. That's a fundamentally different mechanism from
|
|
compressing the **output** of an unmodified command. A proxy-side
|
|
invocation would skip the command-rewrite half — the half that
|
|
generates the largest savings on heavy shell workloads — and only
|
|
catch the output side, which is already covered by
|
|
`log_compressor` and `code_compressor`.
|
|
|
|
## What the proxy does provide
|
|
|
|
Per Phase G PR-G3, the proxy exposes RTK-derived metrics via its
|
|
registry:
|
|
|
|
- `wrap_rtk_invocations_total{tool}` — driven by the wrap-CLI
|
|
polling `rtk gain --format json` and incrementing the registered
|
|
counter by the delta since last poll.
|
|
- `wrap_rtk_tokens_saved_per_session` — emitted at wrap-session
|
|
close.
|
|
|
|
This keeps the operator dashboard single-pane-of-glass without
|
|
re-implementing RTK inside the proxy.
|
|
|
|
## What the wrap CLI does
|
|
|
|
Every `headroom wrap <agent>` subcommand:
|
|
|
|
1. Ensures the RTK binary is installed via `_ensure_rtk_binary()`.
|
|
2. Injects the `<!-- headroom:rtk-instructions -->` block into the
|
|
agent's instruction file (e.g. `AGENTS.md`, `.cursorrules`).
|
|
3. Spawns the proxy and the agent CLI side-by-side.
|
|
4. Polls `rtk gain --format json` on a 5-second memoization window
|
|
and feeds the delta into the proxy's metric registry.
|
|
|
|
See `headroom/cli/wrap/` for the per-agent shims.
|
|
|
|
## Re-litigation policy
|
|
|
|
A change to this architecture should:
|
|
|
|
1. Quote the live-zone-only contract from
|
|
`REALIGNMENT/04-phase-B-live-zone.md` and explain why the
|
|
cache-burst risk is acceptable.
|
|
2. Show measurements (not estimates) that proxy-side RTK adds value
|
|
beyond `log_compressor.rs` on real production traffic.
|
|
3. Have an exit ramp: a CLI flag to disable proxy-side RTK without
|
|
reverting the wrap-CLI integration.
|
|
|
|
Without all three, treat the proposal as a regression and link this
|
|
doc.
|
|
|
|
## References
|
|
|
|
- `REALIGNMENT/09-phase-G-rtk-observability.md` — Phase G plan.
|
|
- `REALIGNMENT/04-phase-B-live-zone.md` — cache hot-zone contract.
|
|
- `headroom/cli/wrap/` — wrap-CLI implementation.
|
|
- `crates/headroom-core/src/transforms/log_compressor.rs` — the
|
|
proxy-side log compressor RTK would parallel.
|
|
- 2026-05-01 user direction message archived in
|
|
`project_compression_realignment_2026_05` memory note.
|