headroom/deploy/beacon/sample-event.json

Ignoring revisions in .git-blame-ignore-revs. Click here to bypass and see the normal blame view.

313 lines
10 KiB
JSON
Raw Permalink Normal View History

fix(telemetry): anonymous compression stats — no prompts, no data (#2728) ## In one line Headroom starts reporting **how well compression is working** — counters and percentages only. **No prompts. No code. No file paths. Nothing about what you're building.** ## Why Right now nobody knows whether compression actually helps real users. You can see your own numbers in `/stats`, but that's it — there's no way to tell whether a given workload compresses well, or why it sometimes doesn't. This closes that loop so we can make compression better for everyone. ## Exactly what gets sent One message per session, and every 5 minutes while you're active: ```json { "session": { "id": "random", "turns": 47, "duration_s": 4210, "seq": 3 }, "tokens": { "original": 890000, "attempted": 410000, "saved": 320000, "tool_saved": 48000, "cache_read": 210000 }, "rates": { "saved_pct": 35.96, "eligible_pct": 46.07, "yield_pct": 78.05, "cache_read_pct": 23.60, "overhead_pct": 1.96 }, "compression": { "transforms": {"crush": 47}, "passthrough_turns": 0 }, "skips": {}, "sources": { "proxy": 47 }, "providers": ["anthropic"], "models": ["claude-sonnet-4-5-20250929"], "failures": 2 } ``` Plus a random install ID, the Headroom version, and OS/architecture (`darwin`, `arm64`). That's the whole thing. A full example lives at `deploy/beacon/sample-event.json`. ## What is never sent - Your prompts or the model's responses - Your code - File paths, project names, repo names - Tool names or MCP server names - Hostname, username, or IP address - Custom or fine-tuned model names (an id like `ft:gpt-4o:acme-corp:…` contains a company name, so only models in a public registry are reported) **This is structural, not a pinky-swear.** Every value in the payload is a number, a fixed word, or a random ID — there is no free-text field anywhere for content to hide in. The receiver (`deploy/beacon/worker.js`, in this repo so you can read it) drops anything not on an explicit allowlist before storing. ## Turning it off Any one of these: ```bash HEADROOM_BEACON=off # or DO_NOT_TRACK=1 # or # offline mode ``` It's on by default, and Headroom says so at startup: ``` Telemetry: anonymous compression stats — never prompts, code, or file paths. Helps us improve compression | Off: HEADROOM_BEACON=off ``` `HEADROOM_TELEMETRY` is a **separate** switch that still only affects local stats. If you had turned that on, this change does not start uploading anything — you answered a different question, and upgrading should not change the answer. ## Why the percentages, not just "tokens saved" "We saved 36%" hides the interesting part. In the example above only **46% of tokens were eligible** for compression at all — the rest is frozen cache prefix and system prompts we deliberately do not touch. Of what we *could* touch, we removed **78%**. Those are two separate problems. Raising eligibility is proxy work; raising yield is compressor work. A single number cannot tell us which to fix. ## Coverage `emit_request_outcome` is a single chokepoint — `handler.metrics.record_request` is called from exactly one place, inside the funnel — so all 30 `RequestOutcome` construction sites are covered: Anthropic, OpenAI, Gemini, Bedrock, batch, streaming, and the long-lived Codex Responses-WS path. The `headroom_compress` MCP path bypassed that funnel and is now wired in separately. It has a different shape (no provider, no upstream latency, and everything handed to the tool is eligible by construction), so `sources` counts turns by origin — MCP turns always read `eligible_pct: 100` and must not drag the proxy's real eligibility ceiling upward. **Subagents.** All subagent traffic through the proxy merges into one session, which is correct for savings and retention but means `turns` conflates fan-out with depth. Fan-out is still derivable — `compression.latency_ms_total / session.duration_s` gives the concurrency ratio (~1x serial, ~4x for four parallel agents), so no extra field is needed. Verified no lost updates under 6-way concurrency (1,200 turns). **Known gap:** `--workers N` gives each process its own aggregator, so one user session becomes up to N. Token totals and fleet rates stay correct; session counts inflate. This matches the existing documented limitation that TOIN state, CostTracker, and the prefix tracker are all per-process. ## Notes for reviewers - **Cumulative snapshots, not deltas.** Every report restates running totals under one session ID, so the highest `seq` per `(install, session)` is the complete session. Dedupe is a window function, and a lost report costs nothing. - **Never breaks the proxy.** Every path swallows its own exceptions; uploads go out on a daemon thread so nothing blocks the request loop. - **Explicit User-Agent is load-bearing.** urllib's default is blocked by Cloudflare (error 1010). Combined with fire-and-forget error handling, that would have failed every upload while looking perfectly healthy. - **The exit flush was broken and is fixed.** `atexit` handed the POST to a daemon thread, and daemon threads are killed before they finish during interpreter shutdown — so nothing was sent. That silently dropped *every session shorter than the 5-minute heartbeat*, plus all short-lived subagent MCP processes. The exit path now posts synchronously with a 2s timeout. - Receiver and query tooling are in `deploy/beacon/`. ## Testing - `python -m headroom.telemetry.session` self-check: dedupe, cumulative totals, dropped-report recovery, payload contains no model id or prompt-derived string, allowlist coverage - 175 telemetry/outcome tests pass; 6 new ones cover the opt-out notice - Verified end to end against a live deployment: client → receiver → storage → query ## Still to do before release The default endpoint currently points at a temporary `workers.dev` URL. It needs to move to a Headroom-owned hostname before this ships in a tagged release — noted inline at `DEFAULT_ENDPOINT`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 05:43:25 -07:00
{
"resourceLogs": [
{
"resource": {
"attributes": [
{
"key": "service.name",
"value": {
"stringValue": "headroom"
}
},
{
"key": "service.version",
"value": {
"stringValue": "0.34.0"
}
},
{
"key": "headroom.install_id",
"value": {
"stringValue": "00000000000000000000000000000000"
}
},
{
"key": "os.type",
"value": {
"stringValue": "darwin"
}
},
{
"key": "host.arch",
"value": {
"stringValue": "arm64"
}
}
]
},
"scopeLogs": [
{
"scope": {
"name": "headroom.telemetry.session"
},
"logRecords": [
{
"timeUnixNano": "1785731364402434048",
"body": {
"kvlistValue": {
"values": [
{
"key": "schema_version",
"value": {
"intValue": "1"
}
},
{
"key": "session",
"value": {
"kvlistValue": {
"values": [
{
"key": "id",
"value": {
"stringValue": "sample0000000001"
}
},
{
"key": "seq",
"value": {
"intValue": "0"
}
},
{
"key": "duration_s",
"value": {
"intValue": "4210"
}
},
{
"key": "turns",
"value": {
"intValue": "47"
}
},
{
"key": "ended",
"value": {
"stringValue": "active"
}
},
{
"key": "final",
"value": {
"boolValue": false
}
}
]
}
}
},
{
"key": "tokens",
"value": {
"kvlistValue": {
"values": [
{
"key": "original",
"value": {
"intValue": "890000"
}
},
{
"key": "attempted",
"value": {
"intValue": "410000"
}
},
{
"key": "input",
"value": {
"intValue": "570000"
}
},
{
"key": "output",
"value": {
"intValue": "41000"
}
},
{
"key": "saved",
"value": {
"intValue": "320000"
}
},
{
"key": "tool_saved",
"value": {
"intValue": "48000"
}
},
{
"key": "cache_read",
"value": {
"intValue": "210000"
}
},
{
"key": "cache_write",
"value": {
"intValue": "30000"
}
},
{
"key": "uncached",
"value": {
"intValue": "650000"
}
}
]
}
}
},
{
"key": "rates",
"value": {
"kvlistValue": {
"values": [
{
"key": "saved_pct",
"value": {
"doubleValue": 35.96
}
},
{
"key": "eligible_pct",
"value": {
"doubleValue": 46.07
}
},
{
"key": "yield_pct",
"value": {
"doubleValue": 78.05
}
},
{
"key": "cache_read_pct",
"value": {
"doubleValue": 23.6
}
},
{
"key": "overhead_pct",
"value": {
"doubleValue": 1.96
}
}
]
}
}
},
{
"key": "compression",
"value": {
"kvlistValue": {
"values": [
{
"key": "transforms",
"value": {
"kvlistValue": {
"values": [
{
"key": "crush",
"value": {
"intValue": "47"
}
}
]
}
}
},
{
"key": "overhead_ms_total",
"value": {
"intValue": "1840"
}
},
{
"key": "latency_ms_total",
"value": {
"intValue": "94000"
}
},
{
"key": "passthrough_turns",
"value": {
"intValue": "0"
}
},
{
"key": "response_cache_hits",
"value": {
"intValue": "3"
}
}
]
}
}
},
{
"key": "skips",
"value": {
"kvlistValue": {
"values": []
}
}
},
{
"key": "providers",
"value": {
"arrayValue": {
"values": [
{
"stringValue": "anthropic"
}
]
}
}
},
{
"key": "models",
"value": {
"arrayValue": {
"values": [
{
"stringValue": "claude-sonnet-4-5-20250929"
}
]
}
}
},
{
"key": "failures",
"value": {
"intValue": "2"
}
fix(beacon): split session failures by status code (#2815) ## Description The session beacon reports `failures` as a single count, incremented whenever a turn ends `>= 500` (`headroom/telemetry/session.py`). Across the current corpus that reads **3,969 failures on 595,445 turns (0.67%)** — and the number cannot answer the only question anyone asks of it: an Anthropic `529` is the provider shedding load and there is nothing to fix; a `500` is usually ours. Today the two are indistinguishable, so diagnosis falls back to inference from time-of-day curves and per-install concentration. This counts the status alongside the total. ```json "failures": 3, "failure_statuses": {"529": 2, "500": 1} ``` Motivating investigation on the live corpus (0.67% of turns, 6% of sessions, 63% of all failures from 48 installs, a 2.5% plateau at 08–11 UTC decaying to 0.03% during the fleet's busiest hour) strongly suggests provider-side 529 after retry exhaustion — but "strongly suggests" is exactly the gap this field closes. ## Type of Change - [x] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [ ] Documentation update - [ ] Performance improvement - [ ] Code refactoring (no functional changes) ## Changes Made - **`headroom/telemetry/session.py`** — `_Session.failure_statuses`, incremented next to `failures` in `record_outcome`. Keys are the bare status string for the 5xx range, `"other"` beyond it. Emitted as a sibling of `failures` in `payload()`. - **`deploy/beacon/worker.js`** — `failure_statuses` added to `ALLOWED_KEYS`. Without this the ingest allowlist silently drops it. - **`deploy/beacon/sample-event.json`** — sample carries the new key in OTLP `kvlistValue` form. ### Why no slug bounding `skips` runs values through `_safe_slug` because they arrive as free strings. A status code is an `int` the proxy itself produced; the `500 <= status < 600` check is what keeps a garbage value from inventing map keys. Nothing here is user-derived, so the field stays content-free. ### Why `schema_version` stays 1 Additive, matching the precedent set by #2796, which added `tokens.tool_saved` and the two `all_layers_*` rates without a bump. Bumping signals a break to consumers when nothing about older rows becomes invalid. ## Testing - [x] Unit tests pass (`pytest`) — the module's own self-check, extended - [x] Linting passes (`ruff check .`) - [ ] Type checking passes (`mypy headroom`) — see note - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text $ python -m headroom.telemetry.session ok $ ruff check headroom/telemetry/session.py All checks passed! $ ruff format --check headroom/telemetry/session.py 1 file already formatted $ mypy --python-version 3.12 headroom/telemetry/session.py Success: no issues found in 1 source file # --python-version 3.12 only to skip a pre-existing numpy-stub syntax error the # repo's python_version = "3.10" triggers locally; unrelated to this diff. $ node --check deploy/beacon/worker.js # ok $ python -c "import json; json.load(open('deploy/beacon/sample-event.json'))" # parses ``` The self-check in `headroom/telemetry/session.py` now records two 529s and one 500 and asserts both the total and the split: ```python assert emitted[-1]["failures"] == 3 assert emitted[-1]["failure_statuses"] == {"529": 2, "500": 1} ``` plus `assert event["failure_statuses"] == {}` on the clean-session path. ## Real Behavior Proof - **Environment:** macOS 25.4.0, Python 3.12 venv, this branch. - **Exact command / steps:** drive `SessionAggregator` with three failing outcomes and encode the payload through the same `_any_value` the wire uses. ```text payload: 3 {'529': 2, '500': 1} otlp : {"kvlistValue": {"values": [{"key": "529", "value": {"intValue": "2"}}, {"key": "500", "value": {"intValue": "1"}}]}} ``` The OTLP form matches `deploy/beacon/sample-event.json` byte-for-byte in shape, and `unwrap()` in `worker.js` turns `kvlistValue` back into a plain object, so it lands in R2 as `{"529": 2, "500": 1}` — the same shape as `skips`, which DuckDB reads as `MAP(VARCHAR, BIGINT)`. - **Observed result:** as above. Verified against the live corpus that schema evolution here is already routine — 3,836 of 3,884 existing rows have `rates.all_layers_saved_pct = NULL` from #2796 landing mid-corpus, and every report still runs. - **Not tested:** the deployed Worker (no staging R2 binding locally); `node --check` covers syntax only. The allowlist addition is one array entry consumed by the existing `pick()`. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [x] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [x] I did **not** edit `CHANGELOG.md` ## Screenshots (if applicable) N/A — wire-format change, covered by the output above. ## Additional Notes **Deploy order matters.** The Worker allowlist drops unknown keys, so `deploy/beacon/worker.js` must be deployed *before* a client release that emits the field — otherwise it is discarded at the door. No corruption either way, just missing data until the Worker catches up. **Old data is unaffected.** R2 objects are immutable NDJSON written per request; nothing rewrites history. The corpus reader already passes `union_by_name = true`, which fills the column with NULL for rows written before this ships.
2026-08-05 17:01:57 -07:00
},
{
"key": "failure_statuses",
"value": {
"kvlistValue": {
"values": [
{
"key": "529",
"value": {
"intValue": "2"
}
}
]
}
}
fix(telemetry): anonymous compression stats — no prompts, no data (#2728) ## In one line Headroom starts reporting **how well compression is working** — counters and percentages only. **No prompts. No code. No file paths. Nothing about what you're building.** ## Why Right now nobody knows whether compression actually helps real users. You can see your own numbers in `/stats`, but that's it — there's no way to tell whether a given workload compresses well, or why it sometimes doesn't. This closes that loop so we can make compression better for everyone. ## Exactly what gets sent One message per session, and every 5 minutes while you're active: ```json { "session": { "id": "random", "turns": 47, "duration_s": 4210, "seq": 3 }, "tokens": { "original": 890000, "attempted": 410000, "saved": 320000, "tool_saved": 48000, "cache_read": 210000 }, "rates": { "saved_pct": 35.96, "eligible_pct": 46.07, "yield_pct": 78.05, "cache_read_pct": 23.60, "overhead_pct": 1.96 }, "compression": { "transforms": {"crush": 47}, "passthrough_turns": 0 }, "skips": {}, "sources": { "proxy": 47 }, "providers": ["anthropic"], "models": ["claude-sonnet-4-5-20250929"], "failures": 2 } ``` Plus a random install ID, the Headroom version, and OS/architecture (`darwin`, `arm64`). That's the whole thing. A full example lives at `deploy/beacon/sample-event.json`. ## What is never sent - Your prompts or the model's responses - Your code - File paths, project names, repo names - Tool names or MCP server names - Hostname, username, or IP address - Custom or fine-tuned model names (an id like `ft:gpt-4o:acme-corp:…` contains a company name, so only models in a public registry are reported) **This is structural, not a pinky-swear.** Every value in the payload is a number, a fixed word, or a random ID — there is no free-text field anywhere for content to hide in. The receiver (`deploy/beacon/worker.js`, in this repo so you can read it) drops anything not on an explicit allowlist before storing. ## Turning it off Any one of these: ```bash HEADROOM_BEACON=off # or DO_NOT_TRACK=1 # or # offline mode ``` It's on by default, and Headroom says so at startup: ``` Telemetry: anonymous compression stats — never prompts, code, or file paths. Helps us improve compression | Off: HEADROOM_BEACON=off ``` `HEADROOM_TELEMETRY` is a **separate** switch that still only affects local stats. If you had turned that on, this change does not start uploading anything — you answered a different question, and upgrading should not change the answer. ## Why the percentages, not just "tokens saved" "We saved 36%" hides the interesting part. In the example above only **46% of tokens were eligible** for compression at all — the rest is frozen cache prefix and system prompts we deliberately do not touch. Of what we *could* touch, we removed **78%**. Those are two separate problems. Raising eligibility is proxy work; raising yield is compressor work. A single number cannot tell us which to fix. ## Coverage `emit_request_outcome` is a single chokepoint — `handler.metrics.record_request` is called from exactly one place, inside the funnel — so all 30 `RequestOutcome` construction sites are covered: Anthropic, OpenAI, Gemini, Bedrock, batch, streaming, and the long-lived Codex Responses-WS path. The `headroom_compress` MCP path bypassed that funnel and is now wired in separately. It has a different shape (no provider, no upstream latency, and everything handed to the tool is eligible by construction), so `sources` counts turns by origin — MCP turns always read `eligible_pct: 100` and must not drag the proxy's real eligibility ceiling upward. **Subagents.** All subagent traffic through the proxy merges into one session, which is correct for savings and retention but means `turns` conflates fan-out with depth. Fan-out is still derivable — `compression.latency_ms_total / session.duration_s` gives the concurrency ratio (~1x serial, ~4x for four parallel agents), so no extra field is needed. Verified no lost updates under 6-way concurrency (1,200 turns). **Known gap:** `--workers N` gives each process its own aggregator, so one user session becomes up to N. Token totals and fleet rates stay correct; session counts inflate. This matches the existing documented limitation that TOIN state, CostTracker, and the prefix tracker are all per-process. ## Notes for reviewers - **Cumulative snapshots, not deltas.** Every report restates running totals under one session ID, so the highest `seq` per `(install, session)` is the complete session. Dedupe is a window function, and a lost report costs nothing. - **Never breaks the proxy.** Every path swallows its own exceptions; uploads go out on a daemon thread so nothing blocks the request loop. - **Explicit User-Agent is load-bearing.** urllib's default is blocked by Cloudflare (error 1010). Combined with fire-and-forget error handling, that would have failed every upload while looking perfectly healthy. - **The exit flush was broken and is fixed.** `atexit` handed the POST to a daemon thread, and daemon threads are killed before they finish during interpreter shutdown — so nothing was sent. That silently dropped *every session shorter than the 5-minute heartbeat*, plus all short-lived subagent MCP processes. The exit path now posts synchronously with a 2s timeout. - Receiver and query tooling are in `deploy/beacon/`. ## Testing - `python -m headroom.telemetry.session` self-check: dedupe, cumulative totals, dropped-report recovery, payload contains no model id or prompt-derived string, allowlist coverage - 175 telemetry/outcome tests pass; 6 new ones cover the opt-out notice - Verified end to end against a live deployment: client → receiver → storage → query ## Still to do before release The default endpoint currently points at a temporary `workers.dev` URL. It needs to move to a Headroom-owned hostname before this ships in a tagged release — noted inline at `DEFAULT_ENDPOINT`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 05:43:25 -07:00
}
]
}
}
}
]
}
]
}
]
}