headroom/scripts
JD Davis adb793bee1
ci: harden PR governance and model cache checks (#1401)
## Description

Hardens two routine PR-review pain points from the recent open-PR sweep:

- PR Governance reruns could keep validating the stale
`pull_request_target` event body even after the live PR description had
been fixed.
- Main CI model-cache misses could surface as dozens of unrelated
memory-test failures instead of one clear cache-preflight failure.

This intentionally avoids PyPI/package-bloat and release/nightly
workflow changes so the PR stays scoped to review and CI stabilization.

## Type of Change

- [x] Bug fix
- [ ] New feature
- [ ] Documentation
- [x] Refactor
- [x] Tests only

## Changes Made

- Added `--body-file` support to `scripts/pr-governance.py` so workflows
can validate the current PR body rather than stale rerun payloads.
- Updated PR Governance to fetch the live PR body via the GitHub API
before validating template fields.
- Added a CI preflight script that loads the default
sentence-transformer model in offline mode and verifies the expected
embedding dimension.
- Wired that preflight into the sharded CI job before pytest starts,
turning missing/corrupt Hugging Face caches into one early, actionable
failure.
- Added workflow/script regression tests for the live-body override and
model-cache preflight placement.

## Testing

- [x] Unit tests
- [x] Lint/static checks
- [ ] Integration tests
- [ ] Manual testing

### Test Output

```text
uv run --with pytest --with pytest-asyncio python -m pytest scripts/tests/test_pr_governance.py scripts/tests/test_pr_health_workflow.py scripts/tests/test_ci_workflow.py -q
9 passed in 0.04s

uv run ruff check scripts/pr-governance.py scripts/ci/verify_hf_model_cache.py scripts/tests/test_pr_governance.py scripts/tests/test_pr_health_workflow.py scripts/tests/test_ci_workflow.py
All checks passed!

python -m py_compile scripts\ci\verify_hf_model_cache.py scripts\pr-governance.py
# passed
```

## Real Behavior Proof

- Environment: Windows 11, Python 3.13.3, isolated worktree
`C:\git\headroom\.worktrees\stabilization-hardening`.
- Exact command / steps: Ran the focused governance/workflow tests, ruff
on touched Python files, and `py_compile` for the executable scripts.
- Observed result: Governance tests prove a stale event body can be
overridden by the live PR body; workflow tests prove CI validates live
PR body and runs the Hugging Face offline-cache preflight before pytest
shards.
- Not tested: Full GitHub CI before PR creation; that will run on this
PR. The new Hugging Face preflight itself is intentionally not run
locally because it depends on the CI-warmed offline model cache.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review
2026-06-26 21:34:34 -07:00
..
ci ci: harden PR governance and model cache checks (#1401) 2026-06-26 21:34:34 -07:00
fixtures feat(scripts): add Codex proxy reconnect-storm repro harness 2026-04-20 22:02:02 +07:00
tests ci: harden PR governance and model cache checks (#1401) 2026-06-26 21:34:34 -07:00
audit_wheel_glibc_symbols.py fix(crusher): shim __libc_single_threaded for glibc < 2.32 + extend audit 2026-05-05 13:59:21 -07:00
build_rust_extension.sh refactor: single-wheel maturin build backend (fixes #355) 2026-05-03 13:16:41 -07:00
changelog-gen.py chore: renormalize line endings to LF 2026-04-24 15:33:30 +02:00
eval_output_shaper.py feat: output-token reduction — verbosity shaper, per-user learning, counterfactual savings (#965) 2026-06-16 21:06:43 -07:00
export_kompress_v2_onnx.py feat: switch Kompress default to kompress-v2-base with weight-only int8 ONNX (#799) 2026-06-09 23:28:40 -07:00
install-git-hooks.sh Fix CI lint failure by formatting PR governance scripts (#933) 2026-06-12 17:11:39 -05:00
install.ps1 feat: headroom wrap opencode / unwrap opencode CLI (#1105) 2026-06-22 11:07:12 -05:00
install.sh feat: headroom wrap opencode / unwrap opencode CLI (#1105) 2026-06-22 11:07:12 -05:00
pr-governance.py ci: harden PR governance and model cache checks (#1401) 2026-06-26 21:34:34 -07:00
README.md feat(scripts): add Codex proxy reconnect-storm repro harness 2026-04-20 22:02:02 +07:00
record_fixtures.py feat(rust): scaffold workspace + parity harness (phase-0) 2026-04-24 13:39:48 -07:00
refresh_model_limits.sh fix(rust): wire ICM compressor into Rust proxy on /v1/messages 2026-05-01 16:44:44 -07:00
replay_codex_ws_load.py fix(tests): ship scripts/replay_codex_ws_load.py so CI can import it 2026-05-14 13:44:41 -07:00
repro_codex_replay.py fix: replace asyncio.timeout with 3.10-compat shim in repro harness 2026-04-20 13:41:02 -05:00
smoke_issue_327.py fix(proxy): remove content-keyed TTL walker that conflated content with positional cache (#327) 2026-05-01 12:04:28 -07:00
sync-plugin-versions.py fix(proxy): lazy-import server to avoid fastapi crash (#442) 2026-06-10 12:44:23 -05:00
validate-workflows.sh ci: scope PR workflow runs by changed paths (#1067) 2026-06-16 19:11:45 -07:00
verify-versions.py fix: make proxy upgrades version-aware 2026-05-09 15:58:27 -07:00
version-sync.py fix: make proxy upgrades version-aware 2026-05-09 15:58:27 -07:00

scripts/

Utility scripts bundled with the Headroom repo. Most are one-off operator tools; a few are runnable as part of development workflows.

Reproducing the reconnect storm

repro_codex_replay.py reproduces the multi-agent Codex reconnect/retry storm against a local Headroom proxy (default http://127.0.0.1:8787), as described in wiki/plans/2026-04-17-codex-proxy-runtime-analysis.md under "Latest Correction". Use it to:

  • Regression-check that /livez stays responsive under a cold-start storm.
  • Empirically tune the Unit 4 pre-upstream semaphore default (HEADROOM_ANTHROPIC_PRE_UPSTREAM_CONCURRENCY).
  • Exercise the Codex WS lifecycle + Anthropic HTTP path simultaneously without needing to replay captured production traffic.

Run

# Default: 8 WS + 4 HTTP clients, 30s storm, p99 /livez must stay <= 500ms.
python scripts/repro_codex_replay.py

# Tighter budget, shorter run:
python scripts/repro_codex_replay.py \
    --url http://127.0.0.1:8787 \
    --ws-clients 16 \
    --anthropic-clients 8 \
    --duration 60 \
    --livez-threshold-ms 100

# Dump the full summary as JSON for downstream tooling:
python scripts/repro_codex_replay.py --json

Exit code:

  • 0 — warmup succeeded (or was skipped), storm ran for the requested duration, and /livez p99 stayed under --livez-threshold-ms.
  • 1 — soft assertion failed, proxy unreachable, or unhandled exception. Proxy-unreachable is detected and reported within ~5 seconds.

Fixtures

The script loads two hand-crafted, fully synthetic JSON fixtures:

  • scripts/fixtures/anthropic_replay_body.json — shape of a large agent reconnect replay /v1/messages?beta=true POST body.
  • scripts/fixtures/codex_response_create_frame.json — first Codex WS frame with the {"type": "response.create", "response": {...}} envelope.

Override via --ws-frame-fixture / --anthropic-body-fixture if you have captured traffic to replay instead.

Interpretation

  • /livez p99 under threshold means the event loop is not starved during the storm. If it rises with the semaphore unbounded (HEADROOM_ANTHROPIC_PRE_UPSTREAM_CONCURRENCY=10000) and drops back under the default, Unit 4's backpressure is working.
  • Codex WS: opened should equal --ws-clients. response.completed typically stays low when upstream auth isn't configured locally — the goal is handshake + relay wiring, not real upstream traffic.
  • Anthropic HTTP: ok_2xx + non_2xx + timed_out + errors should roughly equal attempted. Sustained non-zero timed_out during the storm is the failure signal the plan targets.

A smoke test at tests/test_scripts/test_repro_codex_replay_smoke.py exercises the script against a mock FastAPI server on every PR.

Install scripts

  • install.sh — POSIX installer.
  • install.ps1 — Windows PowerShell installer.

These are generated by the release pipeline; edit with care.