headroom/.github/workflows/ci.yml
Ashish a14ab45cf0
fix(proxy): make budget enforcement actually work (#885)
## Description

`CostTracker._costs` was initialized but never written to, so
`get_period_cost()` always returned `0` and `check_budget()` always
returned "allowed" — the `--budget` flag was a silent no-op.
`_prune_old_costs()` was dead code with zero callers. This makes budget
enforcement actually work: requests are rejected once the configured
limit is reached.

Closes # <!-- no tracked issue; discovered during a proxy-pipeline audit
-->

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- **`headroom/proxy/cost.py`** — `record_tokens()` now computes the
request cost via `estimate_cost()` and appends it to `_costs`,
activating `_prune_old_costs()`. When a call site has no API usage
breakdown (cache/uncached all zero), `tokens_sent` is used as the input
count so input cost is not silently dropped. `COST_RETENTION_HOURS` 24 →
744 so retention covers the longest budget period (monthly sums from the
1st; 24h retention would have under-enforced monthly budgets).
- **`headroom/proxy/outcome.py`** — the request funnel passes
`output_tokens` through to `record_tokens()` so costs include output,
for all providers.
- **`headroom/cli/proxy.py`** — added `--budget-period
[hourly|daily|monthly]` (env `HEADROOM_BUDGET_PERIOD`); it existed in
`ProxyConfig` and the server entry point but was unreachable from the
main CLI. Fixed the `--budget` help text that wrongly said "resets at
midnight UTC".
- **`headroom/cli/main.py`** — minor registration/version plumbing.
- Tests: regression coverage for the full `record_tokens →
get_period_cost → check_budget` chain, the `tokens_sent` fallback, and
the `--budget-period` flag/env wiring.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ pytest tests/test_cost_tracker_counterfactual.py tests/test_request_outcome.py -q
40 passed

$ ruff check headroom/proxy/cost.py headroom/proxy/outcome.py headroom/cli/proxy.py
All checks passed!

$ mypy headroom/proxy/cost.py headroom/proxy/outcome.py headroom/cli/proxy.py --ignore-missing-imports
Success: no issues found
```

## Real Behavior Proof

- Environment: local macOS, Python 3.11, branch `fix/budget-enforcement`
at the PR head commit.
- Exact command / steps: `pytest
tests/test_cost_tracker_counterfactual.py::test_budget_enforced_after_recording_costs
-v` — sets `CostTracker(budget_limit_usd=0.0001)`, records ~$1.50 of
Sonnet input, then asserts `check_budget()` returns not-allowed with
`remaining == 0`.
- Observed result: budget is now enforced — `get_period_cost()` reflects
real spend and `check_budget()` rejects once the limit is exceeded (the
proxy returns HTTP 429 on that path). On `main` the same test fails
because `_costs` is never populated and `check_budget()` always returns
allowed.
- Not tested: live end-to-end rejection against a running proxy with
real upstream traffic; the running proxy needs a restart on this version
to pick up the fix.

```text
$ pytest tests/test_cost_tracker_counterfactual.py::test_budget_enforced_after_recording_costs \
         tests/test_cost_tracker_counterfactual.py::test_budget_input_cost_counted_without_usage_breakdown -v
2 passed
```

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable

## Screenshots (if applicable)

N/A — CLI/backend change with no UI surface. See **Test Output** and
**Real Behavior Proof** above for terminal evidence.

## Additional Notes

- The `ci.yml` coverage-upload change originally added here (commit
`120696e5`) was superseded by an equivalent block the maintainer added
to `main`; the merge from main resolved to main's version. Codecov now
reports all modified lines covered.
- N/A checklist items: no docs or CHANGELOG entry — this is an internal
correctness fix to an existing flag.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-06-15 10:22:27 -05:00

502 lines
18 KiB
YAML

name: CI
# Intelligent + parallel pipeline (cutover from the old 4-version matrix):
# changes — paths-filter; skips heavy work for docs-only changes
# build-wheel — compile the Rust ext ONCE (fast `ci` cargo profile), share via artifact
# lint — ruff + mypy, once
# prefetch-model — download the embedding model ONCE (authenticated), warm shared cache
# test — 4 parallel shards (pytest-split), each a fresh runner VM; run offline
# test-extras / test-agno / build / commitlint / workflow-validation / *-e2e — preserved
#
# Notes: CPU-only torch everywhere (no CUDA stack); test shards run HF_HUB_OFFLINE.
# Multi-version (3.10/3.11/3.13) coverage on main is a planned follow-up.
on:
push:
branches: [main]
pull_request:
branches: [main]
workflow_dispatch:
permissions:
contents: read
concurrency:
group: ci-${{ github.workflow }}-${{ github.ref }}
# Cancel superseded runs on PRs/branches, but never cancel a main build.
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
env:
PY_VERSION: "3.12"
# CPU-only torch — runners have no GPU; the default CUDA wheels pull ~2.5 GB.
PIP_EXTRA_INDEX_URL: https://download.pytorch.org/whl/cpu
jobs:
changes:
runs-on: ubuntu-latest
timeout-minutes: 5
outputs:
code: ${{ steps.filter.outputs.code }}
e2e: ${{ steps.filter.outputs.e2e }}
workflows: ${{ steps.filter.outputs.workflows }}
steps:
- uses: actions/checkout@v6
- uses: dorny/paths-filter@v4
id: filter
with:
filters: |
code:
- 'headroom/**'
- 'crates/**'
- '**/*.rs'
- 'pyproject.toml'
- 'Cargo.toml'
- 'Cargo.lock'
- 'tests/**'
- 'scripts/**'
- '.github/workflows/ci.yml'
e2e:
- 'headroom/**'
- 'crates/**'
- 'docker/**'
- 'Dockerfile'
- 'e2e/**'
- 'scripts/install*'
- 'pyproject.toml'
workflows:
- '.github/workflows/**'
lint:
needs: changes
if: needs.changes.outputs.code == 'true'
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: ${{ env.PY_VERSION }}
- name: Cache pip
uses: actions/cache@v5
with:
path: ~/.cache/pip
key: ${{ runner.os }}-pip-lint-${{ hashFiles('pyproject.toml') }}
restore-keys: ${{ runner.os }}-pip-lint-
- run: python -m pip install --upgrade pip ruff mypy
- name: ruff check
run: ruff check .
- name: ruff format --check
run: ruff format --check .
- name: mypy
run: mypy headroom --ignore-missing-imports
build-wheel:
needs: changes
if: needs.changes.outputs.code == 'true'
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: ${{ env.PY_VERSION }}
- uses: dtolnay/rust-toolchain@1.96.0
- uses: Swatinem/rust-cache@v2
with:
workspaces: ". -> target"
- name: Build wheel once (fast CI cargo profile)
run: |
python -m pip install --upgrade pip maturin
maturin build --profile ci --out dist --interpreter "python${PY_VERSION}"
- uses: actions/upload-artifact@v7
with:
name: headroom-wheel
path: dist/*.whl
retention-days: 1
prefetch-model:
needs: changes
if: needs.changes.outputs.code == 'true'
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/setup-python@v6
with:
python-version: ${{ env.PY_VERSION }}
- name: Cache HuggingFace model
id: hfcache
uses: actions/cache@v5
with:
path: ~/.cache/huggingface
key: ${{ runner.os }}-models-allMiniLM-v2
- name: Fetch all-MiniLM-L6-v2 once (authenticated, resilient)
if: steps.hfcache.outputs.cache-hit != 'true'
env:
HF_TOKEN: ${{ secrets.HF_TOKEN }}
HF_HUB_DISABLE_TELEMETRY: "1"
run: |
python -m pip install --upgrade pip huggingface_hub
for i in 1 2 3 4 5 6; do
if python -c "from huggingface_hub import snapshot_download; snapshot_download('sentence-transformers/all-MiniLM-L6-v2')"; then exit 0; fi
echo "::warning::model fetch attempt $i failed; backing off"; sleep $((i * 30))
done
echo "::error::could not fetch all-MiniLM-L6-v2 from HuggingFace"; exit 1
test:
needs: [changes, build-wheel, prefetch-model]
if: needs.changes.outputs.code == 'true'
runs-on: ubuntu-latest
timeout-minutes: 30
strategy:
fail-fast: false
matrix:
shard: [1, 2, 3, 4]
env:
HF_HUB_OFFLINE: "1"
TRANSFORMERS_OFFLINE: "1"
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: ${{ env.PY_VERSION }}
- name: Cache pip
uses: actions/cache@v5
with:
path: ~/.cache/pip
key: ${{ runner.os }}-pip-${{ env.PY_VERSION }}-${{ hashFiles('pyproject.toml') }}
restore-keys: ${{ runner.os }}-pip-${{ env.PY_VERSION }}-
- name: Restore HuggingFace model cache (warmed by prefetch-model)
uses: actions/cache@v5
with:
path: ~/.cache/huggingface
key: ${{ runner.os }}-models-allMiniLM-v2
- name: Download prebuilt wheel
uses: actions/download-artifact@v8
with:
name: headroom-wheel
path: dist
- name: Install (CPU torch + prebuilt wheel + dev deps, no cargo rebuild)
run: |
python -m pip install --upgrade pip
pip install torch --index-url https://download.pytorch.org/whl/cpu
WHEEL="$(ls dist/*.whl)"
pip install "${WHEEL}[dev]" pytest-split
# cwd's ./headroom source tree shadows the installed wheel; copy the
# compiled extension in so tests import it (no second cargo build).
SITE="$(python -c 'import sysconfig; print(sysconfig.get_path("platlib"))')"
cp "${SITE}/headroom/"_core*.so headroom/
python -c "from headroom._core import DiffCompressor; print('headroom._core OK')"
# Coverage upload: without this, codecov only receives reports from
# the two native-e2e workflows (3 CLI test files total), so head
# coverage reads ~6% and codecov/patch fails for ANY diff not
# exercised by those files — a false negative on every PR. The main
# suite runs here; its coverage must be what codecov sees.
- name: Run test shard ${{ matrix.shard }}/4
run: |
pytest tests scripts/tests \
--splits 4 --group ${{ matrix.shard }} \
--cov=headroom --cov-branch \
--cov-report=xml:coverage-${{ matrix.shard }}.xml \
--cov-report= \
--tb=short -q
- name: Upload coverage shard ${{ matrix.shard }} to Codecov
uses: codecov/codecov-action@v5
with:
files: coverage-${{ matrix.shard }}.xml
flags: python
name: python-shard-${{ matrix.shard }}
# Token is sent so uploads authenticate once the repo is activated on
# Codecov. Until then Codecov may 404 ("Repository not found"); either
# way, coverage upload is reporting-only and must never fail a build
# whose tests pass — so this stays non-blocking.
token: ${{ secrets.CODECOV_TOKEN }}
fail_ci_if_error: false
test-extras:
needs: [changes, build-wheel]
if: needs.changes.outputs.code == 'true'
runs-on: ubuntu-latest
timeout-minutes: 30
env:
FASTEMBED_CACHE_PATH: ${{ github.workspace }}/.fastembed-cache
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: ${{ env.PY_VERSION }}
- name: Cache pip
uses: actions/cache@v5
with:
path: ~/.cache/pip
key: ${{ runner.os }}-pip-extras-${{ hashFiles('pyproject.toml') }}
restore-keys: ${{ runner.os }}-pip-extras-
- name: Cache fastembed model
uses: actions/cache@v5
with:
path: ${{ github.workspace }}/.fastembed-cache
key: ${{ runner.os }}-fastembed-bge-small-v1
- name: Download prebuilt wheel
uses: actions/download-artifact@v8
with:
name: headroom-wheel
path: dist
- name: Install (CPU torch + wheel[dev,relevance])
run: |
python -m pip install --upgrade pip
pip install torch --index-url https://download.pytorch.org/whl/cpu
WHEEL="$(ls dist/*.whl)"
pip install "${WHEEL}[dev,relevance]"
SITE="$(python -c 'import sysconfig; print(sysconfig.get_path("platlib"))')"
cp "${SITE}/headroom/"_core*.so headroom/
python -c "from headroom._core import SmartCrusher; print('headroom._core OK')"
- name: Pre-fetch fastembed model (authenticated, resilient)
env:
HF_TOKEN: ${{ secrets.HF_TOKEN }}
HF_HUB_DISABLE_TELEMETRY: "1"
run: |
for i in 1 2 3 4 5; do
if python -c "from fastembed import TextEmbedding; TextEmbedding('BAAI/bge-small-en-v1.5')"; then exit 0; fi
echo "::warning::fastembed fetch attempt $i failed; backing off"; sleep $((i * 20))
done
echo "::error::could not fetch fastembed model from HuggingFace"; exit 1
- name: Run relevance tests
# Offline so fastembed reads the cache the prefetch step just warmed,
# without an unauthenticated cache-validation HEAD that could 429.
env:
HF_HUB_OFFLINE: "1"
TRANSFORMERS_OFFLINE: "1"
run: pytest tests/test_relevance.py -v
test-agno:
needs: [changes, build-wheel]
if: needs.changes.outputs.code == 'true'
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: ${{ env.PY_VERSION }}
- name: Download prebuilt wheel
uses: actions/download-artifact@v8
with:
name: headroom-wheel
path: dist
- name: Install (CPU torch + wheel[dev,agno])
run: |
python -m pip install --upgrade pip
pip install torch --index-url https://download.pytorch.org/whl/cpu
WHEEL="$(ls dist/*.whl)"
pip install "${WHEEL}[dev,agno]"
SITE="$(python -c 'import sysconfig; print(sysconfig.get_path("platlib"))')"
cp "${SITE}/headroom/"_core*.so headroom/
- name: Run agno tests
run: pytest tests/test_integrations/agno/ -v
test-dashboard-ui:
needs: [changes, build-wheel]
if: needs.changes.outputs.code == 'true'
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: ${{ env.PY_VERSION }}
- name: Download prebuilt wheel
uses: actions/download-artifact@v8
with:
name: headroom-wheel
path: dist
- name: Install (CPU torch + wheel[dev] + playwright)
run: |
python -m pip install --upgrade pip
pip install torch --index-url https://download.pytorch.org/whl/cpu
WHEEL="$(ls dist/*.whl)"
pip install "${WHEEL}[dev]" playwright
SITE="$(python -c 'import sysconfig; print(sysconfig.get_path("platlib"))')"
cp "${SITE}/headroom/"_core*.so headroom/
- name: Install chromium
run: playwright install --with-deps chromium
- name: Run dashboard playwright tests
# Stub-based dashboard tests only (routes fully mocked, no network).
# tests/test_dashboard/test_live_feed.py needs a live proxy on
# localhost:8787 and stays excluded; the main shards keep skipping
# these via importorskip since playwright is not installed there.
env:
HEADROOM_PLAYWRIGHT_ARTIFACT_DIR: ${{ runner.temp }}/playwright-artifacts
run: pytest tests/test_dashboard_*_playwright.py -v
- name: Upload dashboard screenshots
if: always()
uses: actions/upload-artifact@v7
with:
name: dashboard-playwright-artifacts
path: ${{ runner.temp }}/playwright-artifacts
if-no-files-found: ignore
retention-days: 7
commitlint:
if: github.event_name == 'pull_request'
runs-on: ubuntu-latest
timeout-minutes: 5
steps:
- uses: actions/checkout@v6
with:
fetch-depth: 0
- uses: wagoid/commitlint-github-action@v6
with:
configFile: .commitlintrc.json
build:
needs: changes
if: needs.changes.outputs.code == 'true'
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: "3.11"
- name: Cache pip
uses: actions/cache@v5
with:
path: ~/.cache/pip
key: ${{ runner.os }}-pip-build-${{ hashFiles('pyproject.toml') }}
restore-keys: ${{ runner.os }}-pip-build-
- uses: dtolnay/rust-toolchain@1.96.0
- uses: Swatinem/rust-cache@v2
with:
workspaces: ". -> target"
# Smoke check that the SHIPPED build (release profile) + sdist are wired
# right; release.yml's matrix is what actually publishes to PyPI.
- name: Install build tools
run: |
python -m pip install --upgrade pip
pip install 'maturin>=1.5,<2.0' twine
- name: Build wheel + sdist
run: |
maturin sdist --out dist
maturin build --release --out dist
- name: Check package
run: twine check dist/*
workflow-validation:
needs: changes
if: needs.changes.outputs.workflows == 'true'
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v6
- name: Cache actionlint + act
id: tools-cache
uses: actions/cache@v5
with:
path: |
/usr/local/bin/actionlint
/usr/local/bin/act
# Key off the workflow file itself: when someone updates the
# download URLs to a newer tool version, the hash changes and
# the cache busts automatically.
key: ${{ runner.os }}-ci-tools-${{ hashFiles('.github/workflows/ci.yml') }}
- name: Install actionlint
if: steps.tools-cache.outputs.cache-hit != 'true'
run: |
curl -fsSL https://raw.githubusercontent.com/rhysd/actionlint/main/scripts/download-actionlint.bash | bash
sudo mv ./actionlint /usr/local/bin/actionlint
- name: Install act
if: steps.tools-cache.outputs.cache-hit != 'true'
run: |
curl -fsSL https://raw.githubusercontent.com/nektos/act/master/install.sh | sudo bash
sudo install ./bin/act /usr/local/bin/act
- name: Validate workflow files
run: bash scripts/validate-workflows.sh
docker-native-e2e:
needs: changes
if: needs.changes.outputs.e2e == 'true'
runs-on: ubuntu-latest
timeout-minutes: 45
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: "3.11"
- name: Build local Headroom image
run: docker build -t headroom-native-e2e:latest .
- name: Run Docker-native installer e2e
env:
HEADROOM_DOCKER_IMAGE: headroom-native-e2e:latest
run: bash e2e/docker-native-install.sh
- name: Run Docker-native compose smoke test
env:
HEADROOM_IMAGE: headroom-native-e2e:latest
HEADROOM_HOST_HOME: ${{ github.workspace }}
HEADROOM_WORKSPACE: ${{ github.workspace }}
run: |
mkdir -p .headroom .claude .codex .gemini
trap 'docker compose -f docker/docker-compose.native.yml down -v' EXIT
docker compose -f docker/docker-compose.native.yml up -d proxy
for attempt in $(seq 1 30); do
if curl --fail --silent http://127.0.0.1:8787/readyz >/dev/null; then
break
fi
if [ "$attempt" -eq 30 ]; then
docker compose -f docker/docker-compose.native.yml logs proxy
exit 1
fi
sleep 1
done
- name: Run Docker-native wrap e2e
run: |
docker build -f e2e/wrap/Dockerfile -t headroom-wrap-e2e .
docker run --rm headroom-wrap-e2e
- name: Run Docker-native init e2e
run: |
docker build -f e2e/init/Dockerfile -t headroom-init-e2e .
docker run --rm headroom-init-e2e
windows-native-wrapper:
needs: changes
if: needs.changes.outputs.e2e == 'true'
runs-on: windows-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: "3.12"
- name: Install test dependencies
run: |
python -m pip install --upgrade pip
pip install pytest
- name: Run native installer wrapper tests
run: pytest tests/test_install/test_native_installers.py -q
macos-native-wrapper:
needs: changes
if: needs.changes.outputs.e2e == 'true'
runs-on: macos-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: "3.11"
- name: Install bash and test dependencies
run: |
brew install bash
python -m pip install --upgrade pip
python -m pip install --retries 10 --timeout 60 pytest
- name: Run native installer wrapper tests
run: |
BASH_PREFIX="$(brew --prefix bash)"
export PATH="$BASH_PREFIX/bin:$PATH"
pytest tests/test_install/test_native_installers.py -q