mirror of
https://github.com/headroomlabs-ai/headroom.git
synced 2026-08-10 14:27:00 -04:00
## Description Adds the Phase 0 `agent-evals/` nested project for benchmarking coding-agent task accuracy with and without Headroom's proxy/compression path. The project is intentionally separate from the published `headroom-ai` package and provides the shared A/B framework used by the stacked Phase 1 PR #1040. ## Type of Change - [x] New feature (non-breaking change that adds functionality) - [x] Tests only - [x] Documentation update ## Changes Made - Added a three-arm experiment model for direct provider calls, Headroom passthrough, and Headroom compression. - Added the resumable orchestrator, run manifest/config models, JSON logging, and append-only journal handling. - Added savings capture from Headroom response headers plus scorecard reporting for resolved rate and savings. - Added unit tests and live-test markers for provider-key dependent validation. - Kept the benchmark project isolated from the product package and normal Headroom release wheel. ## Testing - [x] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check`) - [x] Type checking passes (`mypy`) - [x] New tests added for new functionality ### Test Output ```text agent-evals Phase 0 validation from original PR: 74 unit tests passed ruff clean mypy clean CI on this PR: changes and commitlint pass; product CI jobs are skipped because this only changes the nested agent-evals project. ``` ## Real Behavior Proof - Environment: local agent-evals development environment with provider-key dependent live tests skipped unless credentials are present. - Exact command / steps: Ran the Phase 0 unit suite, ruff, and mypy for the nested `agent-evals` project; GitHub CI also ran the repository change detection and commitlint jobs for this PR. - Observed result: The Phase 0 framework tests passed locally, static checks were clean, and GitHub CI reported passing change detection/commitlint for the PR. - Not tested: live provider accuracy claims; those require provider keys and larger benchmark runs and are intentionally covered by live-marked tests and later stacked phases. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review Co-authored-by: JerrettDavis <mxjerrett@gmail.com>
33 lines
882 B
Makefile
33 lines
882 B
Makefile
.PHONY: help install lint fmt typecheck test test-live gate
|
|
|
|
help:
|
|
@echo "make install - pip install -e .[dev,stats]"
|
|
@echo "make lint - ruff check src tests"
|
|
@echo "make fmt - ruff format src tests"
|
|
@echo "make typecheck - mypy src"
|
|
@echo "make test - pytest (unit only; live/real_llm skipped without keys)"
|
|
@echo "make test-live - pytest incl. live tests (requires provider keys)"
|
|
@echo "make gate - lint + typecheck + test (the agent-evals push gate)"
|
|
|
|
install:
|
|
pip install -e ".[dev,stats]"
|
|
|
|
lint:
|
|
ruff check src tests
|
|
|
|
fmt:
|
|
ruff format src tests
|
|
|
|
typecheck:
|
|
mypy src
|
|
|
|
# Unit-only by default: live tests self-skip when keys are absent, but we also exclude
|
|
# them here so a no-key machine never even collects them.
|
|
test:
|
|
pytest -m "not live and not real_llm"
|
|
|
|
test-live:
|
|
pytest
|
|
|
|
gate: lint typecheck test
|
|
@echo "✅ agent-evals gate PASSED"
|