headroom/wiki/learn.md
Tejas Chopra bd76235f5c
fix(cli): harden all CLI surfaces + fix docs accuracy (#1491)
## Summary

Full CLI audit + documentation accuracy pass. All 5 commits on this
branch:

### CLI Hardening (4 commits)
- **Clean errors instead of tracebacks**: corrupt manifests, missing
Docker, malformed JSONL, bad `--profile`, invalid env-var values all now
raise `click.ClickException` with helpful messages
- **Range validation**: ~25 numeric flags across 10 files now use
`click.IntRange`/`FloatRange` — `--port 0`, `--hours -1`, `--limit 0`
etc. produce clean usage errors instead of silent wrong behavior
- **Flag combination warnings**: conflicting combos (`--no-rate-limit` +
`--rpm`, `--no-optimize` + `--target-ratio`, `--telemetry` +
`--no-telemetry`) emit yellow warnings on stderr
- **`memory --db-path` default fixed**: was resolving to
`headroom_memory.db` (wrong bare file); now uses project store
`./.headroom/memory.db` if present, else `~/.headroom/memory.db`
- **`memory list --search` + filters**: `--scope`/`--session`/`--since`
were silently ignored when `--search` was also set; now filters are
applied to search results
- **`learn --verbosity --apply` now works**: the output shaper is off by
default (`HEADROOM_OUTPUT_SHAPER`); `--apply` now hot-enables it via
`POST /admin/runtime-env` on a running proxy, or prints explicit `export
HEADROOM_OUTPUT_SHAPER=1` instructions when no proxy is running
- **`perf --hours` overflow**: `1e9` hours no longer raises
`OverflowError`; treated as "all data"
- **`evals memory --categories` invalid input**: `abc,1,2` now raises
`BadParameter` instead of a raw `ValueError` traceback

### Documentation (1 commit, 20 files)

Corrected factual errors found by 3 parallel audit agents across root
docs, wiki, and the published Fumadocs site:

**Critical (caused runtime errors or wrong behavior if followed):**
- `simulation.mdx`: `plan.transforms_applied` -> `plan.transforms`;
`plan.savings_percent` -> computed from available fields (both raised
`AttributeError`)
- `shared-context.mdx`: `import { SharedContext } from "headroom"` ->
`"headroom-ai"` (5x `ImportError`)
- `claude-code-azure-foundry.mdx`: `pip install headroom` -> `pip
install headroom-ai`
- `api-reference.mdx` + `configuration.mdx`: `from headroom import
GoogleProvider` -> `from headroom.providers import GoogleProvider`
- `ccr.mdx`: CCR TTL default 300s -> 1800s (30 min)

**Fabricated flags removed:**
- `wiki/proxy.md` + `wiki/cli.md`: `--no-intelligent-context`,
`--no-intelligent-scoring`, `--no-compress-first` (none exist); replaced
with real CCR flags
- `wiki/configuration.md`: `--no-ccr-responses`, `--no-ccr-expansion`
(none exist); replaced with real flags
- `wiki/troubleshooting.md`, `wiki/metrics.md`,
`docs/troubleshooting.mdx`: `headroom proxy --log-level debug` (flag
doesn't exist)

**Stale content corrected:**
- `llms.txt`: telemetry stated as enabled-by-default (it's opt-in); wrap
list had 5 tools (now 11)
- `README.md`: compatibility matrix added 5 missing `wrap` targets;
`unwrap`, `doctor`, `init`/`install`, savings-analytics now mentioned
- `SECURITY.md`: supported version table showed 0.2.x (current: 0.27.x)
- `wiki/learn.md`: 5 missing flags added; verbosity shaper-off behavior
documented
- `wiki/quickstart.md`: "Configuration Reference" linked to `api.md`
(wrong) -> `configuration.md`
- `CacheAlignerConfig.enabled` default corrected: `True` -> `False`
- `opencode.mdx`: `--port` default wrong ("random") -> 8787; `openai`
backend removed
- `CONTRIBUTING.md`: broken Markdown table cell fixed
- `docs/meta.json`: `claude-code-azure-foundry` added to nav (was
unreachable orphan page)
- `configuration.mdx`: SDK modes vs proxy `--mode` now clearly
distinguished

## Test plan

- [x] `python -m pytest tests/ -x -q` — 857 passed, 0 failures
- [x] 41-combination CLI smoke test (all flag combos across 8 commands)
— 0 tracebacks
- [x] `ruff check` on all modified Python files — clean
- [x] Docs changes are removals/corrections of fabricated or stale
content; no new claims introduced
2026-06-27 14:48:43 -07:00

8.8 KiB

Headroom Learn

Offline failure learning for coding agents. Analyzes past conversations, finds what went wrong, correlates it with what eventually worked, and writes specific project-level learnings that prevent the same mistakes next session.

Quick Start

# See recommendations for current project (dry-run, no changes)
headroom learn

# Write recommendations to CLAUDE.local.md (gitignored, personal default)
headroom learn --apply

# Write to the shared team file instead
headroom learn --apply --target CLAUDE.md

# Analyze a specific project
headroom learn --project ~/my-project --apply

# Analyze all projects
headroom learn --all --apply

How It Works

Past Sessions → Plugin → Analyzer → Writer → Agent-native context file
                  │           │          │
                  │           │          └─ Writes marker-delimited sections
                  │           │             (replaced on re-run, not duplicated)
                  │           │
                  │           └─ LLM-based analysis: finds failure patterns,
                  │              success correlations, and actionable rules
                  │
                  └─ Plugin reads agent-specific logs:
                     • Claude Code: ~/.claude/projects/*.jsonl
                     • Codex:       ~/.codex/sessions/*.json
                     • Gemini CLI:  ~/.gemini/tmp/*/chats/session-*.json

Success Correlation

The core innovation. Instead of cataloging failures ("Read failed 5 times"), Headroom finds what the model did to fix each failure:

  • Failed: Read axion-formats/src/main/java/.../FirstClassEntity.java
  • Then succeeded: Read axion-scala-common/src/main/scala/.../FirstClassEntity.scala
  • Learning: "FirstClassEntity is at axion-scala-common/, not axion-formats/"

This produces specific, actionable corrections — not generic advice.

What It Learns

1. Environment Facts → CLAUDE.md

Which runtime commands work vs fail.

### Environment
- **Python**: use `uv run python` (not `python3` — modules not available outside venv)

2. File Path Corrections → CLAUDE.md

Wrong paths the model keeps guessing, with the correct locations.

### File Path Corrections
- `axion-common/src/.../AxionSparkConstants.scala`
  → actually at `axion-spark-common/src/.../AxionSparkConstants.scala`

3. Search Scope → CLAUDE.md

Which directories to search in (narrow paths fail, broader ones work).

### Search Scope
- Don't search `axion-model/` → use `axion/` (the repo root)

4. Command Patterns → CLAUDE.md

How commands should (and shouldn't) be run.

### Command Patterns
- **user_prefers_manual**: User rejected gradle 18 times — show the command, don't execute
- **python_runtime**: Use `uv run python` not `python3` (ModuleNotFoundError)

5. Known Large Files → CLAUDE.md

Files that need offset/limit with Read.

### Known Large Files
- `proxy/server.py` (~8000 lines) — always use offset/limit

6. Retry Prevention → MEMORY.md

Specific suggestions derived from actual corrections.

7. Permission Notes → MEMORY.md

Commands repeatedly rejected — model should suggest them to the user instead.

Where Learnings Go

Pattern Claude Code Codex Gemini CLI
Environment, paths, commands CLAUDE.local.md (default) or CLAUDE.md (with --target CLAUDE.md) AGENTS.md GEMINI.md
Retry patterns, permissions MEMORY.md instructions.md GEMINI.md

Output files are agent-native: Claude Code writes to CLAUDE.local.md by default (gitignored, personal); pass --target CLAUDE.md for the shared team file. Codex uses AGENTS.md, Gemini uses GEMINI.md. The same learnings, written to the format each agent reads.

Marker-Based Updates

Headroom manages a clearly-delimited section in each file:

<!-- headroom:learn:start -->
## Headroom Learned Patterns
*Auto-generated by `headroom learn` — do not edit manually*
...
<!-- headroom:learn:end -->

On re-run, only the content between markers is replaced. Your existing file content is preserved.

Architecture (Plugin System)

Headroom Learn uses a plugin architecture where each agent is a self-contained plugin:

Plugin Registry (auto-discovered)
├── ClaudeCodePlugin  →  Analyzer (LLM)  →  ClaudeCodeWriter  →  CLAUDE.md / MEMORY.md
├── CodexPlugin       →  Analyzer (LLM)  →  CodexWriter       →  AGENTS.md / instructions.md
├── GeminiPlugin      →  Analyzer (LLM)  →  GeminiWriter      →  GEMINI.md
└── (your plugin)     →  Analyzer (LLM)  →  (your writer)     →  (your file)

Plugins bundle scanning, detection, and writing for one agent. Built-in plugins are auto-discovered from headroom.learn.plugins.*. External plugins register via the headroom.learn_plugin entry point.

The Analyzer is shared — it uses an LLM (Sonnet, GPT-4o, or Gemini Flash) to find patterns. Same analysis for any agent.

Adding Support for a New Agent

  1. Create headroom/learn/plugins/myagent.py
  2. Implement LearnPlugin + ConversationScanner (scanner + writer + detection)
  3. Add plugin = MyAgentPlugin() at module scope
  4. Done — headroom learn --agent myagent works automatically

Or install an external plugin: pip install headroom-learn-cursor (registers via entry point).

CLI Reference

headroom learn [OPTIONS]

Options:
  --project PATH               Project directory (default: current directory)
  --all                        Analyze all discovered projects (mutually exclusive with --project)
  --apply                      Write recommendations (default: dry-run)
  --target TEXT                Context file to write (default: CLAUDE.local.md for Claude Code)
  --main-only                  Write only to the main context file, skip MEMORY.md
  --agent [auto|claude|codex|gemini]
                               Which agent to analyze (default: auto-detect)
  --model TEXT                 LLM for analysis (default: auto from API keys or CLI)
  --workers / -j INTEGER       Parallel analysis workers (min 1, default: auto)
  --verbosity                  Analyze verbosity level instead of failure patterns
  --llm-judge                  Use an LLM to score verbosity quality (requires --verbosity)

Verbosity learning (--verbosity)

headroom learn --verbosity analyzes past sessions to infer the ideal output verbosity level for your project and writes a verbosity.json profile.

Important: the output shaper is off by default. Running --verbosity --apply will either:

  • Hot-enable the output shaper on a running proxy (POST /admin/runtime-env), OR
  • Print instructions to set HEADROOM_OUTPUT_SHAPER=1 before headroom wrap ...

To keep the shaper on across proxy restarts, add export HEADROOM_OUTPUT_SHAPER=1 to your shell profile before starting the proxy.

Flag interactions:

  • --all and --project are mutually exclusive
  • --llm-judge requires --verbosity
  • --verbosity --all --apply is rejected (verbosity persists a single global level)

Supported Agents

Agent Scanner Writer Output Files
Claude Code Reads ~/.claude/projects/*.jsonl ClaudeCodeWriter CLAUDE.md, MEMORY.md
OpenAI Codex Reads ~/.codex/sessions/*.json CodexWriter AGENTS.md, instructions.md
Gemini CLI Reads ~/.gemini/tmp/*/chats/session-*.json GeminiWriter GEMINI.md

LLM Backend Selection

headroom learn needs an LLM to analyze your sessions. It picks one automatically using this priority:

Priority Source Example
1 --model flag headroom learn --model gpt-4o
2 API key env var ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY
3 HEADROOM_LEARN_CLI env var export HEADROOM_LEARN_CLI=gemini
4 Auto-detect installed CLIs Checks PATH for claude, gemini, codex

Using without an API key

If you use Claude Code, Gemini CLI, or Codex via subscription (no raw API key), headroom learn can call them directly:

# Auto-detects claude in PATH — no API key needed
headroom learn

# Explicitly select a CLI backend
headroom learn --model gemini-cli

# Pin a CLI via environment variable
export HEADROOM_LEARN_CLI=codex
headroom learn

Valid values for HEADROOM_LEARN_CLI: claude, gemini, codex.

Real-World Results

Tested on 67,583 tool calls across 23 projects:

Metric Value
Failure rate 7.5% (5,066 failures)
Corrections extracted 164 per project (avg)
Specific path corrections 22 (axion project)
Search scope corrections 24 (axion project)
Command patterns learned 5 (axion project)
Estimated preventable waste ~27 MB across corpus