## Description Repo hygiene for a public OSS project: removes committed `node_modules`, stray/internal/draft markdown, and commercial-surface references — keeping every real doc (the published docs site, the wiki guides, and all component READMEs) intact. Every file was content-audited before removal, and load-bearing files were verified against the code/CI and kept. Net: **1,695 files changed, +23 / −266,409** (the deletions are dominated by a committed `node_modules` tree). Closes # (no tracking issue) ## Type of Change - [ ] Bug fix (non-breaking change that fixes an issue) - [ ] New feature (non-breaking change that adds functionality) - [ ] Breaking change (fix or feature that would cause existing functionality to change) - [x] Documentation update - [x] Code refactoring (no functional changes) ## Changes Made **Removed (verified to have no code/CI dependencies):** - `examples/vercel-ai-sdk-pr/` — 1,649 committed `node_modules` files (zero example source); `node_modules/` added to `.gitignore`. - `docs/spec/` (23 draft "Living Specification" files — orphaned, `1.0.0-draft`, drifted from the code), `docs/superpowers/` (2 agent plans), `docs/proposals/` (2 internal/commercial memos). - 6 orphan `docs/*.md` (auth-modes, bedrock, claude-code-vertex-headroom, cortex-code, output-token-reduction-guide, rtk-loop-weighting). - `PR.md` (committed PR draft), `ENTERPRISE.md`, `.github/FUNDING.yml`. **Content scrubs:** - Removed unreleased "Headroom Cloud" / `api.headroom.ai` / `hr_` references from `configuration.mdx`, `wiki/configuration.md`, `wiki/typescript-sdk.md`, `sdk/typescript/README.md` (reworded to neutral, accurate phrasing). - Dropped a stale "awaiting maintainer before merge" line from `plugins/headroom-oauth2/SPEC.md`; tidied `.gitignore` comments (kept the protective `headroom-managed/` ignore rule). - Fixed the now-dangling links into removed files (README nav/`output-token-reduction` link, `scripts/README`, `wiki/vertex`). **Explicitly KEPT (load-bearing — would orphan in-code citations if removed):** - `.changelog.md` — consumed by `.github/workflows/release.yml` (read as the release-notes file). - `REALIGNMENT/`, `docs/observability.md`, `docs/rtk-architecture.md`, `wiki/plans/`, `TESTING-copilot-subscription.md` — referenced by the Rust core / Python / tests as design docs. ## Testing - [ ] Unit tests pass (`pytest`) - [x] Linting passes (`ruff check`) - [ ] Type checking passes (`mypy`) - [x] New tests added for new functionality - [x] Manual testing performed ### Test Output ```text # Docs/markdown + .gitignore only — no Python/Rust source changed, so the # behavioral test suite is unaffected. Verified the cleanup did not orphan # references or break the published docs site: $ git ls-files 'docs/content/docs/*.mdx' | wc -l # published site intact 42 $ # meta.json nav unchanged; no published page removed. $ grep -rnI "Headroom Cloud|api.headroom.ai|'hr_" $(git ls-files '*.md' '*.mdx') >>> none $ # dangling refs to removed files (excl pre-existing P0/P2 spec stubs that $ # never existed in git): none remaining. ``` ## Real Behavior Proof - Environment: macOS, local git clone of the repo (markdown/.gitignore changes only — no runtime). - Exact command / steps: 4 read-only content-audit agents classified every `.md`/`.mdx` file; each removal candidate was cross-checked against the codebase (`grep` for citations in `.rs`/`.py`/tests, workflows, and configs); only files with no dependents were removed; the tree was re-grepped after removal to confirm no new dangling references; verified the published docs site page count (`git ls-files 'docs/content/docs/*.mdx' | wc -l` = 42, unchanged). - Observed result: the 42-page published docs site and all wiki guides are untouched; no source or workflow references a removed file; `.changelog.md` (consumed by release.yml) and the code-cited design docs were detected as dependencies and kept; the committed `node_modules` tree is removed and `node_modules/` is gitignored so it can't be re-committed; zero "Headroom Cloud"/`headroom.dev` references remain. - Not tested: N/A — no executable code changed (only markdown, `.mdx`, and `.gitignore`), so the behavioral test suite is unaffected. ## Review Readiness - [x] I have performed a self-review - [x] This PR is ready for human review ## Checklist - [x] My code follows the project's style guidelines - [x] I have performed a self-review of my code - [x] I have commented my code, particularly in hard-to-understand areas - [x] I have made corresponding changes to the documentation - [x] My changes generate no new warnings - [ ] I have added tests that prove my fix is effective or that my feature works - [x] New and existing unit tests pass locally with my changes - [ ] I have updated the CHANGELOG.md if applicable ## Additional Notes - This branch deletes `.github/FUNDING.yml` while PR #1526 edits it — the two will be sequenced at merge (delete wins). - A follow-up option (not in this PR): also remove the internal design docs that are currently cited by the code (`REALIGNMENT/`, `docs/observability.md`, `docs/rtk-architecture.md`, `wiki/plans/`) — that requires scrubbing ~15–20 in-code citations so nothing dangles, so it's deliberately deferred. - Untracked local working files (`benchmarks/hf_pilot/`, `tools/copilot-test/`) are intentionally left out of git (not committed).
7.5 KiB
TypeScript SDK
The Headroom TypeScript SDK lets any JavaScript or TypeScript application compress LLM messages before sending them to a model. It saves tokens, reduces costs, and fits more context into every request.
Install
npm install headroom-ai
Requires a running Headroom proxy.
Quick Start
import { compress } from 'headroom-ai';
const result = await compress(messages, { model: 'gpt-4o' });
console.log(`Saved ${result.tokensSaved} tokens`);
const response = await openai.chat.completions.create({
model: 'gpt-4o',
messages: result.messages,
});
How It Works
The TypeScript SDK is an HTTP client. When you call compress(), it sends your messages to the Headroom proxy's POST /v1/compress endpoint. The proxy runs the full compression pipeline (SmartCrusher, ContentRouter, CacheAligner, etc.) and returns compressed messages. No compression logic runs in Node.js — all the heavy lifting happens in the proxy.
Your TypeScript App
│
│ compress(messages)
▼
headroom-ai (npm) ← HTTP client
│
│ POST /v1/compress
▼
Headroom Proxy / Cloud ← compression pipeline (Python)
│
│ compressed messages
▼
Your TypeScript App
│
│ openai.chat.completions.create(compressed)
▼
LLM Provider
Core API: compress()
import { compress } from 'headroom-ai';
const result = await compress(messages, {
model: 'gpt-4o', // model name (for token counting)
baseUrl: 'http://localhost:8787', // proxy URL (default)
apiKey: 'your-api-key', // optional, for authenticated endpoints
timeout: 30000, // ms (default)
fallback: true, // return uncompressed if proxy down (default)
retries: 1, // retry on transient errors (default)
});
result.messages // compressed messages (same format as input)
result.tokensBefore // original token count
result.tokensAfter // compressed token count
result.tokensSaved // tokens removed
result.compressionRatio // tokensAfter / tokensBefore
result.transformsApplied // e.g. ['router:smart_crusher:0.35']
result.compressed // false if fallback kicked in
Messages use standard OpenAI chat format: { role, content, tool_calls?, tool_call_id? }.
Environment Variables
Instead of passing options, set environment variables:
HEADROOM_BASE_URL— proxy URL (default:http://localhost:8787)HEADROOM_API_KEY— optional API key for authenticated endpoints
Reusable Client
For apps making many calls, create a client once and reuse it:
import { HeadroomClient } from 'headroom-ai';
const client = new HeadroomClient({
baseUrl: 'http://localhost:8787',
apiKey: 'your-api-key',
});
const r1 = await client.compress(messages1, { model: 'gpt-4o' });
const r2 = await client.compress(messages2, { model: 'gpt-4o' });
Framework Adapters
Vercel AI SDK
The Headroom middleware plugs directly into Vercel AI SDK's wrapLanguageModel():
import { headroomMiddleware } from 'headroom-ai/vercel-ai';
import { wrapLanguageModel, generateText } from 'ai';
import { openai } from '@ai-sdk/openai';
const model = wrapLanguageModel({
model: openai('gpt-4o'),
middleware: headroomMiddleware(),
});
// All calls through this model are automatically compressed
const { text } = await generateText({ model, messages });
The middleware intercepts messages in the transformParams hook, converts Vercel's internal format to OpenAI format, compresses via the proxy, and converts back. Your app code doesn't change.
You can also compress Vercel messages directly:
import { compressVercelMessages } from 'headroom-ai/vercel-ai';
const result = await compressVercelMessages(modelMessages, { model: 'gpt-4o' });
// result.messages is in Vercel ModelMessage[] format
OpenAI SDK
Wrap your OpenAI client to auto-compress messages on every chat.completions.create() call:
import { withHeadroom } from 'headroom-ai/openai';
import OpenAI from 'openai';
const client = withHeadroom(new OpenAI());
// Messages are compressed before sending — transparent to your code
const response = await client.chat.completions.create({
model: 'gpt-4o',
messages: longConversation,
});
Only chat.completions.create() is intercepted. All other methods (embeddings, images, audio) pass through unchanged.
Anthropic SDK
Same pattern for the Anthropic client:
import { withHeadroom } from 'headroom-ai/anthropic';
import Anthropic from '@anthropic-ai/sdk';
const client = withHeadroom(new Anthropic());
const response = await client.messages.create({
model: 'claude-sonnet-4-5-20250929',
messages: longConversation,
max_tokens: 1024,
});
Only messages.create() is intercepted. The adapter converts between Anthropic's content block format and OpenAI format automatically.
Error Handling
import { compress, HeadroomConnectionError, HeadroomAuthError } from 'headroom-ai';
try {
const result = await compress(messages, { model: 'gpt-4o', fallback: false });
} catch (error) {
if (error instanceof HeadroomAuthError) {
// Invalid API key (401)
} else if (error instanceof HeadroomConnectionError) {
// Proxy unreachable
}
}
With fallback: true (the default), connection errors and 5xx responses return the original messages uncompressed instead of throwing. Auth errors (401) and client errors (400) always throw.
Fallback Behavior
By default, compress() never blocks your app. If the proxy is unreachable:
| Scenario | fallback: true (default) |
fallback: false |
|---|---|---|
| Proxy unreachable | Returns uncompressed, compressed: false |
Throws HeadroomConnectionError |
| Proxy 503 error | Returns uncompressed after retries | Throws HeadroomCompressError |
| Invalid API key (401) | Throws HeadroomAuthError |
Throws HeadroomAuthError |
| Bad request (400) | Throws HeadroomCompressError |
Throws HeadroomCompressError |
Zero Dependencies
The headroom-ai package has no runtime dependencies. Framework SDKs (Vercel AI, OpenAI, Anthropic) are optional peer dependencies — only install what you use.
OpenClaw Plugin
The TypeScript SDK powers the headroom-openclaw plugin for OpenClaw agents. The plugin uses HeadroomClient internally to compress context during the assemble() lifecycle hook. The preferred install flow is headroom wrap openclaw; the direct plugin command is openclaw plugins install --dangerously-force-unsafe-install headroom-ai/openclaw. See the plugin source for details.
Comparison with Python SDK
| Feature | Python SDK | TypeScript SDK |
|---|---|---|
compress() |
Native (runs locally) | HTTP client (calls proxy) |
| Proxy | Built-in server | Connects to proxy |
| Vercel AI SDK | N/A | Middleware adapter |
| OpenAI SDK | HeadroomClient wrapper |
withHeadroom() wrapper |
| Anthropic SDK | HeadroomClient wrapper |
withHeadroom() wrapper |
| LangChain | HeadroomChatModel |
Use compress() directly |
| Memory system | Full (SQLite + HNSW) | Not yet (use proxy) |
| MCP server | Built-in | Not yet |
| CLI tools | headroom proxy, headroom wrap, etc. |
N/A (use Python CLI) |