headroom/pyproject.toml
github-actions[bot] aea3c35177
chore: release main (#1441)
🤖 I have created a release *beep* *boop*
---


<details><summary>0.28.0</summary>

##
[0.28.0](https://github.com/headroomlabs-ai/headroom/compare/v0.27.0...v0.28.0)
(2026-06-29)


### Features

* add --disable-kompress-fallback to restore legacy PASSTHROUGH fallback
([#1185](https://github.com/headroomlabs-ai/headroom/issues/1185))
([f309244](f309244a77))
* add first-class OpenCode support (wrap, learn, mcp install)
([#559](https://github.com/headroomlabs-ai/headroom/issues/559))
([91cd210](91cd2102d7))
* add HEADROOM_KEEPALIVE_EXPIRY to keep upstream connections warm
([#1124](https://github.com/headroomlabs-ai/headroom/issues/1124))
([85786b3](85786b33a3))
* **azure-foundry:** derive upstream URL from ANTHROPIC_FOUNDRY_RESOURCE
([#1138](https://github.com/headroomlabs-ai/headroom/issues/1138))
([e5031b0](e5031b0121))
* **cache:** attribute prompt-cache misses to TTL lapse vs prefix change
([#1313](https://github.com/headroomlabs-ai/headroom/issues/1313))
([#1343](https://github.com/headroomlabs-ai/headroom/issues/1343))
([4658721](4658721ea0))
* **code:** add Perl support to code-aware compressor
([#1125](https://github.com/headroomlabs-ai/headroom/issues/1125))
([f39858c](f39858c233))
* headroom wrap opencode / unwrap opencode CLI
([#1105](https://github.com/headroomlabs-ai/headroom/issues/1105))
([b4571cc](b4571cc346))
* **learn:** weight loops in Headroom Learn + RTK-loop eval
([#1160](https://github.com/headroomlabs-ai/headroom/issues/1160))
([14e8dc4](14e8dc4c84))
* **learn:** write per-project learnings to CLAUDE.local.md by default
([#1115](https://github.com/headroomlabs-ai/headroom/issues/1115))
([ced75e4](ced75e4718))
* **proxy:** add request timeout config
([#738](https://github.com/headroomlabs-ai/headroom/issues/738))
([c0745d4](c0745d4161))
* **proxy:** pilot hardening — inbound auth, security headers, audit
log, air-gap switch
([#1537](https://github.com/headroomlabs-ai/headroom/issues/1537))
([546ab55](546ab553dc))
* **proxy:** support glob patterns in exclude_tools
([#870](https://github.com/headroomlabs-ai/headroom/issues/870))
([#1259](https://github.com/headroomlabs-ai/headroom/issues/1259))
([a2159c0](a2159c0b66))
* **read-maturation:** activity-based hold-back Read maturation
(Mechanism B)
([#1068](https://github.com/headroomlabs-ai/headroom/issues/1068))
([723b80c](723b80c091))
* **savings:** durable savings ledger + headroom savings command
([#1127](https://github.com/headroomlabs-ai/headroom/issues/1127))
([978ffa0](978ffa0a6a))
* **wrap:** add --1m to preserve the 1M context window on wrap claude
([#1158](https://github.com/headroomlabs-ai/headroom/issues/1158))
([#1351](https://github.com/headroomlabs-ai/headroom/issues/1351))
([b50d9c1](b50d9c17ce))
* **wrap:** make tokensave the primary coding-task compressor, Serena
the backup
([#1230](https://github.com/headroomlabs-ai/headroom/issues/1230))
([dca9853](dca9853ed9))


### Bug Fixes

* **agent-evals:** Phase 0 — coding-agent accuracy A/B framework
([#1037](https://github.com/headroomlabs-ai/headroom/issues/1037))
([84f9871](84f9871e30))
* **agno:** tolerate streaming tool-call SDK objects in parser
([#1312](https://github.com/headroomlabs-ai/headroom/issues/1312))
([#1336](https://github.com/headroomlabs-ai/headroom/issues/1336))
([5986c22](5986c2260f))
* **bedrock:** add boto3 1.41 + CRT for aws login credentials
([#1486](https://github.com/headroomlabs-ai/headroom/issues/1486))
([4db3bc9](4db3bc91d9))
* bump codebase-memory-mcp to v0.8.1
([#1284](https://github.com/headroomlabs-ai/headroom/issues/1284))
([530318b](530318b425))
* **ccr:** make headroom_retrieve a hash-only full-content lookup
([#1532](https://github.com/headroomlabs-ai/headroom/issues/1532))
([c2fc4d3](c2fc4d3753))
* **ccr:** propagate --no-ccr-marker flag to all compressors
([#1022](https://github.com/headroomlabs-ai/headroom/issues/1022))
([#1197](https://github.com/headroomlabs-ai/headroom/issues/1197))
([0c9b42a](0c9b42a919))
* **ccr:** skip Anthropic marker emission when tool injection is
deferred
([#1273](https://github.com/headroomlabs-ai/headroom/issues/1273))
([2cae13d](2cae13dd79))
* **ci:** extend gitleaks allowlist to cover test fixtures + verified
examples
([#1539](https://github.com/headroomlabs-ai/headroom/issues/1539))
([d2565a6](d2565a6983))
* **ci:** guarantee model present in test shards to end cache-miss
flakiness
([#1399](https://github.com/headroomlabs-ai/headroom/issues/1399))
([2e29c72](2e29c7223f))
* **ci:** normalize Windows CRLF line endings in PR governance script
([#1012](https://github.com/headroomlabs-ai/headroom/issues/1012))
([5194388](5194388b66))
* **cli:** add explicit UTF-8 encoding to file I/O in wrap commands
([#1126](https://github.com/headroomlabs-ai/headroom/issues/1126))
([#1164](https://github.com/headroomlabs-ai/headroom/issues/1164))
([a0cb798](a0cb7982e3))
* **cli:** fall back gracefully when embedding-server sidecar is absent
([#1206](https://github.com/headroomlabs-ai/headroom/issues/1206))
([38f1404](38f1404432))
* **cli:** harden all CLI surfaces + fix docs accuracy
([#1491](https://github.com/headroomlabs-ai/headroom/issues/1491))
([bd76235](bd76235f5c))
* **cli:** wire --http2/--no-http2 (HEADROOM_HTTP2) into proxy command
([#1373](https://github.com/headroomlabs-ai/headroom/issues/1373))
([e06b616](e06b61671f))
* **cli:** wire --rpm/--tpm and HEADROOM_RPM/HEADROOM_TPM to the Click
proxy command
([#1375](https://github.com/headroomlabs-ai/headroom/issues/1375))
([8aab8f2](8aab8f22cb))
* **code:** slice tree-sitter byte offsets as UTF-8
([#1332](https://github.com/headroomlabs-ai/headroom/issues/1332))
([8238402](82384022bd))
* **code:** validate Python compressed syntax
([#1302](https://github.com/headroomlabs-ai/headroom/issues/1302))
([cbd361d](cbd361de2a))
* **code:** verify a real parse in tree-sitter availability check
([#1231](https://github.com/headroomlabs-ai/headroom/issues/1231))
([#1299](https://github.com/headroomlabs-ai/headroom/issues/1299))
([5e0bb69](5e0bb69725))
* **codex:** retag threads on init so Codex Desktop history stays
visible ([#961](https://github.com/headroomlabs-ai/headroom/issues/961))
([#1349](https://github.com/headroomlabs-ai/headroom/issues/1349))
([e6bbc40](e6bbc40b11))
* **codex:** stop pinning Codex memory MCP to one project db
([#1269](https://github.com/headroomlabs-ai/headroom/issues/1269))
([ad7993b](ad7993bf15))
* **dashboard:** include RTK stats in the historical tab
([#1324](https://github.com/headroomlabs-ai/headroom/issues/1324))
([35939c3](35939c3536))
* **deps:** remediate dependency CVEs and publish SBOM
([#1509](https://github.com/headroomlabs-ai/headroom/issues/1509))
([5771a80](5771a8020e))
* **docker:** persist session history across container revisions
([#1118](https://github.com/headroomlabs-ai/headroom/issues/1118))
([5912d65](5912d65674))
* **gemini:** offload compression to the executor
([#1382](https://github.com/headroomlabs-ai/headroom/issues/1382))
([615848e](615848eba4))
* **gemini:** resolve Google model capabilities through ModelRegistry
([#1276](https://github.com/headroomlabs-ai/headroom/issues/1276))
([17ecad9](17ecad9d89))
* **install:** guard install_agent_ensure against duplicate runtime
spawns
([#1301](https://github.com/headroomlabs-ai/headroom/issues/1301))
([8da0b4e](8da0b4e565))
* **install:** repair macOS launchd restart/start lifecycle
([#1290](https://github.com/headroomlabs-ai/headroom/issues/1290))
([da1a397](da1a3973ed))
* **install:** stop duplicating ENTRYPOINT in persistent-docker runtime
command ([#833](https://github.com/headroomlabs-ai/headroom/issues/833))
([#1348](https://github.com/headroomlabs-ai/headroom/issues/1348))
([feedead](feedead077))
* **io:** use UTF-8 with locale fallback and preserve line endings on
config/text I/O
([#1498](https://github.com/headroomlabs-ai/headroom/issues/1498))
([1baa04e](1baa04ef65))
* **kompress:** hard override keeps must-keep tokens regardless of model
score ([#1400](https://github.com/headroomlabs-ai/headroom/issues/1400))
([42612c8](42612c86df))
* **langchain:** disable streaming on wrapped model during ainvoke()
([#1287](https://github.com/headroomlabs-ai/headroom/issues/1287))
([3590046](359004646b))
* **mcp:** register managed installs with a resolvable headroom command
([#1386](https://github.com/headroomlabs-ai/headroom/issues/1386))
([22def93](22def93177))
* **mcp:** report correct savings_percent in headroom_compress
([#1106](https://github.com/headroomlabs-ai/headroom/issues/1106))
([f216e43](f216e43055))
* **opencode:** write local MCP config
([#1381](https://github.com/headroomlabs-ai/headroom/issues/1381))
([6c83790](6c83790680))
* **packaging:** move hnswlib to optional [vector] extra so [all] needs
no C++ toolchain
([#1499](https://github.com/headroomlabs-ai/headroom/issues/1499))
([80fa086](80fa086660))
* patch rtk hook script to use absolute path after register_claude_hooks
([#571](https://github.com/headroomlabs-ai/headroom/issues/571))
([b618d2d](b618d2d11a))
* **perf:** surface RTK/CLI context-tool savings in perf and the session
card ([#1433](https://github.com/headroomlabs-ai/headroom/issues/1433))
([9362747](93627471b7))
* **proxy:** add --protect-tool-results to prevent lossy compression of
exact-output Bash results
([#1374](https://github.com/headroomlabs-ai/headroom/issues/1374))
([51d4bcf](51d4bcfc11))
* **proxy:** add an Anthropic buffered read-timeout override
([#1331](https://github.com/headroomlabs-ai/headroom/issues/1331))
([3be2526](3be2526b76))
* **proxy:** add versionless Vertex AI routes for Claude Code
compatibility
([#1321](https://github.com/headroomlabs-ai/headroom/issues/1321))
([bb3e040](bb3e040a46))
* **proxy:** bind before eager preload so a hung compressor load can't
block startup
([#1500](https://github.com/headroomlabs-ai/headroom/issues/1500))
([d5ac07f](d5ac07fc45))
* **proxy:** build SSL contexts for custom CA bundles
([#1134](https://github.com/headroomlabs-ai/headroom/issues/1134))
([561ba17](561ba17ec2))
* **proxy:** forward request-id headers on the streaming path
([#1100](https://github.com/headroomlabs-ai/headroom/issues/1100))
([#1258](https://github.com/headroomlabs-ai/headroom/issues/1258))
([3d59df7](3d59df7be8))
* **proxy:** gate CCR retrieve/compress endpoints to loopback
([#1338](https://github.com/headroomlabs-ai/headroom/issues/1338))
([acafb2d](acafb2d0f6))
* **proxy:** honor force_kompress routing profile
([#996](https://github.com/headroomlabs-ai/headroom/issues/996))
([b4682d6](b4682d6f91))
* **proxy:** keep large compression results on the critical path
([#296](https://github.com/headroomlabs-ai/headroom/issues/296))
([#1352](https://github.com/headroomlabs-ai/headroom/issues/1352))
([90734b6](90734b691a))
* **proxy:** offload /v1/compress to the compression executor to stop
blocking the loop
([#1501](https://github.com/headroomlabs-ai/headroom/issues/1501))
([27e010e](27e010e38f))
* **proxy:** preserve Responses memory continuations with store=false
([#1103](https://github.com/headroomlabs-ai/headroom/issues/1103))
([cdfeeac](cdfeeacc63))
* **proxy:** queue mid-turn user messages on non-Bedrock streaming path
([#1377](https://github.com/headroomlabs-ai/headroom/issues/1377))
([b09f027](b09f027062))
* **proxy:** register interceptor in explicit transforms list when
HEADROOM_INTERCEPT_ENABLED
([#1376](https://github.com/headroomlabs-ai/headroom/issues/1376))
([55c700c](55c700c686))
* **proxy:** report real input tokens on streaming message_start
([#1132](https://github.com/headroomlabs-ai/headroom/issues/1132))
([#1305](https://github.com/headroomlabs-ai/headroom/issues/1305))
([70cc96a](70cc96a386))
* **proxy:** retry upstream 429 with Retry-After on both forwarders
([#1329](https://github.com/headroomlabs-ai/headroom/issues/1329))
([90bee89](90bee89243))
* **proxy:** retry upstream 529 overloaded like 429 on both forwarders
([#1495](https://github.com/headroomlabs-ai/headroom/issues/1495))
([547b15d](547b15dab2))
* **proxy:** stop re-compressing headroom_retrieve output and emitting
unredeemable markers
([#1323](https://github.com/headroomlabs-ai/headroom/issues/1323))
([43494ff](43494ff526))
* **proxy:** strip Codex lite header from OpenAI WebSockets
([#1543](https://github.com/headroomlabs-ai/headroom/issues/1543))
([5d3803a](5d3803a21c))
* **read-lifecycle:** persist STALE Read originals in the CCR store
([#1488](https://github.com/headroomlabs-ai/headroom/issues/1488))
([9157173](9157173018))
* recover persistent proxy feature checks and reject non-Copilot
exchange URL
([#1465](https://github.com/headroomlabs-ai/headroom/issues/1465))
([16c638b](16c638bc21))
* remove agents.md
([#1540](https://github.com/headroomlabs-ai/headroom/issues/1540))
([a7d3360](a7d3360a05))
* respect COPILOT_PROVIDER_TYPE env var when provider_type is auto
([#549](https://github.com/headroomlabs-ai/headroom/issues/549))
([24cf256](24cf256e50))
* restore token-mode compression on frozen prefixes
([#1489](https://github.com/headroomlabs-ai/headroom/issues/1489))
([8e0dadf](8e0dadfe02))
* **router:** degrade to pure-Python detection on native panic
([#1123](https://github.com/headroomlabs-ai/headroom/issues/1123))
([#1260](https://github.com/headroomlabs-ai/headroom/issues/1260))
([a00fb67](a00fb6761e))
* **rtk:** stop hook registration timing out on a forked daemon
([#1314](https://github.com/headroomlabs-ai/headroom/issues/1314))
([9758817](9758817979))
* **smart-crusher:** honor enable_ccr_marker on the opaque-blob path
([#1130](https://github.com/headroomlabs-ai/headroom/issues/1130))
([27d6f8e](27d6f8e2a7))
* **subscription:** only reset 5h contribution on real rollover, not API
jitter
([#1255](https://github.com/headroomlabs-ai/headroom/issues/1255))
([8d6c175](8d6c175d60))
* **subscription:** run transcript token scan off the event loop
([#1263](https://github.com/headroomlabs-ai/headroom/issues/1263))
([f03021f](f03021f1b6))
* surface output reduction without a restart, and explain $0.00 savings
on Python 3.14
([#1296](https://github.com/headroomlabs-ai/headroom/issues/1296))
([c30ec4c](c30ec4cda8))
* **tests:** reset whole headroom logger subtree so caplog stays
deterministic
([#1117](https://github.com/headroomlabs-ai/headroom/issues/1117))
([fda4670](fda4670ef8))
* **tls:** add HEADROOM_TLS_STRICT=0 toggle for corporate SSL inspection
([#1308](https://github.com/headroomlabs-ai/headroom/issues/1308))
([#1341](https://github.com/headroomlabs-ai/headroom/issues/1341))
([52068dd](52068dd650))
* **tokenizers:** price CJK/Kana/Hangul at ~1 token per char in
EstimatingTokenCounter
([#1093](https://github.com/headroomlabs-ai/headroom/issues/1093))
([a35fe86](a35fe86e87))
* **transforms:** gate tool string output from lossy compression
([#1307](https://github.com/headroomlabs-ai/headroom/issues/1307))
([#1387](https://github.com/headroomlabs-ai/headroom/issues/1387))
([c6c921a](c6c921a7c1))
* **websocket:** harden responses websocket origin handling
([#1481](https://github.com/headroomlabs-ai/headroom/issues/1481))
([c632023](c632023cc1))
* **windows:** pin UTF-8 encoding on text-mode subprocess calls
([#1311](https://github.com/headroomlabs-ai/headroom/issues/1311))
([d633e81](d633e8172c))
* **wrap:** add Copilot unwrap command
([#1251](https://github.com/headroomlabs-ai/headroom/issues/1251))
([b4fde0c](b4fde0c3a4))
* **wrap:** isolate proxy stdio from proxy.log on Windows
([#1191](https://github.com/headroomlabs-ai/headroom/issues/1191))
([959ab0d](959ab0de47))
* **wrap:** keep agent savings opt-in
([#1294](https://github.com/headroomlabs-ai/headroom/issues/1294))
([b829ceb](b829ceba84))
* **wrap:** show the dashboard URL when the proxy is already running
([#1313](https://github.com/headroomlabs-ai/headroom/issues/1313))
([b0146c4](b0146c4ccd))


### Performance Improvements

* **compression:** take large cold-start contexts off the synchronous
kompress path
([#1171](https://github.com/headroomlabs-ai/headroom/issues/1171))
([#1298](https://github.com/headroomlabs-ai/headroom/issues/1298))
([6c68ff4](6c68ff4e9f))
</details>

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-06-29 12:53:17 -07:00

471 lines
17 KiB
TOML
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

[build-system]
requires = ["maturin>=1.5,<2.0"]
build-backend = "maturin"
[project]
name = "headroom-ai"
version = "0.28.0"
description = "The Context Optimization Layer for LLM Applications - Cut costs by 50-90%"
readme = "README.md"
license = "Apache-2.0"
requires-python = ">=3.10"
authors = [
{ name = "Headroom Contributors" }
]
maintainers = [
{ name = "Headroom Contributors" }
]
keywords = [
"llm",
"openai",
"anthropic",
"claude",
"gpt",
"context",
"token",
"optimization",
"compression",
"caching",
"proxy",
"ai",
"machine-learning",
]
classifiers = [
"Development Status :: 4 - Beta",
"Intended Audience :: Developers",
"License :: OSI Approved :: Apache Software License",
"Operating System :: OS Independent",
"Programming Language :: Python :: 3",
"Programming Language :: Python :: 3.10",
"Programming Language :: Python :: 3.11",
"Programming Language :: Python :: 3.12",
"Programming Language :: Python :: 3.13",
"Programming Language :: Python :: 3.14",
"Topic :: Scientific/Engineering :: Artificial Intelligence",
"Topic :: Software Development :: Libraries :: Python Modules",
"Typing :: Typed",
]
dependencies = [
# Core: lightweight compression (SmartCrusher, ContentRouter, CCR, TOIN)
"tiktoken>=0.5.0", # Tokenizer for all compressors
"pydantic>=2.0.0", # Config and data models
# litellm's own metadata pins requires-python <3.14, and headroom only uses it for
# model registry / pricing / non-core providers — all lazily imported and
# ImportError-guarded. Marking it 3.14-optional lets headroom install on Python 3.14
# (core compression + the Anthropic proxy path never import litellm). See GH #956.
"litellm>=1.86.2,<2.0; python_version < '3.14'", # model registry, pricing, providers (lazy)
"click>=8.1.0", # CLI framework
"rich>=13.0.0", # Rich terminal output
"opentelemetry-api>=1.24.0", # Safe no-op OTEL API for instrumentation
"ast-grep-cli>=0.30.0", # AST-aware code slicing (CodeCompressor); binary wheel
"tomli>=2.0.0; python_version < '3.11'", # tomllib backport for helper scripts
]
[project.optional-dependencies]
# Proxy server (most common install: pip install headroom-ai[proxy])
proxy = [
"fastapi>=0.100.0",
"uvicorn>=0.23.0,<1.0",
"httpx[http2]>=0.24.0",
"openai>=2.14.0", # OpenAI API format support
"mcp>=1.0.0", # MCP server (headroom_compress, retrieve, stats)
"magika>=0.6.0", # ML content detection for ContentRouter
"zstandard>=0.20.0", # Decompress zstd request bodies (Codex, etc.)
"websockets>=13.0", # WebSocket proxy for /v1/responses (Codex gpt-5.4+)
"onnxruntime>=1.16.0", # Kompress ONNX INT8 text compression (no torch needed)
"transformers>=4.30.0,<6.0", # Tokenizer only (for Kompress)
"watchdog>=4.0.0", # File watcher for live code graph reindexing (--code-graph)
"sqlite-vec>=0.1.6", # Vector index for memory (--memory). Lightweight, no torch.
]
# Production ASGI/WSGI server — Unix-only (gunicorn does not support Windows).
# Kept separate from [proxy] so that dev, CI, and Windows users are not forced
# to install a non-functional package. Production deployments should use:
# pip install headroom-ai[proxy,proxy-prod]
proxy-prod = [
"headroom-ai[proxy]",
"gunicorn>=21.0.0; sys_platform != 'win32'",
]
# AST-based code compression (tree-sitter)
# NOTE: cap below 1.0. tree-sitter-language-pack 1.x is a breaking rewrite whose
# get_language()/get_parser() return the pack's own binding types instead of
# standalone tree_sitter.Language/Parser, so _get_parser() in
# transforms/code_compressor.py fails and code compression silently no-ops.
# The 0.x line (>=0.10,<1.0) returns standalone tree_sitter objects as expected.
code = [
"tree-sitter-language-pack>=0.10.0,<1.0",
"tree-sitter>=0.25.2,<0.26",
]
# ML-based compression with Kompress (ModernBERT).
# (The legacy [llmlingua] extra was removed in 0.9.x — no live code path used it.
# Use [ml] for the supported ML compression dependencies.)
ml = [
"torch>=2.12.1",
"transformers>=4.30.0,<6.0",
# transformers >= 5.x requires huggingface-hub >= 1.5.0,<2.0; pinning
# the floor here prevents Kompress from silently falling back to
# "unavailable" when a sibling install (e.g. `pip install
# strands-agents`) drags huggingface-hub backwards.
"huggingface-hub>=1.5.0,<2.0",
]
# Memory system (hierarchical memory with vector search).
# Uses the pure-Python sqlite-vec backend by default (VectorBackend.AUTO ->
# SQLITE_VEC), so no C++ toolchain is required. The optional HNSW backend lives
# in the [vector] extra below; installing it here would make `[all]` (which pulls
# [memory]) fail on any machine without a compiler — see #1368.
memory = [
"sqlite-vec>=0.1.6",
"sentence-transformers>=2.2.0,<6.0",
]
# Optional HNSW vector backend. Needs a C++ toolchain to build hnswlib, so it is
# kept out of [memory] and [all]; opt in with `pip install headroom-ai[vector]`
# and select it via MemoryConfig(vector_backend=VectorBackend.HNSW). The default
# sqlite-vec backend needs no compiler.
vector = [
"hnswlib>=0.8.0",
]
# Qdrant + Neo4j memory backend helpers
memory-stack = [
"mem0ai>=2.0.0,<3.0",
"qdrant-client>=1.9.0,<2.0",
"neo4j>=5.20.0,<7.0",
]
# Apple-Silicon GPU (MPS) offload for the memory embedder. Opt in at runtime with
# HEADROOM_EMBEDDER_RUNTIME=pytorch_mps. macOS-only; intentionally excluded from [all].
pytorch-mps = [
"torch>=2.12.1; sys_platform == 'darwin'",
"sentence-transformers>=2.2.0; sys_platform == 'darwin'",
]
# Semantic relevance scoring with embeddings.
# Uses `fastembed` (BAAI/bge-small-en-v1.5 by default — 33M params,
# 384 dims, ~30 MB int8-quantized ONNX). Same library + model used by
# the Rust SmartCrusher (`fastembed` crate), giving byte-equal embeddings
# across the language boundary. Replaced sentence-transformers in
# Stage 3c.1 — fastembed is faster (~2-3x), smaller (no torch
# dependency), and outranks all-MiniLM-L6-v2 on MTEB by ~6 points.
relevance = [
"fastembed>=0.4.0",
"numpy>=1.24.0",
]
# Image compression (ML-based routing + OCR)
#
# OCR backend uses ONNX Runtime regardless of Python version. The
# rapidocr ecosystem split into two flavors after 1.4.x:
# * rapidocr-onnxruntime 1.4.x — bundled-ORT package, capped at
# Python <3.13 by its requires-python metadata. Drop-in for our
# existing v1 tuple-shaped API call.
# * rapidocr 3.x — engine-agnostic core, supports Python 3.13+.
# Returns a RapidOCROutput dataclass (txts, scores, boxes, ...).
# Needs `onnxruntime` installed separately to use the ORT backend.
#
# `headroom/image/compressor.py` adapts both API shapes at runtime via
# a try/except cascade. See issue #372 for context.
image = [
"pillow>=10.0.0",
"sentencepiece>=0.1.99", # Required by SigLIP tokenizer (SiglipTokenizer)
# Python 3.63.12: keep the proven ORT-bundled package directly.
# ~15 MB ONNX models auto-downloaded on first use.
"rapidocr-onnxruntime>=1.4.0,<2; python_version<'3.13'",
# Python 3.13+: rapidocr-onnxruntime is unavailable (its wheels
# declare requires-python<3.13). Use the successor `rapidocr` 3.x
# core + `onnxruntime` engine; same ORT backend, just split into
# two packages. Total install size and inference speed unchanged.
"rapidocr>=3.0,<4; python_version>='3.13'",
"onnxruntime>=1.7,<2; python_version>='3.13'",
]
# Report generation
reports = [
"jinja2>=3.0.0",
]
# Binary spreadsheet ingestion (.xlsx / .xls -> tabular text)
spreadsheet = [
"openpyxl>=3.1.0", # .xlsx
"xlrd>=2.0.1", # legacy .xls
]
# OpenTelemetry metrics export
otel = [
"opentelemetry-sdk>=1.24.0",
"opentelemetry-exporter-otlp-proto-http>=1.24.0",
]
# any-llm multi-provider backend (requires Python 3.11+)
anyllm = [
"any-llm-sdk>=1.0.0; python_version >= '3.11'",
]
# LangChain integration
langchain = [
"langchain-core>=1.3.3,<4.0",
"langchain-openai>=1.1.14,<2.0",
]
# Agno agent framework integration
agno = [
"agno>=1.0.0",
]
# AWS Strands Agents SDK integration
strands = [
"strands-agents>=0.1.0",
]
# MCP server for Claude Code integration
mcp = [
"mcp>=1.0.0",
"httpx>=0.24.0",
]
# Voice filler detection
voice = [
"onnxruntime>=1.16.0",
"transformers>=4.30.0,<6.0",
"torch>=2.12.1",
]
# Voice training (includes voice deps + training extras)
voice-train = [
"headroom-ai[voice]",
"datasets>=2.14.0",
"accelerate>=0.20.0",
]
# Evaluation framework
evals = [
"datasets>=2.14.0",
"sentence-transformers>=2.2.0,<6.0",
"numpy>=1.24.0",
"scikit-learn>=1.3.0",
"anthropic>=0.18.0",
"openai>=1.0.0",
]
# AWS Bedrock backend
bedrock = [
# `aws login` (IAM Identity Provider / console-login, DPoP) requires
# boto3 >= 1.41.0 AND the AWS Common Runtime (CRT) per AWS docs
# ("Boto3 1.41.0 or later with CRT"). CRT is a separate install — pull it
# via the botocore [crt] extra (awscrt). Without it, resolving `aws login`
# credentials raises botocore's MissingDependencyException.
"boto3>=1.41.0",
"botocore[crt]>=1.41.0",
]
# HTML content extraction
html = [
"trafilatura>=1.6.0",
]
# Comprehensive LLM benchmarks
benchmark = [
"lm-eval[api]>=0.4.0",
"openai>=1.0.0",
"anthropic>=0.18.0",
]
# Development dependencies
dev = [
"pytest>=7.0.0",
"pytest-cov>=4.0.0",
"pytest-asyncio>=0.21.0",
"ruff>=0.1.0",
"mypy>=1.0.0",
"pre-commit>=3.0.0",
"openai>=1.0.0",
"anthropic>=0.18.0",
"litellm>=1.86.2,<2.0; python_version < '3.14'", # see core deps note (GH #956)
"fastapi>=0.100.0",
"uvicorn>=0.23.0,<1.0",
"httpx[http2]>=0.24.0",
"websockets>=13.0",
"opentelemetry-sdk>=1.24.0",
"opentelemetry-exporter-otlp-proto-http>=1.24.0",
"ollama>=0.4.0",
"langchain-ollama>=0.2.0",
"hnswlib>=0.8.0",
"sqlite-vec>=0.1.6",
"sentence-transformers>=2.2.0,<6.0",
"numpy>=1.24.0",
"openpyxl>=3.1.0", # exercises spreadsheet_ingest (.xlsx) in the test suite
]
# All optional dependencies (everything you need)
#
# `benchmark` is deliberately EXCLUDED from `[all]`. It installs the
# EleutherAI lm-evaluation-harness (lm-eval), which headroom invokes as an
# external subprocess (`python -m lm_eval`) — it is never imported as a
# library, so it is not a true runtime dependency. lm-eval pulls two
# transitive deps with unpatchable High CVEs (sqlitedict CVE-2024-35515,
# nltk CVE-2026-54293 via rouge-score), neither of which has an upstream
# fix. Keeping `benchmark` out of `[all]` means `pip install
# headroom-ai[all]` is CVE-free; researchers who need the accuracy harness
# opt in explicitly with `pip install headroom-ai[benchmark]`.
all = [
"headroom-ai[proxy,code,ml,memory,relevance,image,reports,otel,evals,voice,html,mcp,spreadsheet]",
]
[project.scripts]
headroom = "headroom.cli:main"
[project.urls]
Homepage = "https://headroom-docs.vercel.app"
Documentation = "https://headroom-docs.vercel.app/docs"
Repository = "https://github.com/chopratejas/headroom"
Issues = "https://github.com/chopratejas/headroom/issues"
Changelog = "https://github.com/chopratejas/headroom/blob/main/CHANGELOG.md"
# llms.txt convention (llmstxt.org) — point AI agents / LLM crawlers
# at the auto-generated docs index so they can resolve install paths
# and entry points without a follow-up fetch.
"AI / LLM Index" = "https://headroom-docs.vercel.app/llms.txt"
# Maturin builds a single wheel containing both the Python source under
# `headroom/` AND the compiled Rust extension `headroom/_core.so` (cdylib
# from `crates/headroom-py`). One `pip install headroom-ai` ships everything
# atomically — no separate `headroom-core-py` package, no chicken-and-egg,
# no PIP_FIND_LINKS plumbing. Phase A0's runtime fail-loud check still
# exists but only fires if someone forces an sdist install on a platform
# without a wheel and the rust toolchain isn't available to compile it.
# Constrain transitive dependencies that have CVEs requiring minimum versions.
# These packages don't appear as direct headroom deps but are pulled in
# transitively; the floor pins below ensure uv resolves to patched versions.
[tool.uv]
constraint-dependencies = [
# GHSA-5239-wwwm-4pmq (Low) — transitive via rich; fix at 2.20.0
"pygments>=2.20.0",
# GHSA-4xgf-cpjx-pc3j (Medium) — transitive via mcp; fix at 2.14.2
"pydantic-settings>=2.14.2",
# GHSA-mv93-w799-cj2w + 4 others (High) — transitive via lm-eval; fix at 3.1.50
"gitpython>=3.1.50",
# GHSA-f4xh-w4cj-qxq8 (High) — transitive via langchain-core; fix at 0.8.18
"langsmith>=0.9.0",
]
# Pin the project's package index to public PyPI. Without this, `uv lock`
# inherits the developer's user-level `~/.config/uv/uv.toml` index
# setting — including private/internal mirrors like
# `pypi.netflix.net/simple` — and bakes those URLs into uv.lock, which
# then breaks CI on every public runner that can't reach the mirror.
# Declaring the index in pyproject.toml makes the project authoritative
# regardless of who runs `uv lock`.
[[tool.uv.index]]
name = "pypi"
url = "https://pypi.org/simple/"
default = true
[tool.maturin]
# Where the Python package lives. With `python-source = "."` and the
# package directory `headroom/` at repo root, maturin includes every file
# under `headroom/` in the wheel — that picks up the dashboard HTML
# templates and bundled YAML configs. `LICENSE` and `NOTICE` are listed
# explicitly because maturin sdists do not get the package-directory
# treatment wheels do, and PEP 639 auto-discovery emits both files into
# `License-File:` metadata — PyPI rejects sdists whose declared license
# files are missing from the tarball with `400 License-File X does not
# exist in distribution file`.
include = [
{ path = "LICENSE", format = "sdist" },
{ path = "NOTICE", format = "sdist" },
]
python-source = "."
module-name = "headroom._core"
# The cdylib source lives under `crates/headroom-py`. Maturin invokes
# `cargo build` with this manifest to produce `_core.cdylib`, then injects
# the resulting `.so` into the wheel at `headroom/_core.so`.
manifest-path = "crates/headroom-py/Cargo.toml"
features = ["extension-module"]
# Forbid building without the cdylib feature — bare `cargo build` won't
# produce a usable Python extension. Maturin's default `bindings` is "pyo3"
# which is correct here (see `crates/headroom-py/src/`).
bindings = "pyo3"
[tool.ruff]
target-version = "py310"
line-length = 100
[tool.ruff.lint]
select = [
"E", # pycodestyle errors
"W", # pycodestyle warnings
"F", # pyflakes
"I", # isort
"B", # flake8-bugbear
"C4", # flake8-comprehensions
"UP", # pyupgrade
]
ignore = [
"E501", # line too long (handled by formatter)
"B008", # do not perform function calls in argument defaults
"B905", # zip without strict parameter
]
[tool.ruff.lint.isort]
known-first-party = ["headroom"]
[tool.ruff.format]
quote-style = "double"
indent-style = "space"
[tool.mypy]
python_version = "3.10"
warn_return_any = true
warn_unused_configs = true
disallow_untyped_defs = true
ignore_missing_imports = true
# Per-module overrides for modules with dynamic typing patterns
[[tool.mypy.overrides]]
module = [
"headroom.proxy.server",
"headroom.proxy.cost",
"headroom.proxy.prometheus_metrics",
"headroom.proxy.semantic_cache",
"headroom.proxy.rate_limiter",
"headroom.proxy.request_logger",
"headroom.proxy.helpers",
"headroom.integrations.langchain",
"headroom.integrations.mcp",
"headroom.ccr.mcp_server",
"headroom.relevance.embedding",
"headroom.reporting.generator",
]
disallow_untyped_defs = false
[[tool.mypy.overrides]]
module = [
"headroom.tokenizers.*",
"headroom.providers.litellm",
"headroom.providers.google",
]
disallow_untyped_defs = false
warn_return_any = false
# Handler mixins use self.* from HeadroomProxy via duck typing — mypy can't resolve these
[[tool.mypy.overrides]]
module = ["headroom.proxy.handlers.*"]
disallow_untyped_defs = false
ignore_errors = true
# Ignore third-party stubs with syntax errors
[[tool.mypy.overrides]]
module = ["mlx.*"]
ignore_errors = true
[tool.pytest.ini_options]
testpaths = ["tests"]
python_files = ["test_*.py"]
python_functions = ["test_*"]
addopts = "-v --tb=short"
asyncio_mode = "auto"
filterwarnings = [
# pyo3 Unsendable parsers emit an unraisable warning when GC drops them on a
# test-teardown thread; this is a test-harness artifact, not a production issue
# (production threads are long-lived and drop their parsers on themselves).
"ignore::pytest.PytestUnraisableExceptionWarning",
]
markers = [
"slow: slow tests (model loads, large fixtures)",
"real_llm: tests that hit real LLM APIs; skipped unless explicitly enabled",
"live: opt-in multi-turn tests that hit real upstream APIs; require provider keys",
]
[tool.coverage.run]
source = ["headroom"]
branch = true
omit = [
"headroom/cli.py",
"*/tests/*",
]
[tool.coverage.report]
exclude_lines = [
"pragma: no cover",
"def __repr__",
"raise NotImplementedError",
"if TYPE_CHECKING:",
"if __name__ == .__main__.:",
]