headroom/pyproject.toml
Tejas Chopra 93f2d7a2da
chore: release main (#2792)
🤖 I have created a release *beep* *boop*
---


<details><summary>0.35.0</summary>

##
[0.35.0](https://github.com/headroomlabs-ai/headroom/compare/v0.34.0...v0.35.0)
(2026-08-12)


### Features

* **beacon:** allowlist the routing summary key
([#2818](https://github.com/headroomlabs-ai/headroom/issues/2818))
([7940c05](7940c05ebf))
* **beacon:** hourly R2 compaction, per-strategy savings, and a stack
that reports
([#2853](https://github.com/headroomlabs-ai/headroom/issues/2853))
([e0870ef](e0870ef931))
* **cli,pricing:** add CLI extension seam and prompt-cache TTL pricing
([#2802](https://github.com/headroomlabs-ai/headroom/issues/2802))
([6ec3e34](6ec3e3478a))


### Bug Fixes

* **anthropic:** strip first-party tool search on custom upstreams
([#2539](https://github.com/headroomlabs-ai/headroom/issues/2539))
([7f6950b](7f6950be34))
* **backends/anyllm:** convert Anthropic tools and tool_choice to OpenAI
shape
([0d6866b](0d6866b91a))
* **backends/anyllm:** stream tool_use blocks and map finish_reason on
the streaming path
([e4904e2](e4904e23a6))
* **backends/litellm:** None-guard core token counts in OpenAI usage
block ([#2324](https://github.com/headroomlabs-ai/headroom/issues/2324))
([12f9f58](12f9f58cb3))
* **beacon:** report all-layers savings, not context-compression only
([#2796](https://github.com/headroomlabs-ai/headroom/issues/2796))
([e9a24f3](e9a24f3ec1))
* **beacon:** split session failures by status code
([#2815](https://github.com/headroomlabs-ai/headroom/issues/2815))
([2954e37](2954e37048))
* **cache:** bound compression cache bookkeeping
([0ae948c](0ae948c151))
* **cache:** enforce Anthropic's 1h-before-5m cache_control ordering
before forwarding
([#2941](https://github.com/headroomlabs-ai/headroom/issues/2941))
([3752458](3752458022))
* **cache:** mirror client cache_control positions instead of
single-marker consolidation
([def3d76](def3d76e5a))
* **cache:** stabilize Anthropic block-growing lineages
([#2917](https://github.com/headroomlabs-ai/headroom/issues/2917))
([1a04c95](1a04c957f5))
* **ccr:** avoid injecting tool on chat streaming
([d0c1f5b](d0c1f5b8ad))
* **ccr:** preserve exact SQLite TTL boundary
([#2669](https://github.com/headroomlabs-ai/headroom/issues/2669))
([d0a86d4](d0a86d409f))
* **ccr:** report embedded hashes from compress endpoint
([#717](https://github.com/headroomlabs-ai/headroom/issues/717))
([685ebe4](685ebe457d))
* **ccr:** resolve &lt;&lt;ccr:...&gt;&gt; markers inline when no
retrieve-tool path exists
([#2512](https://github.com/headroomlabs-ai/headroom/issues/2512))
([ce8ce83](ce8ce8313f))
* **ccr:** tolerate null/malformed OpenAI data in response handling
([#2467](https://github.com/headroomlabs-ai/headroom/issues/2467))
([e583e08](e583e082d8))
* **ci:** publish latest from the root Docker manifest
([#2252](https://github.com/headroomlabs-ai/headroom/issues/2252))
([5568d73](5568d738af))
* **claude:** stop forcing tool search on Foundry
([#2477](https://github.com/headroomlabs-ai/headroom/issues/2477))
([7981396](798139608c))
* **cli/update:** let install ownership win over bare /.dockerenv so
venv installs self-update
([#2830](https://github.com/headroomlabs-ai/headroom/issues/2830))
([7092b53](7092b53c46))
* **codex:** route alpha search through the Codex backend
([#2538](https://github.com/headroomlabs-ai/headroom/issues/2538))
([a540eb2](a540eb2c61))
* **content-router:** protect custom-tag blocks before mixed-content
section split
([d7bc1e2](d7bc1e275f))
* **deps:** bump h2 to 4.4.1 for CVE-2026-71554
([#2839](https://github.com/headroomlabs-ai/headroom/issues/2839))
([564e0a8](564e0a8d0f))
* **deps:** enforce audited transitive dependency floors
([#2791](https://github.com/headroomlabs-ai/headroom/issues/2791))
([64e2039](64e203931b))
* **doctor:** flag `ollama launch claude` proxy bypass instead of
misdirecting
([#2566](https://github.com/headroomlabs-ai/headroom/issues/2566))
([7f24d69](7f24d695ee))
* emit SSE ping before message_start on Bedrock streaming path (issue
[#902](https://github.com/headroomlabs-ai/headroom/issues/902))
([#1080](https://github.com/headroomlabs-ai/headroom/issues/1080))
([4dab254](4dab254d52))
* **gemini:** resolve native CCR retrieval calls
([#2253](https://github.com/headroomlabs-ai/headroom/issues/2253))
([2483f57](2483f57002))
* **health:** label kompress as degraded/optional when not yet loaded
([#2865](https://github.com/headroomlabs-ai/headroom/issues/2865))
([8949371](89493714d2))
* **image:** decouple routing types from trained_router so importing the
compressor doesn't import torch
([#2513](https://github.com/headroomlabs-ai/headroom/issues/2513))
([#2537](https://github.com/headroomlabs-ai/headroom/issues/2537))
([d7cf981](d7cf981093))
* **install/windows:** register persistent-task from S4U hidden XML
([#2453](https://github.com/headroomlabs-ai/headroom/issues/2453))
([#2459](https://github.com/headroomlabs-ai/headroom/issues/2459))
([1edaeb8](1edaeb8b76))
* **install:** don't crash the PowerShell installer when $PROFILE is
unset ([#2469](https://github.com/headroomlabs-ai/headroom/issues/2469))
([fc5c4e2](fc5c4e239c))
* **install:** trust Docker bridge for dashboard metadata
([e044139](e044139001))
* **install:** use --userns=keep-id under Podman so bind-mount writes
don't fail
([#2846](https://github.com/headroomlabs-ai/headroom/issues/2846))
([3488f8d](3488f8d4b5))
* **learn/gemini:** stop double-counting session tokens
([#2230](https://github.com/headroomlabs-ai/headroom/issues/2230))
([29d8a5e](29d8a5e563))
* **learn/grok:** detect a Windows absolute project path
([#2283](https://github.com/headroomlabs-ai/headroom/issues/2283))
([e240df2](e240df2b69))
* **learn:** stop classifying a successful exit code 0 as an error
([#2289](https://github.com/headroomlabs-ai/headroom/issues/2289))
([a24fe7d](a24fe7dcbf))
* **litellm:** add async_post_call_success_hook to HeadroomCallback
([#1322](https://github.com/headroomlabs-ai/headroom/issues/1322))
([3107994](3107994aed))
* **litellm:** don't forward a caller key the target cannot accept
([#2883](https://github.com/headroomlabs-ai/headroom/issues/2883))
([2f2950a](2f2950a626))
* **memory:** bound the TrafficLearner pending-pattern accumulator
(memory leak)
([#2579](https://github.com/headroomlabs-ai/headroom/issues/2579))
([1f5feff](1f5fefffd3))
* **memory:** close DirectMem0 resources
([6596182](65961827cf))
* **memory:** close MCP backend on shutdown
([4bd8ecd](4bd8ecd1e3))
* **memory:** don't crash inline memory extraction on a non-object
&lt;memory&gt; block
([#2470](https://github.com/headroomlabs-ai/headroom/issues/2470))
([e00c6ff](e00c6ff81c))
* **memory:** keep vector metadata in sync
([#2295](https://github.com/headroomlabs-ai/headroom/issues/2295))
([c471800](c471800e8e))
* **memory:** make explicit-project and user store keys
collision-resistant
([#2231](https://github.com/headroomlabs-ai/headroom/issues/2231))
([f840d5f](f840d5f2fe))
* **memory:** skip &lt;system-reminder&gt; blocks when building the
retrieval query
([#2195](https://github.com/headroomlabs-ai/headroom/issues/2195))
([#2541](https://github.com/headroomlabs-ai/headroom/issues/2541))
([4e5a67a](4e5a67a342))
* **memory:** sync FTS5 and vector indexes on CLI
delete/edit/prune/purge
([fd4628d](fd4628d821))
* **oauth2:** make repository lint checks pass
([c85abf7](c85abf7a87))
* **observability:** aggregate tool savings in OTEL
([#2936](https://github.com/headroomlabs-ai/headroom/issues/2936))
([941c25d](941c25d31e))
* **onnx:** stop ONNX thread pools from spinning idle cores
([#2495](https://github.com/headroomlabs-ai/headroom/issues/2495))
([#2540](https://github.com/headroomlabs-ai/headroom/issues/2540))
([5c561bd](5c561bd913))
* **openai:** skip Responses tool-search deferral for clients that
cannot execute it
([#2696](https://github.com/headroomlabs-ai/headroom/issues/2696))
([54ea28d](54ea28d983))
* **opencode:** ship the transport hook-shim so wheel installs route
Node child traffic
([702dbc5](702dbc5902))
* **providers/anthropic:** don't crash token estimation on null
tool_calls
([#2472](https://github.com/headroomlabs-ai/headroom/issues/2472))
([08466f3](08466f3cae))
* **providers/openai:** bound tiktoken vocab loads with the guarded
loader
([#2554](https://github.com/headroomlabs-ai/headroom/issues/2554))
([0805e8e](0805e8e410))
* **proxy/anthropic:** inject headroom_retrieve whenever a CCR marker is
present, not only for new markers
([#2848](https://github.com/headroomlabs-ai/headroom/issues/2848))
([3808f60](3808f60ca6))
* **proxy/anthropic:** None-guard usage token counts on the direct
buffered path
([#2434](https://github.com/headroomlabs-ai/headroom/issues/2434))
([2b5ee7c](2b5ee7cde8))
* **proxy/anthropic:** run tool-search history repair after turn hooks
([c6f9948](c6f99482e1))
* **proxy/batch:** don't crash an OpenAI batch on a valid-JSON
non-object line
([#2316](https://github.com/headroomlabs-ai/headroom/issues/2316))
([1f2c681](1f2c681c0b))
* **proxy/bedrock:** report uncached input tokens from backend usage,
not the live-zone count
([#2318](https://github.com/headroomlabs-ai/headroom/issues/2318))
([c19e412](c19e412b33))
* **proxy/gemini:** keep streaming-parity baseline so eligible_pct can't
exceed 100
([#2824](https://github.com/headroomlabs-ai/headroom/issues/2824))
([b97c7c6](b97c7c6e99))
* **proxy/metrics:** cap client-supplied model label cardinality
([#2480](https://github.com/headroomlabs-ai/headroom/issues/2480))
([e24a7e6](e24a7e66b9))
* **proxy/metrics:** escape label values in the Prometheus export
([#2463](https://github.com/headroomlabs-ai/headroom/issues/2463))
([6a53861](6a53861063))
* **proxy/openai:** don't crash the Responses memory tool loops on null
arguments
([#2273](https://github.com/headroomlabs-ai/headroom/issues/2273))
([a30db2c](a30db2cae4))
* **proxy/openai:** feed Codex WS traffic into the traffic learner
([#2334](https://github.com/headroomlabs-ai/headroom/issues/2334))
([f669149](f669149769))
* **proxy/openai:** run response hooks on Responses, and bill their
re-drives
([#2872](https://github.com/headroomlabs-ai/headroom/issues/2872))
([675d13f](675d13f08d))
* **proxy:** allow settings routes for trusted gateway/dashboard clients
([#2491](https://github.com/headroomlabs-ai/headroom/issues/2491))
([a5b0a8f](a5b0a8f4cc))
* **proxy:** cache litellm model resolution to stop repeated Provider
List spam
([99f07e7](99f07e7bbd))
* **proxy:** cancel periodic TOIN task on shutdown
([739fdef](739fdef423))
* **proxy:** close the upstream stream when a streaming body is never
consumed
([0951663](0951663562))
* **proxy:** compress cache-mode cold starts and tag prefix-mismatch
passthrough
([#2365](https://github.com/headroomlabs-ai/headroom/issues/2365))
([aaeba0a](aaeba0a319))
* **proxy:** emit request log timestamps in UTC
([620028f](620028fa18))
* **proxy:** enable tool search by default and repair poisoned
transcripts
([#2807](https://github.com/headroomlabs-ai/headroom/issues/2807))
([0237cbf](0237cbffbb))
* **proxy:** gate mid-turn message coalescing to Claude Code clients
([#1643](https://github.com/headroomlabs-ai/headroom/issues/1643))
([a4bd2e6](a4bd2e62a5))
* **proxy:** give each Codex /v1/responses WS turn a unique request_id
([#2164](https://github.com/headroomlabs-ai/headroom/issues/2164))
([d02df10](d02df10758))
* **proxy:** graceful shutdown and reliable Ctrl+C exit
([#621](https://github.com/headroomlabs-ai/headroom/issues/621))
([17cdb18](17cdb185bc))
* **proxy:** guard telemetry and TOIN endpoints
([cde1513](cde1513c91))
* **proxy:** include tool_search_deferral savings in the savings ledger
([12149f7](12149f7446))
* **proxy:** pass through cross-region prefixed Bedrock model IDs
directly
([#2330](https://github.com/headroomlabs-ai/headroom/issues/2330))
([64cb46e](64cb46e24b))
* **proxy:** port session-sticky beta headers to the Rust proxy
([#2381](https://github.com/headroomlabs-ai/headroom/issues/2381))
([f6398a6](f6398a6476))
* **proxy:** preserve merged session and quarantine contracts
([#2943](https://github.com/headroomlabs-ai/headroom/issues/2943))
([039cd24](039cd2431a))
* **proxy:** preserve signed Anthropic thinking blocks on outbound
re-serialize
([#2254](https://github.com/headroomlabs-ai/headroom/issues/2254))
([dc163bc](dc163bcd1c))
* **proxy:** stop discarding compressed Codex WS later-frame payloads
([#2823](https://github.com/headroomlabs-ai/headroom/issues/2823))
([4ec416d](4ec416df88))
* **proxy:** time-cap the compression timeout-debt quarantine
([#2360](https://github.com/headroomlabs-ai/headroom/issues/2360))
([#2412](https://github.com/headroomlabs-ai/headroom/issues/2412))
([c5a08d2](c5a08d22e0))
* **proxy:** unwrap Hermes tool_call bridge in tool name map
([#2717](https://github.com/headroomlabs-ai/headroom/issues/2717))
([a97b824](a97b82413b))
* publish headroom-opencode in release workflow
([#2372](https://github.com/headroomlabs-ai/headroom/issues/2372))
([7859154](78591545ce))
* **settings:** accept documented HEADROOM_* env names as settings keys
([#2833](https://github.com/headroomlabs-ai/headroom/issues/2833))
([de9e052](de9e0523da))
* **subscription:** dedup transcript usage by message id
([#2340](https://github.com/headroomlabs-ai/headroom/issues/2340) token
inflation)
([#2408](https://github.com/headroomlabs-ai/headroom/issues/2408))
([74275b7](74275b7c3e))
* **toin:** bound private query and pattern retention
([8cd1380](8cd138039e))
* **tokenizer:** coerce non-string tool_call fields before counting
([#2801](https://github.com/headroomlabs-ai/headroom/issues/2801))
([b6f9877](b6f9877c78))
* **tokenizer:** price CJK in the Rust fixed-ratio estimator (Python
parity)
([#2260](https://github.com/headroomlabs-ai/headroom/issues/2260))
([6840153](6840153473))
* **transforms/adaptive-sizer:** honor max_k on small-input fast path
([#2319](https://github.com/headroomlabs-ai/headroom/issues/2319))
([8a90523](8a90523209))
* **transforms/smart_crusher:** don't crash on a tool call with a null
function
([#2232](https://github.com/headroomlabs-ai/headroom/issues/2232))
([3bb02f8](3bb02f8f75))
* Vertex model pricing shows $0.00 for versioned model names and
vertex:anthropic provider
([#2517](https://github.com/headroomlabs-ai/headroom/issues/2517))
([eb5b5e4](eb5b5e4198))
* **wrap/claude:** keep --1m effective when an explicit --model is
passed through
([c093bf1](c093bf11eb))
* **wrap/opencode:** verify the opencode binary before mutating config
([ae38486](ae384862a4))
* **wrap/serena:** install Serena from the serena-agent PyPI wheel, not
the git source
([d7b25ae](d7b25ae3bb))
* **wrap:** honor Copilot OAuth wire-api override and model default
([#2387](https://github.com/headroomlabs-ai/headroom/issues/2387))
([1db6d88](1db6d88ab4))
* **wrap:** serialize shared proxy startup
([#2946](https://github.com/headroomlabs-ai/headroom/issues/2946))
([e540d64](e540d64feb))
* **wrap:** stop the launch cwd from shadowing the installed package in
the proxy subprocess
([#2843](https://github.com/headroomlabs-ai/headroom/issues/2843))
([c49be26](c49be269a1))


### Performance Improvements

* cut hot-path latency 27% (token-count memo, startup preloads, JSON
scan memo)
([#2838](https://github.com/headroomlabs-ai/headroom/issues/2838))
([53af90d](53af90d68c))
* **proxy:** bound upstream calls and hot-path costs
([#2852](https://github.com/headroomlabs-ai/headroom/issues/2852))
([f624d3a](f624d3a00a))
* **subscription:** skip transcripts older than the window in
compute_window_tokens
([#2861](https://github.com/headroomlabs-ai/headroom/issues/2861))
([91d6bf3](91d6bf33cd))


### Dependencies

* bump brace-expansion from 5.0.7 to 5.0.9 in /docs
([#2751](https://github.com/headroomlabs-ai/headroom/issues/2751))
([56ee57b](56ee57be98))
* bump bytesize from 1.3.3 to 2.4.2
([#2286](https://github.com/headroomlabs-ai/headroom/issues/2286))
([6448545](6448545a7f))
* bump hf-hub from 0.4.3 to 0.5.0
([#2285](https://github.com/headroomlabs-ai/headroom/issues/2285))
([4925bf6](4925bf6a82))
* bump next from 16.2.10 to 16.3.0 in /docs
([#2750](https://github.com/headroomlabs-ai/headroom/issues/2750))
([0fd0b99](0fd0b996a4))
* bump postcss from 8.5.19 to 8.5.25 in /plugins/openclaw
([#2749](https://github.com/headroomlabs-ai/headroom/issues/2749))
([cd60ee9](cd60ee9ae8))
* bump postcss from 8.5.19 to 8.5.25 in /plugins/opencode
([#2748](https://github.com/headroomlabs-ai/headroom/issues/2748))
([ff4e016](ff4e0167bb))
* bump postcss from 8.5.19 to 8.5.25 in /sdk/typescript
([#2747](https://github.com/headroomlabs-ai/headroom/issues/2747))
([267c2bd](267c2bdcb5))
* bump postcss from 8.5.19 to 8.5.26 in /docs
([#2881](https://github.com/headroomlabs-ai/headroom/issues/2881))
([e6e5826](e6e5826423))
* bump ruff from 0.15.17 to 0.15.22 in the pip-minor-patch group
([#2501](https://github.com/headroomlabs-ai/headroom/issues/2501))
([ecf130d](ecf130d3ac))
* bump rusqlite from 0.32.1 to 0.40.1
([#2287](https://github.com/headroomlabs-ai/headroom/issues/2287))
([522faa1](522faa1a59))
* bump the cargo-minor-patch group across 1 directory with 22 updates
([#2916](https://github.com/headroomlabs-ai/headroom/issues/2916))
([148d860](148d8605e2))
</details>

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

---------

Co-authored-by: JD Davis <mxjerrett@gmail.com>
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-12 19:02:51 -05:00

515 lines
20 KiB
TOML
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

[build-system]
requires = ["maturin>=1.5,<2.0"]
build-backend = "maturin"
[project]
name = "headroom-ai"
version = "0.35.0"
description = "The Context Optimization Layer for LLM Applications - Cut costs by 50-90%"
readme = "README.md"
license = "Apache-2.0"
requires-python = ">=3.10"
authors = [
{ name = "Headroom Contributors" }
]
maintainers = [
{ name = "Headroom Contributors" }
]
keywords = [
"llm",
"openai",
"anthropic",
"claude",
"gpt",
"context",
"token",
"optimization",
"compression",
"caching",
"proxy",
"ai",
"machine-learning",
]
classifiers = [
"Development Status :: 4 - Beta",
"Intended Audience :: Developers",
"License :: OSI Approved :: Apache Software License",
"Operating System :: OS Independent",
"Programming Language :: Python :: 3",
"Programming Language :: Python :: 3.10",
"Programming Language :: Python :: 3.11",
"Programming Language :: Python :: 3.12",
"Programming Language :: Python :: 3.13",
"Programming Language :: Python :: 3.14",
"Topic :: Scientific/Engineering :: Artificial Intelligence",
"Topic :: Software Development :: Libraries :: Python Modules",
"Typing :: Typed",
]
dependencies = [
# Core: lightweight compression (SmartCrusher, ContentRouter, CCR, TOIN)
"tiktoken>=0.5.0", # Tokenizer for all compressors
"pydantic>=2.0.0", # Config and data models
# litellm's own metadata pins requires-python <3.14, and headroom only uses it for
# model registry / pricing / non-core providers — all lazily imported and
# ImportError-guarded. Marking it 3.14-optional lets headroom install on Python 3.14
# (core compression + the Anthropic proxy path never import litellm). See GH #956.
"litellm>=1.86.2,<2.0; python_version < '3.14'", # model registry, pricing, providers (lazy)
"click>=8.3.3", # CLI framework; PYSEC-2026-2132 fix (command injection in click.edit())
"rich>=13.0.0", # Rich terminal output
"opentelemetry-api>=1.24.0", # Safe no-op OTEL API for instrumentation
# AST-aware code slicing (CodeCompressor); binary wheel. 0.44.1 is excluded:
# that PyPI release was a compromised supply-chain build shipping an
# info-stealer `sg.exe` (Trojan:Win64/Lazy!MTB) alongside the real binary
# (GH #2332). The `!=` keeps every other release installable, including a
# future patched one.
"ast-grep-cli>=0.30.0,!=0.44.1",
"pyyaml>=6.0", # omp wrap: parse/merge omp's models.yml registry
"tomli>=2.0.0; python_version < '3.11'", # tomllib backport for helper scripts
"tomlkit>=0.13.0,<1.0", # Loss-minimizing Codex config recovery
]
[project.optional-dependencies]
# Proxy server (most common install: pip install headroom-ai[proxy])
proxy = [
"fastapi>=0.100.0",
"uvicorn>=0.23.0,<1.0",
# LiteLLM provider backends (e.g. openrouter) expect orjson at runtime but
# litellm only declares it under its own [proxy] extra (GH #2056).
"orjson>=3.9.14; platform_python_implementation != 'PyPy'",
"httpx[http2]>=0.24.0",
"openai>=2.14.0", # OpenAI API format support
"mcp>=1.28.1,<2.0.0", # MCP server (headroom_compress, retrieve, stats)
"magika>=0.6.0", # ML content detection for ContentRouter
"zstandard>=0.20.0", # Decompress zstd request bodies (Codex, etc.)
"websockets>=13.0", # WebSocket proxy for /v1/responses (Codex gpt-5.4+)
"onnxruntime>=1.16.0", # Kompress ONNX INT8 text compression (no torch needed)
"transformers>=5.5.0,<6.0", # Tokenizer only (for Kompress)
"watchdog>=4.0.0", # File watcher for live code graph reindexing (--code-graph)
"sqlite-vec>=0.1.6", # Vector index for memory (--memory). Lightweight, no torch.
]
# Production ASGI/WSGI server — Unix-only (gunicorn does not support Windows).
# Kept separate from [proxy] so that dev, CI, and Windows users are not forced
# to install a non-functional package. Production deployments should use:
# pip install headroom-ai[proxy,proxy-prod]
proxy-prod = [
"headroom-ai[proxy]",
"gunicorn>=21.0.0; sys_platform != 'win32'",
]
# AST-based code compression (tree-sitter)
# tree-sitter-language-pack >=1.0 removed the bundled tree-sitter package and
# switched to an incompatible internal node API (.kind vs .type, callable
# root_node, etc.). Pin to <1.0.0 so that tree-sitter>=0.25.2 is pulled in
# as a transitive dependency and the existing code_compressor.py node-walk
# logic continues to work. See issue #1216.
code = [
"tree-sitter-language-pack>=0.10.0,<1.0",
"tree-sitter>=0.25.2,<0.27",
]
# ML-based compression with Kompress (ModernBERT).
# (The legacy [llmlingua] extra was removed in 0.9.x — no live code path used it.
# Use [ml] for the supported ML compression dependencies.)
ml = [
# PyTorch does not publish wheels for macOS 15 x86_64 at this floor, which
# makes `headroom-ai[all]` unsatisfiable on Intel Macs (#1931).
"torch>=2.12.1; sys_platform != 'darwin' or platform_machine != 'x86_64'",
"transformers>=5.5.0,<6.0",
# transformers >= 5.x requires huggingface-hub >= 1.5.0,<2.0; pinning
# the floor here prevents Kompress from silently falling back to
# "unavailable" when a sibling install (e.g. `pip install
# strands-agents`) drags huggingface-hub backwards.
"huggingface-hub>=1.5.0,<2.0",
]
# Memory system (hierarchical memory with vector search).
# Uses the pure-Python sqlite-vec backend by default (VectorBackend.AUTO ->
# SQLITE_VEC), so no C++ toolchain is required. The optional HNSW backend lives
# in the [vector] extra below; installing it here would make `[all]` (which pulls
# [memory]) fail on any machine without a compiler — see #1368.
memory = [
"sqlite-vec>=0.1.6",
"sentence-transformers>=2.2.0,<6.0; sys_platform != 'darwin' or platform_machine != 'x86_64'",
]
# Optional HNSW vector backend. Needs a C++ toolchain to build hnswlib, so it is
# kept out of [memory] and [all]; opt in with `pip install headroom-ai[vector]`
# and select it via MemoryConfig(vector_backend=VectorBackend.HNSW). The default
# sqlite-vec backend needs no compiler.
vector = [
"hnswlib>=0.8.0",
]
# Qdrant + Neo4j memory backend helpers
memory-stack = [
"mem0ai>=2.0.0,<3.0",
"qdrant-client>=1.9.0,<2.0",
"neo4j>=5.20.0,<7.0",
]
# Apple-Silicon GPU (MPS) offload for the memory embedder. Opt in at runtime with
# HEADROOM_EMBEDDER_RUNTIME=pytorch_mps. macOS-only; intentionally excluded from [all].
pytorch-mps = [
"torch>=2.12.1; sys_platform == 'darwin'",
"sentence-transformers>=2.2.0; sys_platform == 'darwin'",
]
# Semantic relevance scoring with embeddings.
# Uses `fastembed` (BAAI/bge-small-en-v1.5 by default — 33M params,
# 384 dims, ~30 MB int8-quantized ONNX). Same library + model used by
# the Rust SmartCrusher (`fastembed` crate), giving byte-equal embeddings
# across the language boundary. Replaced sentence-transformers in
# Stage 3c.1 — fastembed is faster (~2-3x), smaller (no torch
# dependency), and outranks all-MiniLM-L6-v2 on MTEB by ~6 points.
relevance = [
"fastembed>=0.4.0",
"numpy>=1.24.0",
]
# Image compression (ML-based routing + OCR)
#
# OCR backend uses ONNX Runtime regardless of Python version. The
# rapidocr ecosystem split into two flavors after 1.4.x:
# * rapidocr-onnxruntime 1.4.x — bundled-ORT package, capped at
# Python <3.13 by its requires-python metadata. Drop-in for our
# existing v1 tuple-shaped API call.
# * rapidocr 3.x — engine-agnostic core, supports Python 3.13+.
# Returns a RapidOCROutput dataclass (txts, scores, boxes, ...).
# Needs `onnxruntime` installed separately to use the ORT backend.
#
# `headroom/image/compressor.py` adapts both API shapes at runtime via
# a try/except cascade. See issue #372 for context.
image = [
"pillow>=12.3.0", # PYSEC-2026-2253/2254/2255/2256/2257 fixes (decompression-bomb + cmd-injection)
"sentencepiece>=0.1.99", # Required by SigLIP tokenizer (SiglipTokenizer)
# Python 3.63.12: keep the proven ORT-bundled package directly.
# ~15 MB ONNX models auto-downloaded on first use.
"rapidocr-onnxruntime>=1.4.0,<2; python_version<'3.13'",
# Python 3.13+: rapidocr-onnxruntime is unavailable (its wheels
# declare requires-python<3.13). Use the successor `rapidocr` 3.x
# core + `onnxruntime` engine; same ORT backend, just split into
# two packages. Total install size and inference speed unchanged.
"rapidocr>=3.0,<4; python_version>='3.13'",
"onnxruntime>=1.7,<2; python_version>='3.13'",
]
# Report generation
reports = [
"jinja2>=3.0.0",
]
# Binary spreadsheet ingestion (.xlsx / .xls -> tabular text)
spreadsheet = [
"openpyxl>=3.1.0", # .xlsx
"xlrd>=2.0.1", # legacy .xls
]
# OpenTelemetry metrics export
otel = [
"opentelemetry-sdk>=1.24.0",
"opentelemetry-exporter-otlp-proto-http>=1.24.0",
]
# any-llm multi-provider backend (requires Python 3.11+)
anyllm = [
"any-llm-sdk>=1.0.0; python_version >= '3.11'",
]
# LangChain integration
langchain = [
"langchain-core>=1.3.3,<4.0",
"langchain-openai>=1.1.14,<2.0",
]
# Agno agent framework integration
agno = [
"agno>=1.0.0",
]
# AWS Strands Agents SDK integration
strands = [
"strands-agents>=0.1.0",
]
# CrewAI agent framework integration
crewai = [
"crewai>=1.0",
]
# AutoGen agent framework integration
autogen = [
"autogen-agentchat>=0.7",
]
# MCP server for Claude Code integration
mcp = [
"mcp>=1.28.1,<2.0.0",
"httpx>=0.24.0",
"starlette>=0.27.0",
"uvicorn>=0.23.0,<1.0",
]
# Voice filler detection
voice = [
"onnxruntime>=1.16.0",
"transformers>=5.5.0,<6.0",
"torch>=2.12.1; sys_platform != 'darwin' or platform_machine != 'x86_64'",
]
# Voice training (includes voice deps + training extras)
voice-train = [
"headroom-ai[voice]",
"datasets>=2.14.0",
"accelerate>=0.20.0",
]
# Evaluation framework
evals = [
"datasets>=2.14.0",
"sentence-transformers>=2.2.0,<6.0; sys_platform != 'darwin' or platform_machine != 'x86_64'",
"numpy>=1.24.0",
"scikit-learn>=1.3.0",
"anthropic>=0.18.0",
"openai>=1.0.0",
]
# AWS Bedrock backend
bedrock = [
# `aws login` (IAM Identity Provider / console-login, DPoP) requires
# boto3 >= 1.41.0 AND the AWS Common Runtime (CRT) per AWS docs
# ("Boto3 1.41.0 or later with CRT"). CRT is a separate install — pull it
# via the botocore [crt] extra (awscrt). Without it, resolving `aws login`
# credentials raises botocore's MissingDependencyException.
"boto3>=1.41.0",
"botocore[crt]>=1.41.0",
]
# HTML content extraction
html = [
"trafilatura>=1.6.0",
]
# Development dependencies
dev = [
"pytest>=7.0.0",
"pytest-cov>=4.0.0",
"pytest-asyncio>=0.21.0",
"ruff==0.15.22",
"mypy>=1.0.0",
"pre-commit>=3.0.0",
"openai>=1.0.0",
"anthropic>=0.18.0",
"litellm>=1.86.2,<2.0; python_version < '3.14'", # see core deps note (GH #956)
"fastapi>=0.100.0",
"uvicorn>=0.23.0,<1.0",
"httpx[http2]>=0.24.0",
"websockets>=13.0",
"opentelemetry-sdk>=1.24.0",
"opentelemetry-exporter-otlp-proto-http>=1.24.0",
"ollama>=0.4.0",
"langchain-ollama>=0.2.0",
"hnswlib>=0.8.0",
"sqlite-vec>=0.1.6",
"sentence-transformers>=2.2.0,<6.0",
"numpy>=1.24.0",
"openpyxl>=3.1.0", # exercises spreadsheet_ingest (.xlsx) in the test suite
"respx>=0.20.0", # HTTP mock transport for passthrough handler tests
]
# All optional dependencies (everything you need).
#
# The EleutherAI lm-evaluation-harness is intentionally not exposed as a
# project extra. Headroom invokes it as an external subprocess
# (`python -m lm_eval`), and the harness currently pulls sqlitedict
# CVE-2024-35515 with no upstream fix. Keeping it out of locked project extras
# prevents repository scanners from flagging production installs; researchers
# who need standard accuracy benchmarks can install `lm-eval[api]` in their
# benchmark environment separately.
all = [
"headroom-ai[proxy,code,ml,memory,relevance,image,reports,otel,evals,voice,html,mcp,spreadsheet]",
]
# Sandbox: a lean proxy with ALL torch-free capability — for running Headroom in
# a locked-down/low-resource sandbox and offloading heavy ML elsewhere.
#
# = [all] MINUS:
# - image (SigLIP/OCR — excluded by request)
# - ml (torch — the PyTorch Kompress backend; ONNX path in [proxy] still
# runs Kompress locally with no torch, or offload it entirely via
# HEADROOM_KOMPRESS_ENDPOINT)
# - voice (excluded by request)
# - memory + evals (both pull sentence-transformers -> torch, i.e. the very
# ML weight a sandbox avoids; evals is a dev/test harness, not a
# runtime feature). Opt back in explicitly if you accept torch:
# pip install headroom-ai[sandbox,memory]
#
# Everything kept here is torch-free: code-aware compression (tree-sitter),
# embedding relevance (fastembed), HTML/spreadsheet ingestion, reports, OTel.
sandbox = [
"headroom-ai[proxy,code,relevance,reports,otel,html,mcp,spreadsheet]",
]
[project.scripts]
headroom = "headroom.cli:main"
[project.urls]
Homepage = "https://headroom-docs.vercel.app"
Documentation = "https://headroom-docs.vercel.app/docs"
Repository = "https://github.com/chopratejas/headroom"
Issues = "https://github.com/chopratejas/headroom/issues"
Changelog = "https://github.com/chopratejas/headroom/blob/main/CHANGELOG.md"
# llms.txt convention (llmstxt.org) — point AI agents / LLM crawlers
# at the auto-generated docs index so they can resolve install paths
# and entry points without a follow-up fetch.
"AI / LLM Index" = "https://headroom-docs.vercel.app/llms.txt"
# Maturin builds a single wheel containing both the Python source under
# `headroom/` AND the compiled Rust extension `headroom/_core.so` (cdylib
# from `crates/headroom-py`). One `pip install headroom-ai` ships everything
# atomically — no separate `headroom-core-py` package, no chicken-and-egg,
# no PIP_FIND_LINKS plumbing. Phase A0's runtime fail-loud check still
# exists but only fires if someone forces an sdist install on a platform
# without a wheel and the rust toolchain isn't available to compile it.
# Constrain transitive dependencies that have CVEs requiring minimum versions.
# These packages don't appear as direct headroom deps but are pulled in
# transitively; the floor pins below ensure uv resolves to patched versions.
[tool.uv]
constraint-dependencies = [
# GHSA-5239-wwwm-4pmq (Low) — transitive via rich; fix at 2.20.0
"pygments>=2.20.0",
# GHSA-4xgf-cpjx-pc3j (Medium) — transitive via mcp; fix at 2.14.2
"pydantic-settings>=2.14.2",
# GHSA-mv93-w799-cj2w + 4 others (High) — transitive via agno; fix at 3.1.50
"gitpython>=3.1.50",
# GHSA-f4xh-w4cj-qxq8 (High) — transitive via langchain-core; fix at 0.8.18
"langsmith>=0.9.0",
# CVE-2026-49825 (High, XSS) — transitive via lxml[html-clean]; fix at 0.4.5
"lxml-html-clean>=0.4.5",
# CVE-2026-5241 (High) — direct optional dep for proxy/ml/voice; fix at 5.5.0
"transformers>=5.5.0",
# PYSEC-2026-3447 — transitive dependency; fix at 83.0.0
"setuptools>=83.0.0",
# PYSEC-2026-3545/3546/3547 — transitive HTTP/WebSocket parser fixes
"aiohttp>=3.14.3",
# PYSEC-2026-3552/3553/3554 — PKCS#7 and certificate verification fixes
"cryptography>=50.0.0",
]
# Pin the project's package index to public PyPI. Without this, `uv lock`
# inherits the developer's user-level `~/.config/uv/uv.toml` index
# setting — including private/internal mirrors like
# `pypi.netflix.net/simple` — and bakes those URLs into uv.lock, which
# then breaks CI on every public runner that can't reach the mirror.
# Declaring the index in pyproject.toml makes the project authoritative
# regardless of who runs `uv lock`.
[[tool.uv.index]]
name = "pypi"
url = "https://pypi.org/simple/"
default = true
[tool.maturin]
# Where the Python package lives. With `python-source = "."` and the
# package directory `headroom/` at repo root, maturin includes every file
# under `headroom/` in the wheel — that picks up the dashboard HTML
# templates and bundled YAML configs. `LICENSE` and `NOTICE` are listed
# explicitly because maturin sdists do not get the package-directory
# treatment wheels do, and PEP 639 auto-discovery emits both files into
# `License-File:` metadata — PyPI rejects sdists whose declared license
# files are missing from the tarball with `400 License-File X does not
# exist in distribution file`.
include = [
{ path = "LICENSE", format = "sdist" },
{ path = "NOTICE", format = "sdist" },
]
python-source = "."
module-name = "headroom._core"
# The cdylib source lives under `crates/headroom-py`. Maturin invokes
# `cargo build` with this manifest to produce `_core.cdylib`, then injects
# the resulting `.so` into the wheel at `headroom/_core.so`.
manifest-path = "crates/headroom-py/Cargo.toml"
features = ["extension-module"]
# Forbid building without the cdylib feature — bare `cargo build` won't
# produce a usable Python extension. Maturin's default `bindings` is "pyo3"
# which is correct here (see `crates/headroom-py/src/`).
bindings = "pyo3"
[tool.ruff]
target-version = "py310"
line-length = 100
[tool.ruff.lint]
select = [
"E", # pycodestyle errors
"W", # pycodestyle warnings
"F", # pyflakes
"I", # isort
"B", # flake8-bugbear
"C4", # flake8-comprehensions
"UP", # pyupgrade
]
ignore = [
"E501", # line too long (handled by formatter)
"B008", # do not perform function calls in argument defaults
"B905", # zip without strict parameter
]
[tool.ruff.lint.isort]
known-first-party = ["headroom"]
[tool.ruff.format]
quote-style = "double"
indent-style = "space"
[tool.mypy]
python_version = "3.10"
warn_return_any = true
warn_unused_configs = true
disallow_untyped_defs = true
ignore_missing_imports = true
# Per-module overrides for modules with dynamic typing patterns
[[tool.mypy.overrides]]
module = [
"headroom.proxy.server",
"headroom.proxy.cost",
"headroom.proxy.prometheus_metrics",
"headroom.proxy.semantic_cache",
"headroom.proxy.rate_limiter",
"headroom.proxy.request_logger",
"headroom.proxy.helpers",
"headroom.integrations.langchain",
"headroom.integrations.mcp",
"headroom.ccr.mcp_server",
"headroom.relevance.embedding",
"headroom.reporting.generator",
]
disallow_untyped_defs = false
[[tool.mypy.overrides]]
module = [
"headroom.tokenizers.*",
"headroom.providers.litellm",
"headroom.providers.google",
]
disallow_untyped_defs = false
warn_return_any = false
# Handler mixins use self.* from HeadroomProxy via duck typing — mypy can't resolve these
[[tool.mypy.overrides]]
module = ["headroom.proxy.handlers.*"]
disallow_untyped_defs = false
ignore_errors = true
# Ignore third-party stubs with syntax errors
[[tool.mypy.overrides]]
module = ["mlx.*"]
ignore_errors = true
[tool.pytest.ini_options]
testpaths = ["tests"]
python_files = ["test_*.py"]
python_functions = ["test_*"]
addopts = "-v --tb=short"
asyncio_mode = "auto"
filterwarnings = [
# pyo3 Unsendable parsers emit an unraisable warning when GC drops them on a
# test-teardown thread; this is a test-harness artifact, not a production issue
# (production threads are long-lived and drop their parsers on themselves).
"ignore::pytest.PytestUnraisableExceptionWarning",
]
markers = [
"slow: slow tests (model loads, large fixtures)",
"real_llm: tests that hit real LLM APIs; skipped unless explicitly enabled",
"live: opt-in multi-turn tests that hit real upstream APIs; require provider keys",
]
[tool.coverage.run]
source = ["headroom"]
branch = true
omit = [
"headroom/cli.py",
"*/tests/*",
]
[tool.coverage.report]
exclude_lines = [
"pragma: no cover",
"def __repr__",
"raise NotImplementedError",
"if TYPE_CHECKING:",
"if __name__ == .__main__.:",
]