fix(proxy): keep Kompress warmup off the startup path (#2001)
## Description
Proxy startup can enter cached Kompress native model initialization
before binding its port. On the RHEL/CentOS 7-family environment
reported in #1908, that path terminates in a deterministic
`libarrow.so.2400` jemalloc-thread segfault with no Python traceback.
Cache-only preload still initializes native libraries when the model is
already cached.
Defer Kompress model and tokenizer loading out of startup while
preserving the existing lazy request path and the eager warmups for
non-Kompress components. Closes #1908.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Stop cached and uncached Kompress models from entering native preload
during proxy startup.
- Record enabled Kompress as deferred until its existing lazy request
path needs it.
- Preserve disabled-Kompress routing and non-Kompress eager warmups.
- Add lifecycle, cache-state, disabled-mode, and warmup-preservation
regression coverage.
- Document the startup behavior change in the changelog.
## Testing
- [x] Unit tests pass (`uv run pytest
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py -q`)
- [x] Linting passes (`uv run ruff check
headroom/transforms/content_router.py
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py`)
- [ ] Type checking passes (`uv run mypy headroom`)
- [x] New tests added for new functionality when applicable
- [ ] Manual testing performed
### Test Output
```text
uv run pytest tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py -q
17 passed in 1.28s
uv run ruff check headroom/transforms/content_router.py tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py
All checks passed
```
## Real Behavior Proof
- Environment: Red OS 7.3 or equivalent RHEL/CentOS 7-family system,
glibc 2.17, Python 3.11, pyarrow 24.0.0, onnxruntime 1.27.0, cached
Kompress model
- Exact command / steps: start `headroom proxy` with Kompress enabled,
wait 30 seconds, query the loopback health endpoint, verify deferred
Kompress warmup in logs, then inspect `journalctl -k` for new
`libarrow.so.2400` or `jemalloc_bg_thd` faults
- Observed result: automated coverage now proves startup avoids the
cached Kompress preload boundary, keeps non-Kompress warmups live, and
preserves `unavailable` status when dependencies are absent
- Not tested: the native reporter-host run
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Screenshots
Not applicable.
## Additional Notes
Focused local validation included
`tests/test_kompress_request_nonblocking.py` alongside the startup
regression suite. The change is scoped to startup warmup; it does not
claim to repair the external `libarrow.so` or jemalloc incompatibility
when Kompress later executes. `HEADROOM_DISABLE_KOMPRESS=1` remains the
supported narrow workaround for hosts that cannot run the native path.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-13 10:09:45 -04:00
|
|
|
"""Startup eager-preload must defer Kompress native loading before binding.
|
|
|
|
|
|
|
|
|
|
Regression for the production crash where ``eager_load_compressors`` entered
|
|
|
|
|
the cached Kompress native stack on the blocking startup/lifespan path.
|
fix(proxy): make Kompress eager preload cache-only so a cold cache can't block startup (#783)
## Description
`ContentRouter.eager_load_compressors()` runs a network
`hf_hub_download` of the Kompress ONNX model on the **blocking
startup/lifespan path**, before the proxy binds its port. On a cold
cache this is unsafe:
- the download can hang long enough to blow the supervisor's bind
timeout, or
- a native crash in the download/ML stack (an **uncatchable `Fatal
Python error: Aborted` / SIGABRT**) kills the interpreter before it ever
`listen()`s.
Either way the supervisor sees "proxy never opened its port" and gives
up. We observed this in the field from the desktop app (process aborted
during `eager_load_compressors -> _load_kompress_onnx ->
hf_hub_download` of `onnx/kompress-int8.onnx`, while the only Python
thread was parked in the HuggingFace download file-lock; the abort came
from a native thread, so `try/except` at the call site cannot catch it).
The eager preload is a latency optimization and must never be able to
block — or kill — startup. This change makes startup preload
**cache-only**: if the model isn't already cached, we defer the download
to first use (off the startup path) and bind the port normally. Warm
starts are unchanged.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `onnx_runtime.hf_hub_download_local_first(...)`: added `allow_network`
(default `True`). When `False`, a cache miss re-raises the local-lookup
error instead of falling back to a network download.
- `kompress_compressor`: added `allow_download` (default `True`)
threaded through `preload()` -> `_load_kompress()` ->
`_load_kompress_onnx()` / `_load_kompress_pytorch()` and the ModernBERT
tokenizer load. Added `KompressModelNotCached`, raised when a cache-only
load misses. Auto-mode no longer falls back to a PyTorch network
download on a cache-only miss — it propagates so the caller can defer.
- `content_router.eager_load_compressors()`: calls
`preload(allow_download=False)`. On `KompressModelNotCached` it logs and
reports the component as `"deferred"` (a status
`warmup.merge_transform_status` already handles gracefully) instead of
letting a cold download run on the startup path.
Default (first-request) loading behavior and warm-start preload are
unchanged.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
New tests in `tests/test_kompress_preload_deferral.py` cover: cache-only
`hf_hub_download_local_first` never hits the network; default still
falls back; cache-only ONNX load raises `KompressModelNotCached`;
auto-mode does **not** trigger a PyTorch download on a cache-only miss;
and `eager_load_compressors` reports `deferred` (cold) / `enabled`
(warm). Existing `_load_kompress` dispatch tests updated for the new
keyword-only param.
> Note on environment: I do not have a clean reproduction of the native
SIGABRT itself (it depends on a specific machine's HF download/ML native
stack), so the "Manual testing performed" box is left unchecked. The
tests target the structural fix — that startup preload can no longer
perform a network download — which is the precondition for the crash.
## Test Output
```
$ uv run pytest -v tests/test_kompress_preload_deferral.py
tests/test_kompress_preload_deferral.py::test_local_first_no_network_when_disallowed PASSED
tests/test_kompress_preload_deferral.py::test_local_first_falls_back_to_network_by_default PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_onnx_cache_miss_raises_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_auto_does_not_pytorch_download_on_cache_miss PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_defers_when_model_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_enabled_when_model_cached PASSED
6 passed in 4.82s
$ uv run pytest tests/test_transforms/test_kompress_compressor.py tests/test_transforms_content_router.py tests/test_onnx_runtime.py tests/test_proxy_warmup.py
63 passed
$ uv run ruff check <changed files> # All checks passed!
$ uv run mypy headroom/onnx_runtime.py headroom/transforms/kompress_compressor.py headroom/transforms/content_router.py
Success: no issues found
```
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (auto-generated from
conventional commits)
## Additional Notes
This contains the cold-start case. A native crash in onnxruntime
*session init* (as opposed to the download) on first request would still
be a separate issue; it is not what was observed here (the abort was
during the HF download), and isolating it would be a larger, separate
change.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 19:53:03 +02:00
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
from __future__ import annotations
|
|
|
|
|
|
|
|
|
|
import pytest
|
|
|
|
|
|
|
|
|
|
from headroom import onnx_runtime
|
|
|
|
|
from headroom.transforms import kompress_compressor as kc
|
|
|
|
|
from headroom.transforms.content_router import ContentRouter, ContentRouterConfig
|
|
|
|
|
from headroom.transforms.kompress_compressor import KompressModelNotCached
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_local_first_no_network_when_disallowed(monkeypatch):
|
|
|
|
|
"""allow_network=False must never fall back to a network download."""
|
|
|
|
|
import huggingface_hub
|
|
|
|
|
from huggingface_hub.errors import LocalEntryNotFoundError
|
|
|
|
|
|
|
|
|
|
calls: list[bool] = []
|
|
|
|
|
|
|
|
|
|
def fake_download(repo_id, filename, **kwargs):
|
|
|
|
|
local_only = kwargs.get("local_files_only", False)
|
|
|
|
|
calls.append(local_only)
|
|
|
|
|
if local_only:
|
|
|
|
|
raise LocalEntryNotFoundError("not cached")
|
|
|
|
|
return "/cache/networked"
|
|
|
|
|
|
|
|
|
|
monkeypatch.setattr(huggingface_hub, "hf_hub_download", fake_download)
|
|
|
|
|
|
|
|
|
|
with pytest.raises(LocalEntryNotFoundError):
|
|
|
|
|
onnx_runtime.hf_hub_download_local_first("org/model", "f.onnx", allow_network=False)
|
|
|
|
|
|
|
|
|
|
# Only the local-only lookup ran; the network branch was never taken.
|
|
|
|
|
assert calls == [True]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_local_first_falls_back_to_network_by_default(monkeypatch):
|
|
|
|
|
"""allow_network=True (default) keeps the historic cold-start behavior."""
|
|
|
|
|
import huggingface_hub
|
|
|
|
|
from huggingface_hub.errors import LocalEntryNotFoundError
|
|
|
|
|
|
|
|
|
|
calls: list[bool] = []
|
|
|
|
|
|
|
|
|
|
def fake_download(repo_id, filename, **kwargs):
|
|
|
|
|
local_only = kwargs.get("local_files_only", False)
|
|
|
|
|
calls.append(local_only)
|
|
|
|
|
if local_only:
|
|
|
|
|
raise LocalEntryNotFoundError("not cached")
|
|
|
|
|
return "/cache/networked"
|
|
|
|
|
|
|
|
|
|
monkeypatch.setattr(huggingface_hub, "hf_hub_download", fake_download)
|
|
|
|
|
|
|
|
|
|
path = onnx_runtime.hf_hub_download_local_first("org/model", "f.onnx")
|
|
|
|
|
assert path == "/cache/networked"
|
|
|
|
|
assert calls == [True, False] # local-only miss, then network download
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_load_kompress_onnx_cache_miss_raises_not_cached(monkeypatch):
|
|
|
|
|
"""A cache-only ONNX load surfaces KompressModelNotCached, not a network call."""
|
|
|
|
|
from huggingface_hub.errors import LocalEntryNotFoundError
|
|
|
|
|
|
|
|
|
|
monkeypatch.setattr(kc, "_kompress_cache", {})
|
|
|
|
|
|
|
|
|
|
def fake_local_first(repo_id, filename, *, allow_network=True):
|
|
|
|
|
assert allow_network is False # eager preload must request cache-only
|
|
|
|
|
raise LocalEntryNotFoundError("not cached")
|
|
|
|
|
|
|
|
|
|
monkeypatch.setattr(kc, "hf_hub_download_local_first", fake_local_first)
|
|
|
|
|
|
|
|
|
|
with pytest.raises(KompressModelNotCached):
|
|
|
|
|
kc._load_kompress_onnx("org/model", allow_download=False)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_load_kompress_auto_does_not_pytorch_download_on_cache_miss(monkeypatch):
|
|
|
|
|
"""Auto mode must propagate the cache miss, not fall back to a PyTorch fetch."""
|
|
|
|
|
monkeypatch.setattr(kc, "_kompress_cache", {})
|
|
|
|
|
monkeypatch.setattr(kc, "_selected_backend", lambda: "auto")
|
|
|
|
|
monkeypatch.setattr(kc, "_is_onnx_available", lambda: True)
|
|
|
|
|
monkeypatch.setattr(kc, "_is_pytorch_available", lambda: True)
|
|
|
|
|
|
|
|
|
|
def onnx_not_cached(model_id, *, use_coreml=False, allow_download=True):
|
|
|
|
|
raise KompressModelNotCached(model_id)
|
|
|
|
|
|
|
|
|
|
def pytorch_should_not_run(*args, **kwargs):
|
|
|
|
|
raise AssertionError("PyTorch fallback must not download on a cache-only miss")
|
|
|
|
|
|
|
|
|
|
monkeypatch.setattr(kc, "_load_kompress_onnx", onnx_not_cached)
|
|
|
|
|
monkeypatch.setattr(kc, "_load_kompress_pytorch", pytorch_should_not_run)
|
|
|
|
|
|
|
|
|
|
with pytest.raises(KompressModelNotCached):
|
|
|
|
|
kc._load_kompress("org/model", allow_download=False)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
class _StubCompressor:
|
|
|
|
|
def __init__(self, *, cached: bool):
|
|
|
|
|
self._cached = cached
|
|
|
|
|
self.preload_calls: list[bool] = []
|
|
|
|
|
|
|
|
|
|
def preload(self, *, allow_download: bool = True) -> str:
|
|
|
|
|
self.preload_calls.append(allow_download)
|
|
|
|
|
if self._cached:
|
|
|
|
|
return "onnx"
|
|
|
|
|
raise KompressModelNotCached("org/model")
|
|
|
|
|
|
|
|
|
|
|
fix(proxy): keep Kompress warmup off the startup path (#2001)
## Description
Proxy startup can enter cached Kompress native model initialization
before binding its port. On the RHEL/CentOS 7-family environment
reported in #1908, that path terminates in a deterministic
`libarrow.so.2400` jemalloc-thread segfault with no Python traceback.
Cache-only preload still initializes native libraries when the model is
already cached.
Defer Kompress model and tokenizer loading out of startup while
preserving the existing lazy request path and the eager warmups for
non-Kompress components. Closes #1908.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Stop cached and uncached Kompress models from entering native preload
during proxy startup.
- Record enabled Kompress as deferred until its existing lazy request
path needs it.
- Preserve disabled-Kompress routing and non-Kompress eager warmups.
- Add lifecycle, cache-state, disabled-mode, and warmup-preservation
regression coverage.
- Document the startup behavior change in the changelog.
## Testing
- [x] Unit tests pass (`uv run pytest
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py -q`)
- [x] Linting passes (`uv run ruff check
headroom/transforms/content_router.py
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py`)
- [ ] Type checking passes (`uv run mypy headroom`)
- [x] New tests added for new functionality when applicable
- [ ] Manual testing performed
### Test Output
```text
uv run pytest tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py -q
17 passed in 1.28s
uv run ruff check headroom/transforms/content_router.py tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py
All checks passed
```
## Real Behavior Proof
- Environment: Red OS 7.3 or equivalent RHEL/CentOS 7-family system,
glibc 2.17, Python 3.11, pyarrow 24.0.0, onnxruntime 1.27.0, cached
Kompress model
- Exact command / steps: start `headroom proxy` with Kompress enabled,
wait 30 seconds, query the loopback health endpoint, verify deferred
Kompress warmup in logs, then inspect `journalctl -k` for new
`libarrow.so.2400` or `jemalloc_bg_thd` faults
- Observed result: automated coverage now proves startup avoids the
cached Kompress preload boundary, keeps non-Kompress warmups live, and
preserves `unavailable` status when dependencies are absent
- Not tested: the native reporter-host run
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Screenshots
Not applicable.
## Additional Notes
Focused local validation included
`tests/test_kompress_request_nonblocking.py` alongside the startup
regression suite. The change is scoped to startup warmup; it does not
claim to repair the external `libarrow.so` or jemalloc incompatibility
when Kompress later executes. `HEADROOM_DISABLE_KOMPRESS=1` remains the
supported narrow workaround for hosts that cannot run the native path.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-13 10:09:45 -04:00
|
|
|
class _FatalPreloadCompressor(_StubCompressor):
|
|
|
|
|
def preload(self, *, allow_download: bool = True) -> str:
|
|
|
|
|
self.preload_calls.append(allow_download)
|
|
|
|
|
raise SystemExit("native Kompress preload")
|
|
|
|
|
|
|
|
|
|
|
fix(proxy): make Kompress eager preload cache-only so a cold cache can't block startup (#783)
## Description
`ContentRouter.eager_load_compressors()` runs a network
`hf_hub_download` of the Kompress ONNX model on the **blocking
startup/lifespan path**, before the proxy binds its port. On a cold
cache this is unsafe:
- the download can hang long enough to blow the supervisor's bind
timeout, or
- a native crash in the download/ML stack (an **uncatchable `Fatal
Python error: Aborted` / SIGABRT**) kills the interpreter before it ever
`listen()`s.
Either way the supervisor sees "proxy never opened its port" and gives
up. We observed this in the field from the desktop app (process aborted
during `eager_load_compressors -> _load_kompress_onnx ->
hf_hub_download` of `onnx/kompress-int8.onnx`, while the only Python
thread was parked in the HuggingFace download file-lock; the abort came
from a native thread, so `try/except` at the call site cannot catch it).
The eager preload is a latency optimization and must never be able to
block — or kill — startup. This change makes startup preload
**cache-only**: if the model isn't already cached, we defer the download
to first use (off the startup path) and bind the port normally. Warm
starts are unchanged.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `onnx_runtime.hf_hub_download_local_first(...)`: added `allow_network`
(default `True`). When `False`, a cache miss re-raises the local-lookup
error instead of falling back to a network download.
- `kompress_compressor`: added `allow_download` (default `True`)
threaded through `preload()` -> `_load_kompress()` ->
`_load_kompress_onnx()` / `_load_kompress_pytorch()` and the ModernBERT
tokenizer load. Added `KompressModelNotCached`, raised when a cache-only
load misses. Auto-mode no longer falls back to a PyTorch network
download on a cache-only miss — it propagates so the caller can defer.
- `content_router.eager_load_compressors()`: calls
`preload(allow_download=False)`. On `KompressModelNotCached` it logs and
reports the component as `"deferred"` (a status
`warmup.merge_transform_status` already handles gracefully) instead of
letting a cold download run on the startup path.
Default (first-request) loading behavior and warm-start preload are
unchanged.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
New tests in `tests/test_kompress_preload_deferral.py` cover: cache-only
`hf_hub_download_local_first` never hits the network; default still
falls back; cache-only ONNX load raises `KompressModelNotCached`;
auto-mode does **not** trigger a PyTorch download on a cache-only miss;
and `eager_load_compressors` reports `deferred` (cold) / `enabled`
(warm). Existing `_load_kompress` dispatch tests updated for the new
keyword-only param.
> Note on environment: I do not have a clean reproduction of the native
SIGABRT itself (it depends on a specific machine's HF download/ML native
stack), so the "Manual testing performed" box is left unchecked. The
tests target the structural fix — that startup preload can no longer
perform a network download — which is the precondition for the crash.
## Test Output
```
$ uv run pytest -v tests/test_kompress_preload_deferral.py
tests/test_kompress_preload_deferral.py::test_local_first_no_network_when_disallowed PASSED
tests/test_kompress_preload_deferral.py::test_local_first_falls_back_to_network_by_default PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_onnx_cache_miss_raises_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_auto_does_not_pytorch_download_on_cache_miss PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_defers_when_model_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_enabled_when_model_cached PASSED
6 passed in 4.82s
$ uv run pytest tests/test_transforms/test_kompress_compressor.py tests/test_transforms_content_router.py tests/test_onnx_runtime.py tests/test_proxy_warmup.py
63 passed
$ uv run ruff check <changed files> # All checks passed!
$ uv run mypy headroom/onnx_runtime.py headroom/transforms/kompress_compressor.py headroom/transforms/content_router.py
Success: no issues found
```
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (auto-generated from
conventional commits)
## Additional Notes
This contains the cold-start case. A native crash in onnxruntime
*session init* (as opposed to the download) on first request would still
be a separate issue; it is not what was observed here (the abort was
during the HF download), and isolating it would be a larger, separate
change.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 19:53:03 +02:00
|
|
|
def _router_kompress_only() -> ContentRouter:
|
|
|
|
|
return ContentRouter(
|
|
|
|
|
ContentRouterConfig(
|
|
|
|
|
enable_kompress=True,
|
|
|
|
|
enable_code_aware=False,
|
|
|
|
|
enable_smart_crusher=False,
|
|
|
|
|
)
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
|
fix(proxy): keep Kompress warmup off the startup path (#2001)
## Description
Proxy startup can enter cached Kompress native model initialization
before binding its port. On the RHEL/CentOS 7-family environment
reported in #1908, that path terminates in a deterministic
`libarrow.so.2400` jemalloc-thread segfault with no Python traceback.
Cache-only preload still initializes native libraries when the model is
already cached.
Defer Kompress model and tokenizer loading out of startup while
preserving the existing lazy request path and the eager warmups for
non-Kompress components. Closes #1908.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Stop cached and uncached Kompress models from entering native preload
during proxy startup.
- Record enabled Kompress as deferred until its existing lazy request
path needs it.
- Preserve disabled-Kompress routing and non-Kompress eager warmups.
- Add lifecycle, cache-state, disabled-mode, and warmup-preservation
regression coverage.
- Document the startup behavior change in the changelog.
## Testing
- [x] Unit tests pass (`uv run pytest
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py -q`)
- [x] Linting passes (`uv run ruff check
headroom/transforms/content_router.py
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py`)
- [ ] Type checking passes (`uv run mypy headroom`)
- [x] New tests added for new functionality when applicable
- [ ] Manual testing performed
### Test Output
```text
uv run pytest tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py -q
17 passed in 1.28s
uv run ruff check headroom/transforms/content_router.py tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py
All checks passed
```
## Real Behavior Proof
- Environment: Red OS 7.3 or equivalent RHEL/CentOS 7-family system,
glibc 2.17, Python 3.11, pyarrow 24.0.0, onnxruntime 1.27.0, cached
Kompress model
- Exact command / steps: start `headroom proxy` with Kompress enabled,
wait 30 seconds, query the loopback health endpoint, verify deferred
Kompress warmup in logs, then inspect `journalctl -k` for new
`libarrow.so.2400` or `jemalloc_bg_thd` faults
- Observed result: automated coverage now proves startup avoids the
cached Kompress preload boundary, keeps non-Kompress warmups live, and
preserves `unavailable` status when dependencies are absent
- Not tested: the native reporter-host run
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Screenshots
Not applicable.
## Additional Notes
Focused local validation included
`tests/test_kompress_request_nonblocking.py` alongside the startup
regression suite. The change is scoped to startup warmup; it does not
claim to repair the external `libarrow.so` or jemalloc incompatibility
when Kompress later executes. `HEADROOM_DISABLE_KOMPRESS=1` remains the
supported narrow workaround for hosts that cannot run the native path.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-13 10:09:45 -04:00
|
|
|
@pytest.mark.parametrize("cache_state", ["cached", "uncached"])
|
|
|
|
|
def test_eager_load_defers_kompress_regardless_of_cache_state(monkeypatch, cache_state):
|
fix(proxy): make Kompress eager preload cache-only so a cold cache can't block startup (#783)
## Description
`ContentRouter.eager_load_compressors()` runs a network
`hf_hub_download` of the Kompress ONNX model on the **blocking
startup/lifespan path**, before the proxy binds its port. On a cold
cache this is unsafe:
- the download can hang long enough to blow the supervisor's bind
timeout, or
- a native crash in the download/ML stack (an **uncatchable `Fatal
Python error: Aborted` / SIGABRT**) kills the interpreter before it ever
`listen()`s.
Either way the supervisor sees "proxy never opened its port" and gives
up. We observed this in the field from the desktop app (process aborted
during `eager_load_compressors -> _load_kompress_onnx ->
hf_hub_download` of `onnx/kompress-int8.onnx`, while the only Python
thread was parked in the HuggingFace download file-lock; the abort came
from a native thread, so `try/except` at the call site cannot catch it).
The eager preload is a latency optimization and must never be able to
block — or kill — startup. This change makes startup preload
**cache-only**: if the model isn't already cached, we defer the download
to first use (off the startup path) and bind the port normally. Warm
starts are unchanged.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `onnx_runtime.hf_hub_download_local_first(...)`: added `allow_network`
(default `True`). When `False`, a cache miss re-raises the local-lookup
error instead of falling back to a network download.
- `kompress_compressor`: added `allow_download` (default `True`)
threaded through `preload()` -> `_load_kompress()` ->
`_load_kompress_onnx()` / `_load_kompress_pytorch()` and the ModernBERT
tokenizer load. Added `KompressModelNotCached`, raised when a cache-only
load misses. Auto-mode no longer falls back to a PyTorch network
download on a cache-only miss — it propagates so the caller can defer.
- `content_router.eager_load_compressors()`: calls
`preload(allow_download=False)`. On `KompressModelNotCached` it logs and
reports the component as `"deferred"` (a status
`warmup.merge_transform_status` already handles gracefully) instead of
letting a cold download run on the startup path.
Default (first-request) loading behavior and warm-start preload are
unchanged.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
New tests in `tests/test_kompress_preload_deferral.py` cover: cache-only
`hf_hub_download_local_first` never hits the network; default still
falls back; cache-only ONNX load raises `KompressModelNotCached`;
auto-mode does **not** trigger a PyTorch download on a cache-only miss;
and `eager_load_compressors` reports `deferred` (cold) / `enabled`
(warm). Existing `_load_kompress` dispatch tests updated for the new
keyword-only param.
> Note on environment: I do not have a clean reproduction of the native
SIGABRT itself (it depends on a specific machine's HF download/ML native
stack), so the "Manual testing performed" box is left unchecked. The
tests target the structural fix — that startup preload can no longer
perform a network download — which is the precondition for the crash.
## Test Output
```
$ uv run pytest -v tests/test_kompress_preload_deferral.py
tests/test_kompress_preload_deferral.py::test_local_first_no_network_when_disallowed PASSED
tests/test_kompress_preload_deferral.py::test_local_first_falls_back_to_network_by_default PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_onnx_cache_miss_raises_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_auto_does_not_pytorch_download_on_cache_miss PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_defers_when_model_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_enabled_when_model_cached PASSED
6 passed in 4.82s
$ uv run pytest tests/test_transforms/test_kompress_compressor.py tests/test_transforms_content_router.py tests/test_onnx_runtime.py tests/test_proxy_warmup.py
63 passed
$ uv run ruff check <changed files> # All checks passed!
$ uv run mypy headroom/onnx_runtime.py headroom/transforms/kompress_compressor.py headroom/transforms/content_router.py
Success: no issues found
```
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (auto-generated from
conventional commits)
## Additional Notes
This contains the cold-start case. A native crash in onnxruntime
*session init* (as opposed to the download) on first request would still
be a separate issue; it is not what was observed here (the abort was
during the HF download), and isolating it would be a larger, separate
change.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 19:53:03 +02:00
|
|
|
router = _router_kompress_only()
|
fix(proxy): keep Kompress warmup off the startup path (#2001)
## Description
Proxy startup can enter cached Kompress native model initialization
before binding its port. On the RHEL/CentOS 7-family environment
reported in #1908, that path terminates in a deterministic
`libarrow.so.2400` jemalloc-thread segfault with no Python traceback.
Cache-only preload still initializes native libraries when the model is
already cached.
Defer Kompress model and tokenizer loading out of startup while
preserving the existing lazy request path and the eager warmups for
non-Kompress components. Closes #1908.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Stop cached and uncached Kompress models from entering native preload
during proxy startup.
- Record enabled Kompress as deferred until its existing lazy request
path needs it.
- Preserve disabled-Kompress routing and non-Kompress eager warmups.
- Add lifecycle, cache-state, disabled-mode, and warmup-preservation
regression coverage.
- Document the startup behavior change in the changelog.
## Testing
- [x] Unit tests pass (`uv run pytest
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py -q`)
- [x] Linting passes (`uv run ruff check
headroom/transforms/content_router.py
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py`)
- [ ] Type checking passes (`uv run mypy headroom`)
- [x] New tests added for new functionality when applicable
- [ ] Manual testing performed
### Test Output
```text
uv run pytest tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py -q
17 passed in 1.28s
uv run ruff check headroom/transforms/content_router.py tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py
All checks passed
```
## Real Behavior Proof
- Environment: Red OS 7.3 or equivalent RHEL/CentOS 7-family system,
glibc 2.17, Python 3.11, pyarrow 24.0.0, onnxruntime 1.27.0, cached
Kompress model
- Exact command / steps: start `headroom proxy` with Kompress enabled,
wait 30 seconds, query the loopback health endpoint, verify deferred
Kompress warmup in logs, then inspect `journalctl -k` for new
`libarrow.so.2400` or `jemalloc_bg_thd` faults
- Observed result: automated coverage now proves startup avoids the
cached Kompress preload boundary, keeps non-Kompress warmups live, and
preserves `unavailable` status when dependencies are absent
- Not tested: the native reporter-host run
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Screenshots
Not applicable.
## Additional Notes
Focused local validation included
`tests/test_kompress_request_nonblocking.py` alongside the startup
regression suite. The change is scoped to startup warmup; it does not
claim to repair the external `libarrow.so` or jemalloc incompatibility
when Kompress later executes. `HEADROOM_DISABLE_KOMPRESS=1` remains the
supported narrow workaround for hosts that cannot run the native path.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-13 10:09:45 -04:00
|
|
|
stub = _StubCompressor(cached=cache_state == "cached")
|
fix(proxy): make Kompress eager preload cache-only so a cold cache can't block startup (#783)
## Description
`ContentRouter.eager_load_compressors()` runs a network
`hf_hub_download` of the Kompress ONNX model on the **blocking
startup/lifespan path**, before the proxy binds its port. On a cold
cache this is unsafe:
- the download can hang long enough to blow the supervisor's bind
timeout, or
- a native crash in the download/ML stack (an **uncatchable `Fatal
Python error: Aborted` / SIGABRT**) kills the interpreter before it ever
`listen()`s.
Either way the supervisor sees "proxy never opened its port" and gives
up. We observed this in the field from the desktop app (process aborted
during `eager_load_compressors -> _load_kompress_onnx ->
hf_hub_download` of `onnx/kompress-int8.onnx`, while the only Python
thread was parked in the HuggingFace download file-lock; the abort came
from a native thread, so `try/except` at the call site cannot catch it).
The eager preload is a latency optimization and must never be able to
block — or kill — startup. This change makes startup preload
**cache-only**: if the model isn't already cached, we defer the download
to first use (off the startup path) and bind the port normally. Warm
starts are unchanged.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `onnx_runtime.hf_hub_download_local_first(...)`: added `allow_network`
(default `True`). When `False`, a cache miss re-raises the local-lookup
error instead of falling back to a network download.
- `kompress_compressor`: added `allow_download` (default `True`)
threaded through `preload()` -> `_load_kompress()` ->
`_load_kompress_onnx()` / `_load_kompress_pytorch()` and the ModernBERT
tokenizer load. Added `KompressModelNotCached`, raised when a cache-only
load misses. Auto-mode no longer falls back to a PyTorch network
download on a cache-only miss — it propagates so the caller can defer.
- `content_router.eager_load_compressors()`: calls
`preload(allow_download=False)`. On `KompressModelNotCached` it logs and
reports the component as `"deferred"` (a status
`warmup.merge_transform_status` already handles gracefully) instead of
letting a cold download run on the startup path.
Default (first-request) loading behavior and warm-start preload are
unchanged.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
New tests in `tests/test_kompress_preload_deferral.py` cover: cache-only
`hf_hub_download_local_first` never hits the network; default still
falls back; cache-only ONNX load raises `KompressModelNotCached`;
auto-mode does **not** trigger a PyTorch download on a cache-only miss;
and `eager_load_compressors` reports `deferred` (cold) / `enabled`
(warm). Existing `_load_kompress` dispatch tests updated for the new
keyword-only param.
> Note on environment: I do not have a clean reproduction of the native
SIGABRT itself (it depends on a specific machine's HF download/ML native
stack), so the "Manual testing performed" box is left unchecked. The
tests target the structural fix — that startup preload can no longer
perform a network download — which is the precondition for the crash.
## Test Output
```
$ uv run pytest -v tests/test_kompress_preload_deferral.py
tests/test_kompress_preload_deferral.py::test_local_first_no_network_when_disallowed PASSED
tests/test_kompress_preload_deferral.py::test_local_first_falls_back_to_network_by_default PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_onnx_cache_miss_raises_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_auto_does_not_pytorch_download_on_cache_miss PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_defers_when_model_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_enabled_when_model_cached PASSED
6 passed in 4.82s
$ uv run pytest tests/test_transforms/test_kompress_compressor.py tests/test_transforms_content_router.py tests/test_onnx_runtime.py tests/test_proxy_warmup.py
63 passed
$ uv run ruff check <changed files> # All checks passed!
$ uv run mypy headroom/onnx_runtime.py headroom/transforms/kompress_compressor.py headroom/transforms/content_router.py
Success: no issues found
```
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (auto-generated from
conventional commits)
## Additional Notes
This contains the cold-start case. A native crash in onnxruntime
*session init* (as opposed to the download) on first request would still
be a separate issue; it is not what was observed here (the abort was
during the HF download), and isolating it would be a larger, separate
change.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 19:53:03 +02:00
|
|
|
monkeypatch.setattr(router, "_get_kompress", lambda: stub)
|
fix(kompress): reject artifacts that fail at run, and prefetch model files at startup (#2740)
## Description
Three cold-start / robustness gaps found while debugging a user report
of **0.12% savings across 722 requests** (49.8M input tokens, 60,920
saved).
Closes #
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
## Changes Made
### 1. The artifact fallback was unreachable for run-time failures
`_create_onnx_session` tries `int8-wo` → `fp32` → `int8`, and its
docstring describes exactly this scenario — but it only skipped a
candidate when `InferenceSession(...)` **construction** threw.
The int8 weight-only artifact carries `MatMulNBits` with `bits=8`. ORT's
CPU kernel only handles 8-bit through the prepacked MLAS path, so a
build or ISA without an 8-bit `SQNBitGemm` kernel falls into
`ComputeBUnpacked`, which hard-asserts `nbits_ == 4`. That raises on
`session.run()` **after** construction succeeded — so the fp32 candidate
was never reached and ML compression was dead for the process lifetime.
The reported log has 207 consecutive failures over three days.
A two-token `_smoke_run` inside the existing candidate loop makes the
fallback fire. `onnxruntime>=1.16.0` is unpinned, so which side of this
an install lands on is a lottery.
### 2. A broken model cost an inference on every request, forever
The per-request handler logged a `WARNING` and passed through with no
latch — 207 identical lines that read as noise rather than "ML
compression is dead". Now latches to passthrough after **3 consecutive**
failures (any success resets the count) with one actionable `ERROR`
naming the artifact override.
### 3. The model download began on the first request, not at startup
#2001 was right to move Kompress off the startup path — on RHEL/CentOS
7-family hosts, entering cached native init before the port binds
segfaults in `libarrow`/jemalloc with no Python traceback (#1908), which
no `try/except` can catch. **This PR does not touch that.**
But #2001 left the ~4-minute *download* on the first request, with every
request in that window silently uncompressed behind one "model not
ready" warning.
Downloading is separable from loading. `prefetch_kompress_artifacts`
resolves the files over plain `huggingface_hub` HTTP and never
constructs an `InferenceSession` or imports `transformers`, so startup
can prefetch bytes without touching the boundary #1908 crashes on.
Native load stays deferred, status stays `deferred`, and a test asserts
no session is constructed during prefetch.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ .venv/bin/ruff check headroom/ tests/test_kompress_failsafe.py tests/test_kompress_preload_deferral.py --exclude headroom/dashboard/templates
All checks passed!
$ .venv/bin/mypy headroom/
Success: no issues found in 508 source files
$ python -m pytest tests/test_kompress_failsafe.py tests/test_kompress_preload_deferral.py \
tests/test_kompress_request_nonblocking.py tests/test_force_kompress_all.py \
tests/test_kompress_must_keep.py tests/test_proxy_disable_kompress.py \
tests/test_proxy_per_provider_kompress.py tests/test_proxy_warmup.py \
tests/test_proxy_eager_preload_bind.py -q
95 passed in 10.51s
```
## Real Behavior Proof
- **Environment:** macOS 26.4 arm64, Python 3.12.6, onnxruntime 1.21.1,
repo `.venv`.
**(1) Fallback chain, against the real HF repo:**
```text
WARNING ONNX artifact 'onnx/kompress-int8-wo.onnx' from chopratejas/kompress-v2-base
is unusable (... nbits_ == 4 was false ...); trying next candidate
SESSION OK -> ['input_ids', 'attention_mask']
SMOKE RUN OK on the selected artifact
```
Also confirmed the default artifact really is 8-bit, by loading the
cached blob: `{'bits': [8], 'block_size': [128]}`.
**(2) Files-only prefetch, with `InferenceSession` patched to raise:**
```text
INFO Kompress: prefetching model artifacts for chopratejas/kompress-v2-base ...
prefetch ok=True in 0.08s, no session constructed
```
- **Not tested / important caveat:** the user's exact failure **cannot
be reproduced on this machine**. On ORT 1.21.1 arm64 the int8-wo
artifact fails at *construction* (`matmul_nbits.cc:115`), which the
pre-existing load-only fallback already caught. Their build fails at
*execution* (`matmul_nbits.cc:442`, `ComputeBUnpacked`). So the run-time
path is pinned with a fake ORT session that constructs fine and then
rejects `run()` — a mechanism test, not a reproduction of their build.
Confirming the fix on their host needs their `onnxruntime` version.
- **Not tested:** no RHEL/CentOS 7 host available to re-verify #1908
non-regression; the argument is structural (prefetch never constructs a
session) and asserted by
`test_prefetch_never_constructs_a_session_or_imports_transformers`.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md`
🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-03 10:42:48 -07:00
|
|
|
# Artifact prefetch is files-only, but it still reaches the network — keep it
|
|
|
|
|
# out of this assertion so the test stays about the native-preload boundary.
|
|
|
|
|
monkeypatch.setattr(router, "_prefetch_kompress_artifacts_async", lambda _cfg: False)
|
fix(proxy): make Kompress eager preload cache-only so a cold cache can't block startup (#783)
## Description
`ContentRouter.eager_load_compressors()` runs a network
`hf_hub_download` of the Kompress ONNX model on the **blocking
startup/lifespan path**, before the proxy binds its port. On a cold
cache this is unsafe:
- the download can hang long enough to blow the supervisor's bind
timeout, or
- a native crash in the download/ML stack (an **uncatchable `Fatal
Python error: Aborted` / SIGABRT**) kills the interpreter before it ever
`listen()`s.
Either way the supervisor sees "proxy never opened its port" and gives
up. We observed this in the field from the desktop app (process aborted
during `eager_load_compressors -> _load_kompress_onnx ->
hf_hub_download` of `onnx/kompress-int8.onnx`, while the only Python
thread was parked in the HuggingFace download file-lock; the abort came
from a native thread, so `try/except` at the call site cannot catch it).
The eager preload is a latency optimization and must never be able to
block — or kill — startup. This change makes startup preload
**cache-only**: if the model isn't already cached, we defer the download
to first use (off the startup path) and bind the port normally. Warm
starts are unchanged.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `onnx_runtime.hf_hub_download_local_first(...)`: added `allow_network`
(default `True`). When `False`, a cache miss re-raises the local-lookup
error instead of falling back to a network download.
- `kompress_compressor`: added `allow_download` (default `True`)
threaded through `preload()` -> `_load_kompress()` ->
`_load_kompress_onnx()` / `_load_kompress_pytorch()` and the ModernBERT
tokenizer load. Added `KompressModelNotCached`, raised when a cache-only
load misses. Auto-mode no longer falls back to a PyTorch network
download on a cache-only miss — it propagates so the caller can defer.
- `content_router.eager_load_compressors()`: calls
`preload(allow_download=False)`. On `KompressModelNotCached` it logs and
reports the component as `"deferred"` (a status
`warmup.merge_transform_status` already handles gracefully) instead of
letting a cold download run on the startup path.
Default (first-request) loading behavior and warm-start preload are
unchanged.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
New tests in `tests/test_kompress_preload_deferral.py` cover: cache-only
`hf_hub_download_local_first` never hits the network; default still
falls back; cache-only ONNX load raises `KompressModelNotCached`;
auto-mode does **not** trigger a PyTorch download on a cache-only miss;
and `eager_load_compressors` reports `deferred` (cold) / `enabled`
(warm). Existing `_load_kompress` dispatch tests updated for the new
keyword-only param.
> Note on environment: I do not have a clean reproduction of the native
SIGABRT itself (it depends on a specific machine's HF download/ML native
stack), so the "Manual testing performed" box is left unchecked. The
tests target the structural fix — that startup preload can no longer
perform a network download — which is the precondition for the crash.
## Test Output
```
$ uv run pytest -v tests/test_kompress_preload_deferral.py
tests/test_kompress_preload_deferral.py::test_local_first_no_network_when_disallowed PASSED
tests/test_kompress_preload_deferral.py::test_local_first_falls_back_to_network_by_default PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_onnx_cache_miss_raises_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_auto_does_not_pytorch_download_on_cache_miss PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_defers_when_model_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_enabled_when_model_cached PASSED
6 passed in 4.82s
$ uv run pytest tests/test_transforms/test_kompress_compressor.py tests/test_transforms_content_router.py tests/test_onnx_runtime.py tests/test_proxy_warmup.py
63 passed
$ uv run ruff check <changed files> # All checks passed!
$ uv run mypy headroom/onnx_runtime.py headroom/transforms/kompress_compressor.py headroom/transforms/content_router.py
Success: no issues found
```
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (auto-generated from
conventional commits)
## Additional Notes
This contains the cold-start case. A native crash in onnxruntime
*session init* (as opposed to the download) on first request would still
be a separate issue; it is not what was observed here (the abort was
during the HF download), and isolating it would be a larger, separate
change.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 19:53:03 +02:00
|
|
|
|
|
|
|
|
status = router.eager_load_compressors()
|
|
|
|
|
|
|
|
|
|
assert status["kompress"] == "deferred"
|
fix(proxy): keep Kompress warmup off the startup path (#2001)
## Description
Proxy startup can enter cached Kompress native model initialization
before binding its port. On the RHEL/CentOS 7-family environment
reported in #1908, that path terminates in a deterministic
`libarrow.so.2400` jemalloc-thread segfault with no Python traceback.
Cache-only preload still initializes native libraries when the model is
already cached.
Defer Kompress model and tokenizer loading out of startup while
preserving the existing lazy request path and the eager warmups for
non-Kompress components. Closes #1908.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Stop cached and uncached Kompress models from entering native preload
during proxy startup.
- Record enabled Kompress as deferred until its existing lazy request
path needs it.
- Preserve disabled-Kompress routing and non-Kompress eager warmups.
- Add lifecycle, cache-state, disabled-mode, and warmup-preservation
regression coverage.
- Document the startup behavior change in the changelog.
## Testing
- [x] Unit tests pass (`uv run pytest
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py -q`)
- [x] Linting passes (`uv run ruff check
headroom/transforms/content_router.py
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py`)
- [ ] Type checking passes (`uv run mypy headroom`)
- [x] New tests added for new functionality when applicable
- [ ] Manual testing performed
### Test Output
```text
uv run pytest tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py -q
17 passed in 1.28s
uv run ruff check headroom/transforms/content_router.py tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py
All checks passed
```
## Real Behavior Proof
- Environment: Red OS 7.3 or equivalent RHEL/CentOS 7-family system,
glibc 2.17, Python 3.11, pyarrow 24.0.0, onnxruntime 1.27.0, cached
Kompress model
- Exact command / steps: start `headroom proxy` with Kompress enabled,
wait 30 seconds, query the loopback health endpoint, verify deferred
Kompress warmup in logs, then inspect `journalctl -k` for new
`libarrow.so.2400` or `jemalloc_bg_thd` faults
- Observed result: automated coverage now proves startup avoids the
cached Kompress preload boundary, keeps non-Kompress warmups live, and
preserves `unavailable` status when dependencies are absent
- Not tested: the native reporter-host run
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Screenshots
Not applicable.
## Additional Notes
Focused local validation included
`tests/test_kompress_request_nonblocking.py` alongside the startup
regression suite. The change is scoped to startup warmup; it does not
claim to repair the external `libarrow.so` or jemalloc incompatibility
when Kompress later executes. `HEADROOM_DISABLE_KOMPRESS=1` remains the
supported narrow workaround for hosts that cannot run the native path.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-13 10:09:45 -04:00
|
|
|
assert stub.preload_calls == []
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_eager_load_keeps_disabled_kompress_disabled(monkeypatch):
|
|
|
|
|
router = ContentRouter(
|
|
|
|
|
ContentRouterConfig(
|
|
|
|
|
enable_kompress=False,
|
|
|
|
|
enable_code_aware=False,
|
|
|
|
|
enable_smart_crusher=False,
|
|
|
|
|
)
|
|
|
|
|
)
|
|
|
|
|
stub = _StubCompressor(cached=True)
|
|
|
|
|
monkeypatch.setattr(router, "_get_kompress", lambda: stub)
|
fix(proxy): make Kompress eager preload cache-only so a cold cache can't block startup (#783)
## Description
`ContentRouter.eager_load_compressors()` runs a network
`hf_hub_download` of the Kompress ONNX model on the **blocking
startup/lifespan path**, before the proxy binds its port. On a cold
cache this is unsafe:
- the download can hang long enough to blow the supervisor's bind
timeout, or
- a native crash in the download/ML stack (an **uncatchable `Fatal
Python error: Aborted` / SIGABRT**) kills the interpreter before it ever
`listen()`s.
Either way the supervisor sees "proxy never opened its port" and gives
up. We observed this in the field from the desktop app (process aborted
during `eager_load_compressors -> _load_kompress_onnx ->
hf_hub_download` of `onnx/kompress-int8.onnx`, while the only Python
thread was parked in the HuggingFace download file-lock; the abort came
from a native thread, so `try/except` at the call site cannot catch it).
The eager preload is a latency optimization and must never be able to
block — or kill — startup. This change makes startup preload
**cache-only**: if the model isn't already cached, we defer the download
to first use (off the startup path) and bind the port normally. Warm
starts are unchanged.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `onnx_runtime.hf_hub_download_local_first(...)`: added `allow_network`
(default `True`). When `False`, a cache miss re-raises the local-lookup
error instead of falling back to a network download.
- `kompress_compressor`: added `allow_download` (default `True`)
threaded through `preload()` -> `_load_kompress()` ->
`_load_kompress_onnx()` / `_load_kompress_pytorch()` and the ModernBERT
tokenizer load. Added `KompressModelNotCached`, raised when a cache-only
load misses. Auto-mode no longer falls back to a PyTorch network
download on a cache-only miss — it propagates so the caller can defer.
- `content_router.eager_load_compressors()`: calls
`preload(allow_download=False)`. On `KompressModelNotCached` it logs and
reports the component as `"deferred"` (a status
`warmup.merge_transform_status` already handles gracefully) instead of
letting a cold download run on the startup path.
Default (first-request) loading behavior and warm-start preload are
unchanged.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
New tests in `tests/test_kompress_preload_deferral.py` cover: cache-only
`hf_hub_download_local_first` never hits the network; default still
falls back; cache-only ONNX load raises `KompressModelNotCached`;
auto-mode does **not** trigger a PyTorch download on a cache-only miss;
and `eager_load_compressors` reports `deferred` (cold) / `enabled`
(warm). Existing `_load_kompress` dispatch tests updated for the new
keyword-only param.
> Note on environment: I do not have a clean reproduction of the native
SIGABRT itself (it depends on a specific machine's HF download/ML native
stack), so the "Manual testing performed" box is left unchecked. The
tests target the structural fix — that startup preload can no longer
perform a network download — which is the precondition for the crash.
## Test Output
```
$ uv run pytest -v tests/test_kompress_preload_deferral.py
tests/test_kompress_preload_deferral.py::test_local_first_no_network_when_disallowed PASSED
tests/test_kompress_preload_deferral.py::test_local_first_falls_back_to_network_by_default PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_onnx_cache_miss_raises_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_auto_does_not_pytorch_download_on_cache_miss PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_defers_when_model_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_enabled_when_model_cached PASSED
6 passed in 4.82s
$ uv run pytest tests/test_transforms/test_kompress_compressor.py tests/test_transforms_content_router.py tests/test_onnx_runtime.py tests/test_proxy_warmup.py
63 passed
$ uv run ruff check <changed files> # All checks passed!
$ uv run mypy headroom/onnx_runtime.py headroom/transforms/kompress_compressor.py headroom/transforms/content_router.py
Success: no issues found
```
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (auto-generated from
conventional commits)
## Additional Notes
This contains the cold-start case. A native crash in onnxruntime
*session init* (as opposed to the download) on first request would still
be a separate issue; it is not what was observed here (the abort was
during the HF download), and isolating it would be a larger, separate
change.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 19:53:03 +02:00
|
|
|
|
fix(proxy): keep Kompress warmup off the startup path (#2001)
## Description
Proxy startup can enter cached Kompress native model initialization
before binding its port. On the RHEL/CentOS 7-family environment
reported in #1908, that path terminates in a deterministic
`libarrow.so.2400` jemalloc-thread segfault with no Python traceback.
Cache-only preload still initializes native libraries when the model is
already cached.
Defer Kompress model and tokenizer loading out of startup while
preserving the existing lazy request path and the eager warmups for
non-Kompress components. Closes #1908.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Stop cached and uncached Kompress models from entering native preload
during proxy startup.
- Record enabled Kompress as deferred until its existing lazy request
path needs it.
- Preserve disabled-Kompress routing and non-Kompress eager warmups.
- Add lifecycle, cache-state, disabled-mode, and warmup-preservation
regression coverage.
- Document the startup behavior change in the changelog.
## Testing
- [x] Unit tests pass (`uv run pytest
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py -q`)
- [x] Linting passes (`uv run ruff check
headroom/transforms/content_router.py
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py`)
- [ ] Type checking passes (`uv run mypy headroom`)
- [x] New tests added for new functionality when applicable
- [ ] Manual testing performed
### Test Output
```text
uv run pytest tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py -q
17 passed in 1.28s
uv run ruff check headroom/transforms/content_router.py tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py
All checks passed
```
## Real Behavior Proof
- Environment: Red OS 7.3 or equivalent RHEL/CentOS 7-family system,
glibc 2.17, Python 3.11, pyarrow 24.0.0, onnxruntime 1.27.0, cached
Kompress model
- Exact command / steps: start `headroom proxy` with Kompress enabled,
wait 30 seconds, query the loopback health endpoint, verify deferred
Kompress warmup in logs, then inspect `journalctl -k` for new
`libarrow.so.2400` or `jemalloc_bg_thd` faults
- Observed result: automated coverage now proves startup avoids the
cached Kompress preload boundary, keeps non-Kompress warmups live, and
preserves `unavailable` status when dependencies are absent
- Not tested: the native reporter-host run
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Screenshots
Not applicable.
## Additional Notes
Focused local validation included
`tests/test_kompress_request_nonblocking.py` alongside the startup
regression suite. The change is scoped to startup warmup; it does not
claim to repair the external `libarrow.so` or jemalloc incompatibility
when Kompress later executes. `HEADROOM_DISABLE_KOMPRESS=1` remains the
supported narrow workaround for hosts that cannot run the native path.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-13 10:09:45 -04:00
|
|
|
status = router.eager_load_compressors()
|
|
|
|
|
|
|
|
|
|
assert "kompress" not in status
|
|
|
|
|
assert stub.preload_calls == []
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_eager_load_reports_unavailable_kompress(monkeypatch):
|
|
|
|
|
router = _router_kompress_only()
|
|
|
|
|
monkeypatch.setattr(router, "_get_kompress", lambda: None)
|
|
|
|
|
|
|
|
|
|
status = router.eager_load_compressors()
|
fix(proxy): make Kompress eager preload cache-only so a cold cache can't block startup (#783)
## Description
`ContentRouter.eager_load_compressors()` runs a network
`hf_hub_download` of the Kompress ONNX model on the **blocking
startup/lifespan path**, before the proxy binds its port. On a cold
cache this is unsafe:
- the download can hang long enough to blow the supervisor's bind
timeout, or
- a native crash in the download/ML stack (an **uncatchable `Fatal
Python error: Aborted` / SIGABRT**) kills the interpreter before it ever
`listen()`s.
Either way the supervisor sees "proxy never opened its port" and gives
up. We observed this in the field from the desktop app (process aborted
during `eager_load_compressors -> _load_kompress_onnx ->
hf_hub_download` of `onnx/kompress-int8.onnx`, while the only Python
thread was parked in the HuggingFace download file-lock; the abort came
from a native thread, so `try/except` at the call site cannot catch it).
The eager preload is a latency optimization and must never be able to
block — or kill — startup. This change makes startup preload
**cache-only**: if the model isn't already cached, we defer the download
to first use (off the startup path) and bind the port normally. Warm
starts are unchanged.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `onnx_runtime.hf_hub_download_local_first(...)`: added `allow_network`
(default `True`). When `False`, a cache miss re-raises the local-lookup
error instead of falling back to a network download.
- `kompress_compressor`: added `allow_download` (default `True`)
threaded through `preload()` -> `_load_kompress()` ->
`_load_kompress_onnx()` / `_load_kompress_pytorch()` and the ModernBERT
tokenizer load. Added `KompressModelNotCached`, raised when a cache-only
load misses. Auto-mode no longer falls back to a PyTorch network
download on a cache-only miss — it propagates so the caller can defer.
- `content_router.eager_load_compressors()`: calls
`preload(allow_download=False)`. On `KompressModelNotCached` it logs and
reports the component as `"deferred"` (a status
`warmup.merge_transform_status` already handles gracefully) instead of
letting a cold download run on the startup path.
Default (first-request) loading behavior and warm-start preload are
unchanged.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
New tests in `tests/test_kompress_preload_deferral.py` cover: cache-only
`hf_hub_download_local_first` never hits the network; default still
falls back; cache-only ONNX load raises `KompressModelNotCached`;
auto-mode does **not** trigger a PyTorch download on a cache-only miss;
and `eager_load_compressors` reports `deferred` (cold) / `enabled`
(warm). Existing `_load_kompress` dispatch tests updated for the new
keyword-only param.
> Note on environment: I do not have a clean reproduction of the native
SIGABRT itself (it depends on a specific machine's HF download/ML native
stack), so the "Manual testing performed" box is left unchecked. The
tests target the structural fix — that startup preload can no longer
perform a network download — which is the precondition for the crash.
## Test Output
```
$ uv run pytest -v tests/test_kompress_preload_deferral.py
tests/test_kompress_preload_deferral.py::test_local_first_no_network_when_disallowed PASSED
tests/test_kompress_preload_deferral.py::test_local_first_falls_back_to_network_by_default PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_onnx_cache_miss_raises_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_auto_does_not_pytorch_download_on_cache_miss PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_defers_when_model_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_enabled_when_model_cached PASSED
6 passed in 4.82s
$ uv run pytest tests/test_transforms/test_kompress_compressor.py tests/test_transforms_content_router.py tests/test_onnx_runtime.py tests/test_proxy_warmup.py
63 passed
$ uv run ruff check <changed files> # All checks passed!
$ uv run mypy headroom/onnx_runtime.py headroom/transforms/kompress_compressor.py headroom/transforms/content_router.py
Success: no issues found
```
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (auto-generated from
conventional commits)
## Additional Notes
This contains the cold-start case. A native crash in onnxruntime
*session init* (as opposed to the download) on first request would still
be a separate issue; it is not what was observed here (the abort was
during the HF download), and isolating it would be a larger, separate
change.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 19:53:03 +02:00
|
|
|
|
fix(proxy): keep Kompress warmup off the startup path (#2001)
## Description
Proxy startup can enter cached Kompress native model initialization
before binding its port. On the RHEL/CentOS 7-family environment
reported in #1908, that path terminates in a deterministic
`libarrow.so.2400` jemalloc-thread segfault with no Python traceback.
Cache-only preload still initializes native libraries when the model is
already cached.
Defer Kompress model and tokenizer loading out of startup while
preserving the existing lazy request path and the eager warmups for
non-Kompress components. Closes #1908.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Stop cached and uncached Kompress models from entering native preload
during proxy startup.
- Record enabled Kompress as deferred until its existing lazy request
path needs it.
- Preserve disabled-Kompress routing and non-Kompress eager warmups.
- Add lifecycle, cache-state, disabled-mode, and warmup-preservation
regression coverage.
- Document the startup behavior change in the changelog.
## Testing
- [x] Unit tests pass (`uv run pytest
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py -q`)
- [x] Linting passes (`uv run ruff check
headroom/transforms/content_router.py
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py`)
- [ ] Type checking passes (`uv run mypy headroom`)
- [x] New tests added for new functionality when applicable
- [ ] Manual testing performed
### Test Output
```text
uv run pytest tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py -q
17 passed in 1.28s
uv run ruff check headroom/transforms/content_router.py tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py
All checks passed
```
## Real Behavior Proof
- Environment: Red OS 7.3 or equivalent RHEL/CentOS 7-family system,
glibc 2.17, Python 3.11, pyarrow 24.0.0, onnxruntime 1.27.0, cached
Kompress model
- Exact command / steps: start `headroom proxy` with Kompress enabled,
wait 30 seconds, query the loopback health endpoint, verify deferred
Kompress warmup in logs, then inspect `journalctl -k` for new
`libarrow.so.2400` or `jemalloc_bg_thd` faults
- Observed result: automated coverage now proves startup avoids the
cached Kompress preload boundary, keeps non-Kompress warmups live, and
preserves `unavailable` status when dependencies are absent
- Not tested: the native reporter-host run
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Screenshots
Not applicable.
## Additional Notes
Focused local validation included
`tests/test_kompress_request_nonblocking.py` alongside the startup
regression suite. The change is scoped to startup warmup; it does not
claim to repair the external `libarrow.so` or jemalloc incompatibility
when Kompress later executes. `HEADROOM_DISABLE_KOMPRESS=1` remains the
supported narrow workaround for hosts that cannot run the native path.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-13 10:09:45 -04:00
|
|
|
assert status["kompress"] == "unavailable"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_non_kompress_warmups_continue_when_kompress_is_deferred(monkeypatch):
|
fix(proxy): make Kompress eager preload cache-only so a cold cache can't block startup (#783)
## Description
`ContentRouter.eager_load_compressors()` runs a network
`hf_hub_download` of the Kompress ONNX model on the **blocking
startup/lifespan path**, before the proxy binds its port. On a cold
cache this is unsafe:
- the download can hang long enough to blow the supervisor's bind
timeout, or
- a native crash in the download/ML stack (an **uncatchable `Fatal
Python error: Aborted` / SIGABRT**) kills the interpreter before it ever
`listen()`s.
Either way the supervisor sees "proxy never opened its port" and gives
up. We observed this in the field from the desktop app (process aborted
during `eager_load_compressors -> _load_kompress_onnx ->
hf_hub_download` of `onnx/kompress-int8.onnx`, while the only Python
thread was parked in the HuggingFace download file-lock; the abort came
from a native thread, so `try/except` at the call site cannot catch it).
The eager preload is a latency optimization and must never be able to
block — or kill — startup. This change makes startup preload
**cache-only**: if the model isn't already cached, we defer the download
to first use (off the startup path) and bind the port normally. Warm
starts are unchanged.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `onnx_runtime.hf_hub_download_local_first(...)`: added `allow_network`
(default `True`). When `False`, a cache miss re-raises the local-lookup
error instead of falling back to a network download.
- `kompress_compressor`: added `allow_download` (default `True`)
threaded through `preload()` -> `_load_kompress()` ->
`_load_kompress_onnx()` / `_load_kompress_pytorch()` and the ModernBERT
tokenizer load. Added `KompressModelNotCached`, raised when a cache-only
load misses. Auto-mode no longer falls back to a PyTorch network
download on a cache-only miss — it propagates so the caller can defer.
- `content_router.eager_load_compressors()`: calls
`preload(allow_download=False)`. On `KompressModelNotCached` it logs and
reports the component as `"deferred"` (a status
`warmup.merge_transform_status` already handles gracefully) instead of
letting a cold download run on the startup path.
Default (first-request) loading behavior and warm-start preload are
unchanged.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
New tests in `tests/test_kompress_preload_deferral.py` cover: cache-only
`hf_hub_download_local_first` never hits the network; default still
falls back; cache-only ONNX load raises `KompressModelNotCached`;
auto-mode does **not** trigger a PyTorch download on a cache-only miss;
and `eager_load_compressors` reports `deferred` (cold) / `enabled`
(warm). Existing `_load_kompress` dispatch tests updated for the new
keyword-only param.
> Note on environment: I do not have a clean reproduction of the native
SIGABRT itself (it depends on a specific machine's HF download/ML native
stack), so the "Manual testing performed" box is left unchecked. The
tests target the structural fix — that startup preload can no longer
perform a network download — which is the precondition for the crash.
## Test Output
```
$ uv run pytest -v tests/test_kompress_preload_deferral.py
tests/test_kompress_preload_deferral.py::test_local_first_no_network_when_disallowed PASSED
tests/test_kompress_preload_deferral.py::test_local_first_falls_back_to_network_by_default PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_onnx_cache_miss_raises_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_auto_does_not_pytorch_download_on_cache_miss PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_defers_when_model_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_enabled_when_model_cached PASSED
6 passed in 4.82s
$ uv run pytest tests/test_transforms/test_kompress_compressor.py tests/test_transforms_content_router.py tests/test_onnx_runtime.py tests/test_proxy_warmup.py
63 passed
$ uv run ruff check <changed files> # All checks passed!
$ uv run mypy headroom/onnx_runtime.py headroom/transforms/kompress_compressor.py headroom/transforms/content_router.py
Success: no issues found
```
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (auto-generated from
conventional commits)
## Additional Notes
This contains the cold-start case. A native crash in onnxruntime
*session init* (as opposed to the download) on first request would still
be a separate issue; it is not what was observed here (the abort was
during the HF download), and isolating it would be a larger, separate
change.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 19:53:03 +02:00
|
|
|
router = _router_kompress_only()
|
|
|
|
|
stub = _StubCompressor(cached=True)
|
|
|
|
|
monkeypatch.setattr(router, "_get_kompress", lambda: stub)
|
fix(kompress): reject artifacts that fail at run, and prefetch model files at startup (#2740)
## Description
Three cold-start / robustness gaps found while debugging a user report
of **0.12% savings across 722 requests** (49.8M input tokens, 60,920
saved).
Closes #
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
## Changes Made
### 1. The artifact fallback was unreachable for run-time failures
`_create_onnx_session` tries `int8-wo` → `fp32` → `int8`, and its
docstring describes exactly this scenario — but it only skipped a
candidate when `InferenceSession(...)` **construction** threw.
The int8 weight-only artifact carries `MatMulNBits` with `bits=8`. ORT's
CPU kernel only handles 8-bit through the prepacked MLAS path, so a
build or ISA without an 8-bit `SQNBitGemm` kernel falls into
`ComputeBUnpacked`, which hard-asserts `nbits_ == 4`. That raises on
`session.run()` **after** construction succeeded — so the fp32 candidate
was never reached and ML compression was dead for the process lifetime.
The reported log has 207 consecutive failures over three days.
A two-token `_smoke_run` inside the existing candidate loop makes the
fallback fire. `onnxruntime>=1.16.0` is unpinned, so which side of this
an install lands on is a lottery.
### 2. A broken model cost an inference on every request, forever
The per-request handler logged a `WARNING` and passed through with no
latch — 207 identical lines that read as noise rather than "ML
compression is dead". Now latches to passthrough after **3 consecutive**
failures (any success resets the count) with one actionable `ERROR`
naming the artifact override.
### 3. The model download began on the first request, not at startup
#2001 was right to move Kompress off the startup path — on RHEL/CentOS
7-family hosts, entering cached native init before the port binds
segfaults in `libarrow`/jemalloc with no Python traceback (#1908), which
no `try/except` can catch. **This PR does not touch that.**
But #2001 left the ~4-minute *download* on the first request, with every
request in that window silently uncompressed behind one "model not
ready" warning.
Downloading is separable from loading. `prefetch_kompress_artifacts`
resolves the files over plain `huggingface_hub` HTTP and never
constructs an `InferenceSession` or imports `transformers`, so startup
can prefetch bytes without touching the boundary #1908 crashes on.
Native load stays deferred, status stays `deferred`, and a test asserts
no session is constructed during prefetch.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ .venv/bin/ruff check headroom/ tests/test_kompress_failsafe.py tests/test_kompress_preload_deferral.py --exclude headroom/dashboard/templates
All checks passed!
$ .venv/bin/mypy headroom/
Success: no issues found in 508 source files
$ python -m pytest tests/test_kompress_failsafe.py tests/test_kompress_preload_deferral.py \
tests/test_kompress_request_nonblocking.py tests/test_force_kompress_all.py \
tests/test_kompress_must_keep.py tests/test_proxy_disable_kompress.py \
tests/test_proxy_per_provider_kompress.py tests/test_proxy_warmup.py \
tests/test_proxy_eager_preload_bind.py -q
95 passed in 10.51s
```
## Real Behavior Proof
- **Environment:** macOS 26.4 arm64, Python 3.12.6, onnxruntime 1.21.1,
repo `.venv`.
**(1) Fallback chain, against the real HF repo:**
```text
WARNING ONNX artifact 'onnx/kompress-int8-wo.onnx' from chopratejas/kompress-v2-base
is unusable (... nbits_ == 4 was false ...); trying next candidate
SESSION OK -> ['input_ids', 'attention_mask']
SMOKE RUN OK on the selected artifact
```
Also confirmed the default artifact really is 8-bit, by loading the
cached blob: `{'bits': [8], 'block_size': [128]}`.
**(2) Files-only prefetch, with `InferenceSession` patched to raise:**
```text
INFO Kompress: prefetching model artifacts for chopratejas/kompress-v2-base ...
prefetch ok=True in 0.08s, no session constructed
```
- **Not tested / important caveat:** the user's exact failure **cannot
be reproduced on this machine**. On ORT 1.21.1 arm64 the int8-wo
artifact fails at *construction* (`matmul_nbits.cc:115`), which the
pre-existing load-only fallback already caught. Their build fails at
*execution* (`matmul_nbits.cc:442`, `ComputeBUnpacked`). So the run-time
path is pinned with a fake ORT session that constructs fine and then
rejects `run()` — a mechanism test, not a reproduction of their build.
Confirming the fix on their host needs their `onnxruntime` version.
- **Not tested:** no RHEL/CentOS 7 host available to re-verify #1908
non-regression; the argument is structural (prefetch never constructs a
session) and asserted by
`test_prefetch_never_constructs_a_session_or_imports_transformers`.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md`
🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-03 10:42:48 -07:00
|
|
|
monkeypatch.setattr(router, "_prefetch_kompress_artifacts_async", lambda _cfg: False)
|
fix(proxy): keep Kompress warmup off the startup path (#2001)
## Description
Proxy startup can enter cached Kompress native model initialization
before binding its port. On the RHEL/CentOS 7-family environment
reported in #1908, that path terminates in a deterministic
`libarrow.so.2400` jemalloc-thread segfault with no Python traceback.
Cache-only preload still initializes native libraries when the model is
already cached.
Defer Kompress model and tokenizer loading out of startup while
preserving the existing lazy request path and the eager warmups for
non-Kompress components. Closes #1908.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Stop cached and uncached Kompress models from entering native preload
during proxy startup.
- Record enabled Kompress as deferred until its existing lazy request
path needs it.
- Preserve disabled-Kompress routing and non-Kompress eager warmups.
- Add lifecycle, cache-state, disabled-mode, and warmup-preservation
regression coverage.
- Document the startup behavior change in the changelog.
## Testing
- [x] Unit tests pass (`uv run pytest
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py -q`)
- [x] Linting passes (`uv run ruff check
headroom/transforms/content_router.py
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py`)
- [ ] Type checking passes (`uv run mypy headroom`)
- [x] New tests added for new functionality when applicable
- [ ] Manual testing performed
### Test Output
```text
uv run pytest tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py -q
17 passed in 1.28s
uv run ruff check headroom/transforms/content_router.py tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py
All checks passed
```
## Real Behavior Proof
- Environment: Red OS 7.3 or equivalent RHEL/CentOS 7-family system,
glibc 2.17, Python 3.11, pyarrow 24.0.0, onnxruntime 1.27.0, cached
Kompress model
- Exact command / steps: start `headroom proxy` with Kompress enabled,
wait 30 seconds, query the loopback health endpoint, verify deferred
Kompress warmup in logs, then inspect `journalctl -k` for new
`libarrow.so.2400` or `jemalloc_bg_thd` faults
- Observed result: automated coverage now proves startup avoids the
cached Kompress preload boundary, keeps non-Kompress warmups live, and
preserves `unavailable` status when dependencies are absent
- Not tested: the native reporter-host run
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Screenshots
Not applicable.
## Additional Notes
Focused local validation included
`tests/test_kompress_request_nonblocking.py` alongside the startup
regression suite. The change is scoped to startup warmup; it does not
claim to repair the external `libarrow.so` or jemalloc incompatibility
when Kompress later executes. `HEADROOM_DISABLE_KOMPRESS=1` remains the
supported narrow workaround for hosts that cannot run the native path.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-13 10:09:45 -04:00
|
|
|
monkeypatch.setattr("headroom.compression.detector._magika_available", lambda: True)
|
|
|
|
|
monkeypatch.setattr("headroom.compression.detector._get_magika", lambda: object())
|
fix(proxy): make Kompress eager preload cache-only so a cold cache can't block startup (#783)
## Description
`ContentRouter.eager_load_compressors()` runs a network
`hf_hub_download` of the Kompress ONNX model on the **blocking
startup/lifespan path**, before the proxy binds its port. On a cold
cache this is unsafe:
- the download can hang long enough to blow the supervisor's bind
timeout, or
- a native crash in the download/ML stack (an **uncatchable `Fatal
Python error: Aborted` / SIGABRT**) kills the interpreter before it ever
`listen()`s.
Either way the supervisor sees "proxy never opened its port" and gives
up. We observed this in the field from the desktop app (process aborted
during `eager_load_compressors -> _load_kompress_onnx ->
hf_hub_download` of `onnx/kompress-int8.onnx`, while the only Python
thread was parked in the HuggingFace download file-lock; the abort came
from a native thread, so `try/except` at the call site cannot catch it).
The eager preload is a latency optimization and must never be able to
block — or kill — startup. This change makes startup preload
**cache-only**: if the model isn't already cached, we defer the download
to first use (off the startup path) and bind the port normally. Warm
starts are unchanged.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- `onnx_runtime.hf_hub_download_local_first(...)`: added `allow_network`
(default `True`). When `False`, a cache miss re-raises the local-lookup
error instead of falling back to a network download.
- `kompress_compressor`: added `allow_download` (default `True`)
threaded through `preload()` -> `_load_kompress()` ->
`_load_kompress_onnx()` / `_load_kompress_pytorch()` and the ModernBERT
tokenizer load. Added `KompressModelNotCached`, raised when a cache-only
load misses. Auto-mode no longer falls back to a PyTorch network
download on a cache-only miss — it propagates so the caller can defer.
- `content_router.eager_load_compressors()`: calls
`preload(allow_download=False)`. On `KompressModelNotCached` it logs and
reports the component as `"deferred"` (a status
`warmup.merge_transform_status` already handles gracefully) instead of
letting a cold download run on the startup path.
Default (first-request) loading behavior and warm-start preload are
unchanged.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed
New tests in `tests/test_kompress_preload_deferral.py` cover: cache-only
`hf_hub_download_local_first` never hits the network; default still
falls back; cache-only ONNX load raises `KompressModelNotCached`;
auto-mode does **not** trigger a PyTorch download on a cache-only miss;
and `eager_load_compressors` reports `deferred` (cold) / `enabled`
(warm). Existing `_load_kompress` dispatch tests updated for the new
keyword-only param.
> Note on environment: I do not have a clean reproduction of the native
SIGABRT itself (it depends on a specific machine's HF download/ML native
stack), so the "Manual testing performed" box is left unchecked. The
tests target the structural fix — that startup preload can no longer
perform a network download — which is the precondition for the crash.
## Test Output
```
$ uv run pytest -v tests/test_kompress_preload_deferral.py
tests/test_kompress_preload_deferral.py::test_local_first_no_network_when_disallowed PASSED
tests/test_kompress_preload_deferral.py::test_local_first_falls_back_to_network_by_default PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_onnx_cache_miss_raises_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_load_kompress_auto_does_not_pytorch_download_on_cache_miss PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_defers_when_model_not_cached PASSED
tests/test_kompress_preload_deferral.py::test_eager_load_enabled_when_model_cached PASSED
6 passed in 4.82s
$ uv run pytest tests/test_transforms/test_kompress_compressor.py tests/test_transforms_content_router.py tests/test_onnx_runtime.py tests/test_proxy_warmup.py
63 passed
$ uv run ruff check <changed files> # All checks passed!
$ uv run mypy headroom/onnx_runtime.py headroom/transforms/kompress_compressor.py headroom/transforms/content_router.py
Success: no issues found
```
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable (auto-generated from
conventional commits)
## Additional Notes
This contains the cold-start case. A native crash in onnxruntime
*session init* (as opposed to the download) on first request would still
be a separate issue; it is not what was observed here (the abort was
during the HF download), and isolating it would be a larger, separate
change.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-11 19:53:03 +02:00
|
|
|
|
|
|
|
|
status = router.eager_load_compressors()
|
|
|
|
|
|
fix(proxy): keep Kompress warmup off the startup path (#2001)
## Description
Proxy startup can enter cached Kompress native model initialization
before binding its port. On the RHEL/CentOS 7-family environment
reported in #1908, that path terminates in a deterministic
`libarrow.so.2400` jemalloc-thread segfault with no Python traceback.
Cache-only preload still initializes native libraries when the model is
already cached.
Defer Kompress model and tokenizer loading out of startup while
preserving the existing lazy request path and the eager warmups for
non-Kompress components. Closes #1908.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Stop cached and uncached Kompress models from entering native preload
during proxy startup.
- Record enabled Kompress as deferred until its existing lazy request
path needs it.
- Preserve disabled-Kompress routing and non-Kompress eager warmups.
- Add lifecycle, cache-state, disabled-mode, and warmup-preservation
regression coverage.
- Document the startup behavior change in the changelog.
## Testing
- [x] Unit tests pass (`uv run pytest
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py -q`)
- [x] Linting passes (`uv run ruff check
headroom/transforms/content_router.py
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py`)
- [ ] Type checking passes (`uv run mypy headroom`)
- [x] New tests added for new functionality when applicable
- [ ] Manual testing performed
### Test Output
```text
uv run pytest tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py -q
17 passed in 1.28s
uv run ruff check headroom/transforms/content_router.py tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py
All checks passed
```
## Real Behavior Proof
- Environment: Red OS 7.3 or equivalent RHEL/CentOS 7-family system,
glibc 2.17, Python 3.11, pyarrow 24.0.0, onnxruntime 1.27.0, cached
Kompress model
- Exact command / steps: start `headroom proxy` with Kompress enabled,
wait 30 seconds, query the loopback health endpoint, verify deferred
Kompress warmup in logs, then inspect `journalctl -k` for new
`libarrow.so.2400` or `jemalloc_bg_thd` faults
- Observed result: automated coverage now proves startup avoids the
cached Kompress preload boundary, keeps non-Kompress warmups live, and
preserves `unavailable` status when dependencies are absent
- Not tested: the native reporter-host run
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Screenshots
Not applicable.
## Additional Notes
Focused local validation included
`tests/test_kompress_request_nonblocking.py` alongside the startup
regression suite. The change is scoped to startup warmup; it does not
claim to repair the external `libarrow.so` or jemalloc incompatibility
when Kompress later executes. `HEADROOM_DISABLE_KOMPRESS=1` remains the
supported narrow workaround for hosts that cannot run the native path.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-13 10:09:45 -04:00
|
|
|
assert status["kompress"] == "deferred"
|
|
|
|
|
assert status["magika"] == "enabled"
|
|
|
|
|
assert stub.preload_calls == []
|
|
|
|
|
|
|
|
|
|
|
fix(kompress): reject artifacts that fail at run, and prefetch model files at startup (#2740)
## Description
Three cold-start / robustness gaps found while debugging a user report
of **0.12% savings across 722 requests** (49.8M input tokens, 60,920
saved).
Closes #
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
## Changes Made
### 1. The artifact fallback was unreachable for run-time failures
`_create_onnx_session` tries `int8-wo` → `fp32` → `int8`, and its
docstring describes exactly this scenario — but it only skipped a
candidate when `InferenceSession(...)` **construction** threw.
The int8 weight-only artifact carries `MatMulNBits` with `bits=8`. ORT's
CPU kernel only handles 8-bit through the prepacked MLAS path, so a
build or ISA without an 8-bit `SQNBitGemm` kernel falls into
`ComputeBUnpacked`, which hard-asserts `nbits_ == 4`. That raises on
`session.run()` **after** construction succeeded — so the fp32 candidate
was never reached and ML compression was dead for the process lifetime.
The reported log has 207 consecutive failures over three days.
A two-token `_smoke_run` inside the existing candidate loop makes the
fallback fire. `onnxruntime>=1.16.0` is unpinned, so which side of this
an install lands on is a lottery.
### 2. A broken model cost an inference on every request, forever
The per-request handler logged a `WARNING` and passed through with no
latch — 207 identical lines that read as noise rather than "ML
compression is dead". Now latches to passthrough after **3 consecutive**
failures (any success resets the count) with one actionable `ERROR`
naming the artifact override.
### 3. The model download began on the first request, not at startup
#2001 was right to move Kompress off the startup path — on RHEL/CentOS
7-family hosts, entering cached native init before the port binds
segfaults in `libarrow`/jemalloc with no Python traceback (#1908), which
no `try/except` can catch. **This PR does not touch that.**
But #2001 left the ~4-minute *download* on the first request, with every
request in that window silently uncompressed behind one "model not
ready" warning.
Downloading is separable from loading. `prefetch_kompress_artifacts`
resolves the files over plain `huggingface_hub` HTTP and never
constructs an `InferenceSession` or imports `transformers`, so startup
can prefetch bytes without touching the boundary #1908 crashes on.
Native load stays deferred, status stays `deferred`, and a test asserts
no session is constructed during prefetch.
## Testing
- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed
### Test Output
```text
$ .venv/bin/ruff check headroom/ tests/test_kompress_failsafe.py tests/test_kompress_preload_deferral.py --exclude headroom/dashboard/templates
All checks passed!
$ .venv/bin/mypy headroom/
Success: no issues found in 508 source files
$ python -m pytest tests/test_kompress_failsafe.py tests/test_kompress_preload_deferral.py \
tests/test_kompress_request_nonblocking.py tests/test_force_kompress_all.py \
tests/test_kompress_must_keep.py tests/test_proxy_disable_kompress.py \
tests/test_proxy_per_provider_kompress.py tests/test_proxy_warmup.py \
tests/test_proxy_eager_preload_bind.py -q
95 passed in 10.51s
```
## Real Behavior Proof
- **Environment:** macOS 26.4 arm64, Python 3.12.6, onnxruntime 1.21.1,
repo `.venv`.
**(1) Fallback chain, against the real HF repo:**
```text
WARNING ONNX artifact 'onnx/kompress-int8-wo.onnx' from chopratejas/kompress-v2-base
is unusable (... nbits_ == 4 was false ...); trying next candidate
SESSION OK -> ['input_ids', 'attention_mask']
SMOKE RUN OK on the selected artifact
```
Also confirmed the default artifact really is 8-bit, by loading the
cached blob: `{'bits': [8], 'block_size': [128]}`.
**(2) Files-only prefetch, with `InferenceSession` patched to raise:**
```text
INFO Kompress: prefetching model artifacts for chopratejas/kompress-v2-base ...
prefetch ok=True in 0.08s, no session constructed
```
- **Not tested / important caveat:** the user's exact failure **cannot
be reproduced on this machine**. On ORT 1.21.1 arm64 the int8-wo
artifact fails at *construction* (`matmul_nbits.cc:115`), which the
pre-existing load-only fallback already caught. Their build fails at
*execution* (`matmul_nbits.cc:442`, `ComputeBUnpacked`). So the run-time
path is pinned with a fake ORT session that constructs fine and then
rejects `run()` — a mechanism test, not a reproduction of their build.
Confirming the fix on their host needs their `onnxruntime` version.
- **Not tested:** no RHEL/CentOS 7 host available to re-verify #1908
non-regression; the argument is structural (prefetch never constructs a
session) and asserted by
`test_prefetch_never_constructs_a_session_or_imports_transformers`.
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md`
🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-03 10:42:48 -07:00
|
|
|
# ── Startup artifact prefetch (files only, never native init) ──────────────────
|
|
|
|
|
# The cold-start cost #2001 left behind: the ~4-minute model download began on the
|
|
|
|
|
# FIRST REQUEST, so every request in that window went silently uncompressed.
|
|
|
|
|
# Prefetching FILES at startup is safe because it is plain huggingface_hub HTTP —
|
|
|
|
|
# it never constructs an InferenceSession, which is the boundary that segfaults in
|
|
|
|
|
# libarrow/jemalloc on RHEL/CentOS 7-family hosts (#1908).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_eager_load_starts_artifact_prefetch_without_native_preload(monkeypatch):
|
|
|
|
|
router = _router_kompress_only()
|
|
|
|
|
stub = _StubCompressor(cached=False)
|
|
|
|
|
monkeypatch.setattr(router, "_get_kompress", lambda: stub)
|
|
|
|
|
prefetch_calls: list[object] = []
|
|
|
|
|
monkeypatch.setattr(
|
|
|
|
|
router,
|
|
|
|
|
"_prefetch_kompress_artifacts_async",
|
|
|
|
|
lambda cfg: (prefetch_calls.append(cfg), True)[1],
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
status = router.eager_load_compressors()
|
|
|
|
|
|
|
|
|
|
assert status["kompress_artifacts"] == "prefetching"
|
|
|
|
|
# The #2001 invariant still holds: no native preload on the startup path.
|
|
|
|
|
assert status["kompress"] == "deferred"
|
|
|
|
|
assert stub.preload_calls == []
|
|
|
|
|
assert len(prefetch_calls) == 1
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_prefetch_never_constructs_a_session_or_imports_transformers(monkeypatch):
|
|
|
|
|
"""The safety property that makes startup prefetch legal at all."""
|
|
|
|
|
monkeypatch.setattr(kc, "_kompress_cache", {})
|
|
|
|
|
requested: list[str] = []
|
|
|
|
|
|
|
|
|
|
def fake_local_first(repo_id, filename, *, allow_network=True):
|
|
|
|
|
requested.append(filename)
|
|
|
|
|
return f"/cache/{filename}"
|
|
|
|
|
|
|
|
|
|
monkeypatch.setattr(kc, "hf_hub_download_local_first", fake_local_first)
|
|
|
|
|
|
|
|
|
|
def explode(*args, **kwargs):
|
|
|
|
|
raise AssertionError("prefetch must not build the model")
|
|
|
|
|
|
|
|
|
|
monkeypatch.setattr(kc, "_load_kompress", explode)
|
|
|
|
|
monkeypatch.setattr(kc, "_load_kompress_onnx", explode)
|
|
|
|
|
|
|
|
|
|
assert kc.prefetch_kompress_artifacts("org/model") is True
|
|
|
|
|
# Stops at the first candidate that resolves — the loader tries the same order.
|
|
|
|
|
assert requested == [kc._onnx_filename_candidates()[0]]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_prefetch_reports_false_when_no_artifact_resolves(monkeypatch):
|
|
|
|
|
monkeypatch.setattr(kc, "_kompress_cache", {})
|
|
|
|
|
|
|
|
|
|
def always_missing(repo_id, filename, *, allow_network=True):
|
|
|
|
|
raise OSError("not found")
|
|
|
|
|
|
|
|
|
|
monkeypatch.setattr(kc, "hf_hub_download_local_first", always_missing)
|
|
|
|
|
|
|
|
|
|
assert kc.prefetch_kompress_artifacts("org/model") is False
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def test_background_prefetch_is_noop_when_model_already_cached(monkeypatch):
|
|
|
|
|
monkeypatch.setattr(kc, "_kompress_cache", {"org/model": object()})
|
|
|
|
|
monkeypatch.setattr(kc, "is_kompress_available", lambda: True)
|
|
|
|
|
|
|
|
|
|
assert kc.ensure_background_prefetch("org/model") is False
|
|
|
|
|
|
|
|
|
|
|
fix(proxy): keep Kompress warmup off the startup path (#2001)
## Description
Proxy startup can enter cached Kompress native model initialization
before binding its port. On the RHEL/CentOS 7-family environment
reported in #1908, that path terminates in a deterministic
`libarrow.so.2400` jemalloc-thread segfault with no Python traceback.
Cache-only preload still initializes native libraries when the model is
already cached.
Defer Kompress model and tokenizer loading out of startup while
preserving the existing lazy request path and the eager warmups for
non-Kompress components. Closes #1908.
## Type of Change
- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)
## Changes Made
- Stop cached and uncached Kompress models from entering native preload
during proxy startup.
- Record enabled Kompress as deferred until its existing lazy request
path needs it.
- Preserve disabled-Kompress routing and non-Kompress eager warmups.
- Add lifecycle, cache-state, disabled-mode, and warmup-preservation
regression coverage.
- Document the startup behavior change in the changelog.
## Testing
- [x] Unit tests pass (`uv run pytest
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py -q`)
- [x] Linting passes (`uv run ruff check
headroom/transforms/content_router.py
tests/test_kompress_preload_deferral.py
tests/test_kompress_request_nonblocking.py`)
- [ ] Type checking passes (`uv run mypy headroom`)
- [x] New tests added for new functionality when applicable
- [ ] Manual testing performed
### Test Output
```text
uv run pytest tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py -q
17 passed in 1.28s
uv run ruff check headroom/transforms/content_router.py tests/test_kompress_preload_deferral.py tests/test_kompress_request_nonblocking.py
All checks passed
```
## Real Behavior Proof
- Environment: Red OS 7.3 or equivalent RHEL/CentOS 7-family system,
glibc 2.17, Python 3.11, pyarrow 24.0.0, onnxruntime 1.27.0, cached
Kompress model
- Exact command / steps: start `headroom proxy` with Kompress enabled,
wait 30 seconds, query the loopback health endpoint, verify deferred
Kompress warmup in logs, then inspect `journalctl -k` for new
`libarrow.so.2400` or `jemalloc_bg_thd` faults
- Observed result: automated coverage now proves startup avoids the
cached Kompress preload boundary, keeps non-Kompress warmups live, and
preserves `unavailable` status when dependencies are absent
- Not tested: the native reporter-host run
## Review Readiness
- [x] I have performed a self-review
- [x] This PR is ready for human review
## Checklist
- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [ ] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I have updated the CHANGELOG.md if applicable
## Screenshots
Not applicable.
## Additional Notes
Focused local validation included
`tests/test_kompress_request_nonblocking.py` alongside the startup
regression suite. The change is scoped to startup warmup; it does not
claim to repair the external `libarrow.so` or jemalloc incompatibility
when Kompress later executes. `HEADROOM_DISABLE_KOMPRESS=1` remains the
supported narrow workaround for hosts that cannot run the native path.
Co-authored-by: JD Davis <mxjerrett@gmail.com>
2026-07-13 10:09:45 -04:00
|
|
|
@pytest.mark.asyncio
|
|
|
|
|
async def test_proxy_startup_does_not_enter_cached_kompress_native_loader(monkeypatch):
|
|
|
|
|
pytest.importorskip("httpx")
|
|
|
|
|
from headroom.proxy.server import HeadroomProxy, ProxyConfig
|
|
|
|
|
|
|
|
|
|
proxy = HeadroomProxy(
|
|
|
|
|
ProxyConfig(
|
|
|
|
|
optimize=True,
|
|
|
|
|
cache_enabled=False,
|
|
|
|
|
rate_limit_enabled=False,
|
|
|
|
|
cost_tracking_enabled=False,
|
|
|
|
|
code_aware_enabled=False,
|
|
|
|
|
)
|
|
|
|
|
)
|
|
|
|
|
router = _router_kompress_only()
|
|
|
|
|
stub = _FatalPreloadCompressor(cached=True)
|
|
|
|
|
monkeypatch.setattr(router, "_get_kompress", lambda: stub)
|
|
|
|
|
proxy.anthropic_pipeline.transforms = [router]
|
|
|
|
|
proxy.openai_pipeline.transforms = [router]
|
|
|
|
|
|
|
|
|
|
await proxy.startup()
|
|
|
|
|
try:
|
|
|
|
|
assert stub.preload_calls == []
|
|
|
|
|
assert proxy.warmup.kompress.info["source_status"] == "deferred"
|
|
|
|
|
finally:
|
|
|
|
|
await proxy.shutdown()
|