headroom/wiki/troubleshooting.md
DM 64783d8824
fix: skip Magika backend on x86 CPUs without AVX2 (#1162)
## Description

Adds a narrow runtime AVX2 guard before initializing the Magika/ONNX
Runtime detector on x86/x86_64. On x86/x86_64 CPUs without AVX2,
Headroom falls back to existing non-Magika detection tiers instead of
crashing during ONNX Runtime initialization. AVX2-capable systems retain
existing behavior.

Refs #1005

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [x] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- Adds a Magika/ONNX Runtime CPU support guard before `Session::new()`.
- Returns a normal Magika init error on x86/x86_64 hosts without AVX2,
allowing the existing detection chain to fall through to non-Magika
tiers.
- Keeps AVX2-capable x86/x86_64 behavior unchanged.
- Does not apply the x86-specific AVX2 gate on non-x86 targets.
- Adds CPU-aware Rust tests and a short troubleshooting note.

## Testing

- [ ] Unit tests pass (`pytest`)
- [ ] Linting passes (`ruff check .`)
- [ ] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

### Test Output

```text
$ cargo test -p headroom-core --lib --locked
833 passed; 0 failed; 1 ignored

$ cargo test --workspace --locked
passed

$ cargo clippy -p headroom-core --locked -- -D warnings
clean
```

## Real Behavior Proof

- Environment: x86_64 Linux host with AVX but no AVX2 (Intel Xeon
E5-2697 v2 on Proxmox), local build from this branch.
- Exact command / steps: `python -X faulthandler -c 'from headroom._core
import detect_content_type; print(detect_content_type("hello world"))'`
- Observed result: before — process exited with `Fatal Python error:
Illegal instruction`; after — command completed successfully returning
`DetectionResult(content_type="text", ...)`, and full `cargo test -p
headroom-core --lib --locked` passed with 833/0/1.
- Not tested: generic no-AVX CPUs, alternate ONNX Runtime builds,
non-x86 platforms.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable

## Screenshots (if applicable)

N/A

## Additional Notes

This partially addresses #1005 by handling one concrete native crash
class: the Magika detector initializes ONNX Runtime through ort/ort-sys,
whose precompiled runtime can contain AVX2-family instructions. On
AVX-only x86_64 hosts, that initialization can SIGILL before Headroom
can fall back.

Scope:

- This does not introduce generic no-AVX wheels.
- This does not redesign Rust-core packaging.
- This does not disable the Rust core globally.
- This only prevents the Magika/ONNX detector tier from loading on
x86/x86_64 CPUs where AVX2 is unavailable.
- Non-Magika detection tiers continue to run.
- On non-x86 targets, this x86-specific AVX2 gate is not applied.

Changelog omitted: small native detector fallback fix with no public API
change.

Co-authored-by: AI Agent <ai-agent@homelab.internal>
2026-06-30 13:34:17 -05:00

12 KiB

Troubleshooting Guide

Solutions for common Headroom issues.


Proxy Server Issues

"Proxy won't start"

Symptom: headroom proxy fails or hangs.

Solutions:

# 1. Check if port is already in use
lsof -i :8787
# If something is using the port, either kill it or use a different port

# 2. Try a different port
headroom proxy --port 8788

# 3. Check for missing dependencies
pip install "headroom-ai[proxy]"

# 4. Run with request logging
headroom proxy --log-file ~/.headroom/logs/proxy.jsonl --log-messages

"Connection refused" when calling proxy

Symptom: curl: (7) Failed to connect to localhost port 8787

Solutions:

# 1. Verify proxy is running
curl http://localhost:8787/health

# 2. Check if proxy started on a different port
ps aux | grep headroom

# 3. Check firewall settings (macOS)
sudo pfctl -s rules | grep 8787

"Upstream rejects a beta token the client no longer sends"

Symptom: The upstream API returns an error referencing a beta feature (anthropic-beta header) even though the client is no longer sending that header.

Cause: Headroom's SessionBetaTracker re-injects any anthropic-beta token seen earlier in the same session to preserve prefix-cache stability. Once a token is in the tracker it persists for the rest of the session. Stopping the token on the client side alone is not sufficient.

Solution: Set HEADROOM_BETA_HEADER_STICKY=disabled to pass the client's header value verbatim without accumulation:

export HEADROOM_BETA_HEADER_STICKY=disabled
headroom proxy ...

Alternatively, restarting the proxy process clears the in-memory tracker. See Session Beta Header Tracking for details.


"Proxy returns errors for some requests"

Symptom: Some requests work, others fail with 502/503.

Solutions:

# 1. Check proxy logs for the actual error
headroom proxy --log-file ~/.headroom/logs/proxy.jsonl --log-messages

# 2. Verify API key is set
echo $OPENAI_API_KEY  # or ANTHROPIC_API_KEY

# 3. Test the underlying API directly
curl https://api.openai.com/v1/models -H "Authorization: Bearer $OPENAI_API_KEY"

SDK Issues

"No token savings"

Symptom: stats['session']['tokens_saved_total'] is 0.

Diagnosis:

# 1. Check mode
stats = client.get_stats()
print(f"Mode: {stats['config']['mode']}")  # Should be "optimize"

# 2. Check transforms are enabled
print(f"SmartCrusher: {stats['transforms']['smart_crusher_enabled']}")

# 3. Check if content meets threshold
# SmartCrusher only compresses tool outputs > 200 tokens by default

Solutions:

# 1. Ensure mode is "optimize"
client = HeadroomClient(
    original_client=OpenAI(),
    provider=OpenAIProvider(),
    default_mode="optimize",  # NOT "audit"
)

# 2. Or override per-request
response = client.chat.completions.create(
    model="gpt-4o",
    messages=messages,
    headroom_mode="optimize",
)

# 3. Lower the compression threshold
config = HeadroomConfig()
config.smart_crusher.min_tokens_to_crush = 100  # Default is 200

Why It Might Be 0:

  • Mode is "audit" (observation only)
  • Messages don't contain tool outputs
  • Tool outputs are below the token threshold
  • Data isn't compressible (high uniqueness)

"Compression too aggressive"

Symptom: LLM responses are missing information that was in tool outputs.

Solutions:

# 1. Keep more items
config = HeadroomConfig()
config.smart_crusher.max_items_after_crush = 50  # Default: 15

# 2. Skip compression for specific tools
response = client.chat.completions.create(
    model="gpt-4o",
    messages=messages,
    headroom_tool_profiles={
        "important_tool": {"skip_compression": True},
    },
)

# 3. Disable SmartCrusher entirely
config.smart_crusher.enabled = False

"High latency"

Symptom: Requests take longer than expected.

Diagnosis:

import time
import logging

logging.basicConfig(level=logging.DEBUG)

start = time.time()
response = client.chat.completions.create(...)
print(f"Total time: {time.time() - start:.2f}s")

# Check logs for:
# - "SmartCrusher" timing
# - "EmbeddingScorer" timing (slow if using embeddings)

Solutions:

# 1. Use BM25 instead of embeddings (faster)
config = HeadroomConfig()
config.smart_crusher.relevance.tier = "bm25"  # Default may use embeddings

# 2. Increase threshold to skip small payloads
config.smart_crusher.min_tokens_to_crush = 500

# 3. Disable transforms you don't need
config.cache_aligner.enabled = False
config.rolling_window.enabled = False

"ValidationError on setup"

Symptom: validate_setup() returns errors.

Common Issues:

result = client.validate_setup()
print(result)

# Provider error:
# {"provider": {"ok": False, "error": "No API key"}}
# → Set OPENAI_API_KEY or pass api_key to OpenAI()

# Storage error:
# {"storage": {"ok": False, "error": "unable to open database"}}
# → Check path permissions, use :memory: for testing

# Config error:
# {"config": {"ok": False, "error": "Invalid mode"}}
# → Use "audit" or "optimize" only

Solutions:

# 1. For testing, use in-memory storage
client = HeadroomClient(
    original_client=OpenAI(),
    provider=OpenAIProvider(),
    store_url="sqlite:///:memory:",  # No file created
)

# 2. For temp directory storage
import tempfile
import os
db_path = os.path.join(tempfile.gettempdir(), "headroom.db")
client = HeadroomClient(
    original_client=OpenAI(),
    provider=OpenAIProvider(),
    store_url=f"sqlite:///{db_path}",
)

Import/Installation Issues

"pip install fails with C++ compilation error"

Symptom: Installation fails with an error like:

RuntimeError: Unsupported compiler -- at least C++11 support is needed!
ERROR: Failed building wheel for hnswlib

Cause: headroom-ai depends on hnswlib, a C++ extension that must be compiled from source. Slim environments (Docker slim images, minimal CI runners) lack the required build tools.

Solutions:

# Linux / Debian-based (including Docker)
apt-get install -y build-essential && pip install headroom-ai

# macOS (Xcode command line tools)
xcode-select --install && pip install headroom-ai

In a Dockerfile, install and remove build tools in one layer to keep the image slim:

FROM python:3.11-slim
RUN apt-get update && apt-get install -y --no-install-recommends build-essential \
    && pip install "headroom-ai[proxy]" \
    && apt-get purge -y build-essential && apt-get autoremove -y \
    && rm -rf /var/lib/apt/lists/*

"ModuleNotFoundError: No module named 'headroom'"

# 1. Check it's installed in the right environment
pip show headroom-ai

# 2. If using virtual environment, ensure it's activated
source venv/bin/activate  # or equivalent

# 3. Reinstall
pip install --upgrade headroom-ai

"ImportError: cannot import name 'X' from 'headroom'"

# Check available imports
import headroom
print(dir(headroom))

# Common imports:
from headroom import (
    HeadroomClient,
    OpenAIProvider,
    AnthropicProvider,
    HeadroomConfig,
    # Exceptions
    HeadroomError,
    ConfigurationError,
    ProviderError,
)

"Missing optional dependency"

# For proxy server
pip install "headroom-ai[proxy]"

# For embedding-based relevance scoring
pip install "headroom-ai[relevance]"

# For everything
pip install "headroom-ai[all]"

Provider-Specific Issues

OpenAI: "Invalid API key"

from openai import OpenAI
import os

# Ensure key is set
api_key = os.environ.get("OPENAI_API_KEY")
if not api_key:
    raise ValueError("OPENAI_API_KEY not set")

client = HeadroomClient(
    original_client=OpenAI(api_key=api_key),
    provider=OpenAIProvider(),
)

Anthropic: "Authentication error"

from anthropic import Anthropic
import os

api_key = os.environ.get("ANTHROPIC_API_KEY")
client = HeadroomClient(
    original_client=Anthropic(api_key=api_key),
    provider=AnthropicProvider(),
)

"Unknown model" warnings

# For custom/fine-tuned models, specify context limit
client = HeadroomClient(
    original_client=OpenAI(),
    provider=OpenAIProvider(),
    model_context_limits={
        "ft:gpt-4o-2024-08-06:my-org::abc123": 128000,
        "my-custom-model": 32000,
    },
)

Debugging Techniques

Enable Full Logging

import logging

# See everything
logging.basicConfig(
    level=logging.DEBUG,
    format="%(asctime)s %(name)s %(levelname)s %(message)s",
)

# Or just Headroom logs
logging.getLogger("headroom").setLevel(logging.DEBUG)

Inspect Transform Results

# Use simulate to see what would happen
plan = client.chat.completions.simulate(
    model="gpt-4o",
    messages=messages,
)

print(f"Tokens: {plan.tokens_before} -> {plan.tokens_after}")
print(f"Transforms: {plan.transforms}")
print(f"Waste signals: {plan.waste_signals}")

# See the actual optimized messages
import json
print(json.dumps(plan.messages_optimized, indent=2))

Check Storage Contents

from datetime import datetime, timedelta

# Get recent metrics
metrics = client.get_metrics(
    start_time=datetime.utcnow() - timedelta(hours=1),
    limit=10,
)

for m in metrics:
    print(f"{m.timestamp}: {m.tokens_input_before} -> {m.tokens_input_after}")
    print(f"  Transforms: {m.transforms_applied}")
    if m.error:
        print(f"  ERROR: {m.error}")

Manual Transform Testing

from headroom import SmartCrusher, Tokenizer
from headroom.config import SmartCrusherConfig
import json

# Test compression directly
config = SmartCrusherConfig()
crusher = SmartCrusher(config)
tokenizer = Tokenizer()

messages = [
    {"role": "tool", "content": json.dumps({"items": list(range(100))}), "tool_call_id": "1"}
]

result = crusher.apply(messages, tokenizer)
print(f"Tokens: {result.tokens_before} -> {result.tokens_after}")
print(f"Compressed content: {result.messages[0]['content'][:200]}...")

"Native detector crashes with illegal instruction"

On some older or virtualized x86_64 CPUs, AVX2 may be unavailable. The Magika/ONNX Runtime detector can require AVX2 through its precompiled runtime binary. Headroom skips that detector tier on x86/x86_64 hosts without AVX2 and falls back to non-Magika detection tiers instead of crashing.

If native startup still fails on an older CPU, set:

export HEADROOM_REQUIRE_RUST_CORE=false

Error Reference

Exception Meaning Solution
ConfigurationError Invalid config values Check config parameters
ProviderError Provider issue (unknown model, etc.) Set model_context_limits
StorageError Database issue Check path/permissions
CompressionError Compression failed Rare - check data format
TokenizationError Token counting failed Check model name
ValidationError Setup validation failed Run validate_setup()

Handling Errors

from headroom import (
    HeadroomClient,
    HeadroomError,
    ConfigurationError,
    StorageError,
)

try:
    client = HeadroomClient(...)
    response = client.chat.completions.create(...)
except ConfigurationError as e:
    print(f"Config issue: {e}")
    print(f"Details: {e.details}")
except StorageError as e:
    print(f"Storage issue: {e}")
    # Headroom continues to work, just without metrics persistence
except HeadroomError as e:
    print(f"Headroom error: {e}")

Getting Help

  1. Enable debug logging and check the output
  2. Use simulate() to see what transforms would apply
  3. Check validate_setup() for configuration issues
  4. File an issue at https://github.com/headroom-sdk/headroom/issues

When filing an issue, include:

  • Headroom version (pip show headroom)
  • Python version
  • Provider (OpenAI/Anthropic)
  • Debug log output
  • Minimal reproduction code