headroom/tests/test_providers
gglucass 0805e8e410
fix(providers/openai): bound tiktoken vocab loads with the guarded loader (#2554)
## Description

`headroom/providers/openai.py::_get_encoding` calls
`tiktoken.get_encoding` directly. tiktoken downloads missing
vocabularies via `requests.get` with **no timeout**, so on a network
that blackholes the vocab CDN (corporate firewall, SSL-intercepting
proxy), whichever thread first counts tokens for an OpenAI model — proxy
startup included — blocks indefinitely.

This is the provider-path hole left by #956: the tokenizer registry
already routes through a bounded loader
(`headroom/tokenizers/tiktoken_counter.py`, worker-thread load +
`HEADROOM_TIKTOKEN_LOAD_TIMEOUT_SECONDS`, default 10s) and falls back to
estimation, but the OpenAI provider path never got the same treatment.

Observed in production (Headroom Desktop fleet, Sentry): a proxy that
never finished booting, with a faulthandler dump wedged in
`tiktoken/registry.py` `get_encoding` on the main thread, reached from
the `headroom` CLI entrypoint via click. The desktop app now also
pre-seeds a persistent `TIKTOKEN_CACHE_DIR`, but the unbounded load
affects every deployment of the proxy, so it should be fixed here too.

Follow-up to #956.

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [ ] Code refactoring (no functional changes)

## Changes Made

- `_get_encoding` now routes through the bounded `load_encoding` from
`headroom.tokenizers.tiktoken_counter` instead of calling
`tiktoken.get_encoding` directly, so a stalled vocab download raises
`TiktokenLoadError` after the timeout instead of hanging the calling
thread.
- `OpenAIProvider.get_token_counter` catches `TiktokenLoadError` and
falls back to `EstimatingTokenCounter`, cached per model so later
requests never re-block on the same failed download — mirroring
`TokenizerRegistry._create_tiktoken`.
- `TIKTOKEN_AVAILABLE` uses `importlib.util.find_spec` (the module-level
`import tiktoken` became unused; same pattern as `LITELLM_AVAILABLE`).
- Two regression tests (`TestGuardedEncodingLoad`) covering the
bounded-raise path and the cached estimation fallback.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed

### Test Output

```text
$ uv run --frozen --extra dev pytest tests/test_providers/ tests/test_tokenizers/
======================= 117 passed, 4 warnings in 27.95s =======================

$ uvx ruff check headroom/providers/openai.py tests/test_providers/test_openai.py
All checks passed!
$ uvx ruff format --check headroom/providers/openai.py tests/test_providers/test_openai.py
2 files already formatted

$ uv run --frozen --extra dev mypy headroom/providers/openai.py
Success: no issues found in 1 source file
```

## Real Behavior Proof

- Environment: macOS (arm64), Python 3.12, uv-managed venv, branch off
`upstream/main` (a6d4921e).
- Exact command / steps: the stalled download is simulated in
`TestGuardedEncodingLoad` by monkeypatching
`tiktoken_counter.load_encoding` to raise `TiktokenLoadError`;
`OpenAITokenCounter("gpt-4o")` then raises the bounded error, and
`OpenAIProvider.get_token_counter("gpt-4o")` returns a working
`EstimatingTokenCounter` and reuses the same instance on the second
call.
- Observed result: tests pass (see output above); with an empty tiktoken
on-disk cache, the encoding load path is identical to the one in the
production faulthandler dump.
- Not tested: an end-to-end run against a genuinely blackholed vocab CDN
(needs a firewalled network); the timeout mechanism itself is #956's
code, already covered by its own tests.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I commented my code, particularly in hard-to-understand areas
- [x] I made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I added tests that prove my fix is effective or that my feature
works
- [x] New and existing unit tests pass locally with my changes
- [ ] I updated CHANGELOG.md if applicable (N/A — release-please
generates it from the PR title; changelog-guard forbids manual edits)

## Screenshots (if applicable)

N/A — no UI change.

## Additional Notes

The estimation fallback is sticky for the process lifetime (per-model
cache + the loader's fail-fast set from #956): a network that recovers
mid-session keeps estimation until restart. That matches the registry's
existing behavior, so no new divergence is introduced.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: JD Davis <jd@jds-macbook-air.tail2a279.ts.net>
2026-08-12 00:20:16 -05:00
..
__init__.py Initial commit: Headroom SDK - LLM context optimization toolkit 2026-01-06 23:16:58 -08:00
test_anthropic.py fix(providers/anthropic): don't crash token estimation on null tool_calls (#2472) 2026-08-12 00:13:14 -05:00
test_cohere.py Prepare for OSS release v0.2.0 2026-01-07 11:36:44 -08:00
test_deepseek.py test(pricing): stop asserting DeepSeek pricing freshness on wall-clock time (#2428) 2026-07-19 22:14:34 -07:00
test_openai.py fix(providers/openai): bound tiktoken vocab loads with the guarded loader (#2554) 2026-08-12 00:20:16 -05:00
test_universal.py fix(providers): stop pricing modern content blocks at zero (#2760) 2026-08-03 22:40:38 -07:00