Extract diagnostic decode policy (#1981)

## Description

Extracts lossy diagnostic byte decoding from `helpers.py` into
`headroom.proxy.diagnostic_decode_policy`. Protocol parsers stay strict
while the diagnostic/logging path has a dedicated, directly tested
policy.

Closes #

## Type of Change

- [ ] Bug fix (non-breaking change that fixes an issue)
- [ ] New feature (non-breaking change that adds functionality)
- [ ] Breaking change (fix or feature that would cause existing
functionality to change)
- [ ] Documentation update
- [ ] Performance improvement
- [x] Code refactoring (no functional changes)

## Changes Made

- Added `diagnostic_decode_policy.py` for UTF-8 diagnostic decoding with
replacement characters.
- Kept `helpers.safe_decode_for_logging` delegating to the extracted
policy for existing callers.
- Added direct tests for valid UTF-8, invalid byte replacement, max-byte
truncation, and helper delegation.
- Carried forward the LiteLLM callback compatibility shim needed for
current mypy on `main`.

## Testing

- [x] Unit tests pass (`pytest`)
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [ ] Manual testing performed

### Test Output

```text
python -m pytest tests\test_diagnostic_decode_policy.py
4 passed in 0.18s

python -m ruff check .
All checks passed!

python -m ruff format --check .
1095 files already formatted

python -m mypy headroom --ignore-missing-imports
Success: no issues found in 409 source files

gitleaks protect --staged --no-banner --redact
no leaks found
```

## Real Behavior Proof

- Environment: Windows, Python 3.13.13, branch
`jd/architecture-slice-30`.
- Exact command / steps: ran focused diagnostic decode policy tests,
ruff, ruff format check, mypy, and staged gitleaks scan.
- Observed result: diagnostic decode behavior is directly covered and
local lint/type/security checks pass.
- Not tested: live upstream error responses; existing helper import path
remains intact.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [ ] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [ ] I have updated the CHANGELOG.md if applicable

## Screenshots (if applicable)

N/A

## Additional Notes

Documentation and changelog updates are N/A for this internal
architecture-only refactor. The push reported existing default-branch
Dependabot alerts; no staged secret leaks were found for this PR.

---------

Co-authored-by: Tejas Chopra <chopratejas@gmail.com>
This commit is contained in:
JD Davis 2026-07-12 16:17:39 +00:00 committed by GitHub
parent 41ce14bd64
commit 7fb9209089
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
3 changed files with 39 additions and 10 deletions

View file

@ -0,0 +1,17 @@
"""Lossy byte decoding policy for diagnostics and logs."""
from __future__ import annotations
import codecs
def safe_decode_for_logging(raw: bytes, *, max_bytes: int | None = None) -> str:
"""Decode bytes to a string for log/diagnostic display only.
Wire/protocol parsers should decode complete protocol frames strictly. This
policy is for already-discarded diagnostics where replacement characters are
preferable to failing the error-reporting path.
"""
blob = raw[:max_bytes] if max_bytes is not None else raw
decoder = codecs.getincrementaldecoder("utf-8")(errors="replace")
return decoder.decode(bytes(blob), final=True)

View file

@ -25,6 +25,7 @@ from typing import TYPE_CHECKING, Any, Literal, cast
from headroom import paths as _paths
from headroom._subprocess import run
from headroom.proxy import (
diagnostic_decode_policy,
request_limit_policy,
sse_byte_buffer_policy,
wire_debug_format_policy,
@ -565,16 +566,7 @@ def safe_decode_for_logging(raw: bytes, *, max_bytes: int | None = None) -> str:
Use ``parse_sse_events_from_byte_buffer`` for SSE parsing instead.
"""
blob = raw[:max_bytes] if max_bytes is not None else raw
# Decode incrementally and represent any invalid bytes as the
# Unicode replacement character (<28>). Implemented via the
# `codecs` incremental decoder so we never reach for the
# forbidden `errors="ignore"`/`errors="replace"` keyword in the
# SSE-bearing modules.
import codecs as _codecs
decoder = _codecs.getincrementaldecoder("utf-8")(errors="replace")
return decoder.decode(bytes(blob), final=True)
return diagnostic_decode_policy.safe_decode_for_logging(raw, max_bytes=max_bytes)
def parse_sse_events_from_byte_buffer(

View file

@ -0,0 +1,20 @@
from __future__ import annotations
from headroom.proxy.diagnostic_decode_policy import safe_decode_for_logging
from headroom.proxy.helpers import safe_decode_for_logging as helper_safe_decode_for_logging
def test_safe_decode_for_logging_decodes_utf8() -> None:
assert safe_decode_for_logging("hello \u2603".encode()) == "hello \u2603"
def test_safe_decode_for_logging_replaces_invalid_bytes() -> None:
assert safe_decode_for_logging(b"ok\xffdone") == "ok\ufffddone"
def test_safe_decode_for_logging_honors_max_bytes_before_decoding() -> None:
assert safe_decode_for_logging(b"abcdef", max_bytes=3) == "abc"
def test_helpers_safe_decode_delegates_to_policy() -> None:
assert helper_safe_decode_for_logging(b"ok\xffdone") == safe_decode_for_logging(b"ok\xffdone")