headroom/tests/test_cli/test_wrap_stale_marker.py
Tejas Chopra f27f235032
fix(wrap): stop concurrent wrap sessions clobbering settings.local.json (#3232)
## Description

Several `headroom wrap` sessions in one project each write the proxy URL
into
`.claude/settings.local.json` and restore it on exit. That
read-modify-write was
unsynchronised. The write itself is atomic so the file never tears, but
the
updates were still lost against each other:

- **Live sessions were silently unrouted.** The first session to exit
deleted the
key while its siblings were still running. They kept working, but their
traffic
  stopped going through the proxy — no error, no warning, no savings.
- **A dead proxy was written back into the project.** A session that
started
second captured the *first* session's proxy URL as "the original", so
its exit
restored a URL pointing at a port that was already gone. Every later
session in
  that project then failed to connect.
- **SIGTERM/SIGHUP never ran the restore at all.** `cleanup` was
registered as the
handler, but a Python signal handler that returns normally does not
unwind the
stack — under PEP 475 the interrupted `waitpid` is simply retried. The
`finally`
block that restores `settings.local.json` never ran, while the handler
had
  already terminated the proxy underneath a child that was still alive.

Closes #3205

## Type of Change

- [x] Bug fix (non-breaking change that fixes an issue)

## Changes Made

- **`_wrap_settings_lock`** — an exclusive OS lock (flock /
`msvcrt.locking`) held
across the settings read-modify-write. A workspace that cannot hold lock
state
  degrades to the previous behaviour rather than failing, matching
  `_proxy_start_lock`.
- **`.headroom_wrap_owners.json`** — a sidecar recording, per env key,
the true
pre-wrap `original` plus the live sessions holding it. The first writer
records
the original; later writers inherit it and are flagged `inherited`, so
no
session restores a value it did not observe first-hand. A session exits
without
restoring while a sibling still holds the key. Dead holders are pruned
with the
same conservative PID+identity liveness the proxy-client markers use, so
a
  SIGKILLed session cannot wedge the key.
- **`unwrap` passes `force=True`** — unwrap is the user explicitly
asking for
their settings back, so it drops every claim instead of deferring to a
live
sibling and silently printing success while leaving the proxy URL in the
file.
- **The #2221 self-heal passes `dead_ports`** — a wrapper process can
outlive its
proxy (proxy alone SIGKILLed). Its claim would otherwise veto the
self-heal and
  leave `ANTHROPIC_BASE_URL` pointing at a port just proven dead.
- **`_rehome_wrap_marker`** — the wrap marker has one slot, won by the
last
writer. When that writer exits while a sibling still owns the key, the
marker is
rewritten to describe the survivor (carrying the record's true
original), so the
survivor keeps its #2221 self-heal record instead of being left with a
marker
  describing a dead process.
- **`_exit_on_signal`** replaces `cleanup` as the SIGTERM/SIGHUP
handler. Raising
`SystemExit` unwinds, so the settings restore actually runs and cleanup
happens
  exactly once from `finally`.
- **`_proxy_start_lock` now shares `_locked_file`** with the new
settings lock
  rather than carrying a second verbatim copy of the platform branches.

## Testing

- [x] Unit tests pass (`pytest`) — full suite, 11518 passed / 588
skipped
- [x] Linting passes (`ruff check .`)
- [x] Type checking passes (`mypy headroom`)
- [x] New tests added for new functionality
- [x] Manual testing performed

`tests/test_wrap_concurrent_settings.py` (14 tests) covers: a sibling
exit leaving
survivors routed, the last session out restoring the true original, a
pre-existing
user URL surviving the whole cycle, three sessions in every exit order,
a crashed
session not wedging the key, forced unwrap past a live session, a holder
that
outlived its proxy not vetoing the self-heal, marker rehoming, and the
signal-handler unwind.

### Test Output

```text
$ uv run pytest tests/test_cli/test_unwrap_claude.py tests/test_cli/test_wrap_claude_base_url.py \
    tests/test_cli/test_wrap_claude_finally_unbound.py tests/test_cli/test_wrap_claude_vertex_proxy_env.py \
    tests/test_cli/test_wrap_claude.py tests/test_cli/test_wrap_dead_marker_selfheal.py \
    tests/test_cli/test_wrap_helpers.py tests/test_cli/test_wrap_stale_marker.py \
    tests/test_cli/test_wrap_persistent.py tests/test_wrap_concurrent_settings.py tests/test_cli_doctor.py -q

tests/test_wrap_concurrent_settings.py ..............                    [ 72%]
tests/test_cli_doctor.py ............................................... [ 89%]
...............................                                          [100%]

============================= 285 passed in 3.01s ==============================

$ uv run pytest tests/ -q
======== 11518 passed, 588 skipped, 6036 warnings in 1831.34s (0:30:31) ========

$ uv run ruff check .
All checks passed!

$ uv run mypy headroom
Success: no issues found in 527 source files
```

## Real Behavior Proof

- **Environment:** macOS 15 (Darwin 25.4.0), Python 3.12.13, repo venv,
Claude
  provider path (`ANTHROPIC_BASE_URL` in `.claude/settings.local.json`).
- **Exact command / steps:** a script spawning **two real OS processes**
— no
  mocks, real PIDs, real files — that call the same
`_write_claude_wrap_base_url` / `_restore_claude_wrap_base_url` helpers
`wrap claude` uses. The project starts with a real user gateway already
set.
Session A (port 8787) starts, session B (port 8788) starts 0.7s later, A
exits
while B is still running, then B exits. Run identically on `main` and on
this
  branch.

**Before (on `main`) — both bugs visible:**

```text
start                       : {"ANTHROPIC_BASE_URL": "https://my-gateway.example.com"}
  session port=8787 started, remembers previous='https://my-gateway.example.com'
  session port=8788 started, remembers previous='http://127.0.0.1:8787'
both sessions running       : {"ANTHROPIC_BASE_URL": "http://127.0.0.1:8788"}
  session port=8787 exited
after FIRST session exits   : {"ANTHROPIC_BASE_URL": "https://my-gateway.example.com"}
  session port=8788 exited
after LAST session exits    : {"ANTHROPIC_BASE_URL": "http://127.0.0.1:8787"}
```

Session B is still running, but after A exits the proxy URL is gone from
under it
— B is unrouted with no error. And the final state is
`http://127.0.0.1:8787`: a
dead proxy left permanently in the user's project, with their real
gateway lost.

**After (this branch):**

```text
start                       : {"ANTHROPIC_BASE_URL": "https://my-gateway.example.com"}
  session port=8787 started, remembers previous='https://my-gateway.example.com'
  session port=8788 started, remembers previous='http://127.0.0.1:8787'
both sessions running       : {"ANTHROPIC_BASE_URL": "http://127.0.0.1:8788"}
  session port=8787 exited
after FIRST session exits   : {"ANTHROPIC_BASE_URL": "http://127.0.0.1:8788"}
  session port=8788 exited
after LAST session exits    : {"ANTHROPIC_BASE_URL": "https://my-gateway.example.com"}
```

B stays routed after A exits, and the last session out restores the
user's real
gateway.

- **Observed result:** matches the intent on both counts — no unrouting,
no dead
  proxy residue, user's pre-existing URL preserved.
- **Not tested:** Windows (`msvcrt.locking`) — the lock and dead-holder
pruning
  are exercised on POSIX only; the Windows branch is the same code path
`_proxy_start_lock` has shipped with. No live end-to-end run against a
real
Anthropic endpoint with two concurrent `claude` CLIs; the proof above
drives the
same helpers out of two real processes instead. Foundry/Vertex key
variants are
covered by unit tests, not by a live run. Real SIGTERM/SIGHUP delivery
to a
running `wrap claude` was not exercised end to end — the handler's
unwind is
covered by a unit test, and full signal delivery would need a spawned
and
  killed subprocess, which the existing #1768 test also declined to do.

## Runtime Rollout Safety

- **Rollout-managed feature(s):** none — this is an unconditional
correctness fix
  on the wrap settings path.
- **Minimum rollout channel:** n/a.
- **Stable/default behavior changed:** yes, three ways. (1) A wrap
session exiting
while a sibling holds the key now leaves the key in place instead of
removing
  it. (2) SIGTERM/SIGHUP now unwinds, so the child CLI is terminated by
  `subprocess.run`'s cleanup rather than being left running against a
  torn-down proxy. (3) Two new sidecar files appear next to
`settings.local.json`: `.headroom_wrap_owners.json` (removed when the
last
holder exits) and `.headroom_wrap_settings.lock` (retained by design —
deleting
  a live lock file creates an inode-replacement race).
- **Kill switch / disable path:** none. A workspace where the lock file
cannot be
created degrades to the previous unsynchronised behaviour automatically.
- **Unsafe override required:** no.
- **Qualification impact:** none beyond the wrap settings path.
- **Rollback path:** revert the commit; the sidecar files are ignored by
older
  versions and can be deleted safely.

## Review Readiness

- [x] I have performed a self-review
- [x] This PR is ready for human review

## Checklist

- [x] My code follows the project's style guidelines
- [x] I have performed a self-review of my code
- [x] I have commented my code, particularly in hard-to-understand areas
- [x] I have made corresponding changes to the documentation
- [x] My changes generate no new warnings
- [x] I have added tests that prove my fix is effective or that my
feature works
- [x] New and existing unit tests pass locally with my changes
- [x] I did **not** edit `CHANGELOG.md`

## Additional Notes

- The ownership record is keyed per env key, so `ANTHROPIC_BASE_URL`,
the
Foundry/Vertex variants and the tool-search entry are tracked
independently.
- Documentation: the behaviour is documented in the helper docstrings
rather than
user-facing docs — the sidecar files are internal state a user never
configures.
- Follow-up worth considering: `.headroom_wrap_settings.lock` is
intentionally
never deleted (matching `_proxy_start_lock`'s retention rationale), so
it stays
in `.claude/` after `unwrap`. Removing it safely needs a separate think
about
  the inode-replacement race.

Co-authored-by: Tejas Chopra <tejas@Tejass-MacBook-Pro.local>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-23 22:36:17 -07:00

70 lines
2.7 KiB
Python

from __future__ import annotations
import json
import signal
from pathlib import Path
import pytest
from headroom.cli import doctor as doctor_cli
from headroom.cli import wrap as wrap_cli
def _settings(tmp_path: Path) -> Path:
return tmp_path / ".claude" / "settings.local.json"
def test_doctor_skips_with_no_marker(tmp_path: Path) -> None:
result = doctor_cli.check_wrap_marker_staleness(_settings(tmp_path))
assert result.status == doctor_cli.SKIP
def test_doctor_passes_with_live_marker(tmp_path: Path) -> None:
path = _settings(tmp_path)
wrap_cli._write_claude_wrap_base_url("http://127.0.0.1:8787", settings_path=path, port=8787)
result = doctor_cli.check_wrap_marker_staleness(path)
assert result.status == doctor_cli.PASS
def test_doctor_flags_stale_wrap_marker(tmp_path: Path) -> None:
path = _settings(tmp_path)
wrap_cli._write_claude_wrap_base_url("http://127.0.0.1:8787", settings_path=path, port=8787)
marker_path = wrap_cli._wrap_marker_path(path)
marker = json.loads(marker_path.read_text(encoding="utf-8"))
marker["pid"] = 999_999_999
marker_path.write_text(json.dumps(marker), encoding="utf-8")
result = doctor_cli.check_wrap_marker_staleness(path)
assert result.status == doctor_cli.WARN
assert "999999999" in result.summary
assert "headroom unwrap claude" in result.summary
def test_claude_command_registers_sighup_next_to_sigterm() -> None:
"""`claude()` must catch SIGHUP (terminal close) the same way it catches
SIGTERM, or a crashed-by-terminal-close wrap session never restores its
base_url (issue #1768). Full signal delivery isn't practical to exercise
via CliRunner (would require spawning/killing a real subprocess), so this
asserts the registration is present in claude()'s source, guarded for
platforms without SIGHUP.
"""
import inspect
src = inspect.getsource(wrap_cli.claude.callback)
assert 'hasattr(signal, "SIGHUP")' in src
assert "signal.signal(signal.SIGHUP, _exit_on_signal)" in src
assert "signal.signal(signal.SIGTERM, _exit_on_signal)" in src
def test_signal_handler_unwinds_so_the_restore_can_run() -> None:
"""Registering `cleanup` directly never achieved what #1768 wanted.
A Python signal handler that returns normally does not unwind the stack --
under PEP 475 the interrupted `waitpid` is simply retried -- so the finally
block that restores settings.local.json never ran, while the handler had
already torn the proxy down under a live child. The handler must raise.
"""
with pytest.raises(SystemExit) as excinfo:
wrap_cli._exit_on_signal(signal.SIGHUP, None)
assert excinfo.value.code == 128 + signal.SIGHUP