* Isolate grammar state during speculative verification
* Align serial grammar verification test
* fix(skippy): trim sampled verification test window
* test(skippy): cover long-context tool verification
Return VerifyWindow MTP drafts through the typed ABI, advance the Skippy ABI mirror, and update Rust callers and tests. Harden ngram/KV state handling, GLM phase and range gates, and Metal dispatch configuration and active-count routing.
Keep the later split wk_b/wv_b graph contract, retain dynamic environment reads required by scoped tests, and avoid a duplicate Metal builder registry because those findings are superseded by later queue policy and existing symbol compilation.
* Enable adaptive verify window for ngram/draft speculation
The adaptive verify window was never enabled on the split-serving path:
to_embedded_openai_args hardcoded adaptive_speculative_window = false. With a
fixed window, an early reject never shrank the window, so a sustained reject
storm kept proposing at full depth and paying the full 2-round-trip recovery
cost per token. On a WAN split this measured as ~40% throughput loss with
N-gram speculation ON versus OFF, despite high per-token acceptance.
Enable the adaptive window whenever speculation actually proposes a window
(ngram or draft mode). The existing shrink_adaptive_window logic then narrows
the window toward the observed accept depth after an early reject, cutting
recovery frequency. Adds a regression test asserting ngram speculation turns
the adaptive window on.
* Replace speculative rollback with positional MTP n-gram pipelining
* Pipeline speculative verify windows across latency
* Fix positional correction and adaptive pipeline depth
* Continuously refill the speculative horizon
* Productionize pipelined MTP n-gram speculation
* Fix speculative docs and UI formatting
* Remove stale speculative projections and fix CI
* Handle fragmented direct-return fallback replies
* Replace speculative repair with fixed-depth positional pipeline
* Expose split-stage compute overlap telemetry
* Lock split topology placement
* Document locked split topology
* Address locked topology review feedback
* Fix SPEED-Bench timing JSONL output
* Bound benchmark telemetry finalization
* Hash SPEED-Bench request and response pairs
* mesh: stop re-applying formation-time RTT gate to operational stage streams
Ported from the WAN lab branch (wip/wan-direct-prediction-return, c340f741),
where it was validated live on a ~26ms WAN split. open_stage_transport_stream
re-applied the formation-time MAX_SPLIT_RTT_MS ceiling to every fresh
operational stream, so per-request direct-return sinks were rejected under
normal WAN RTT jitter while pooled forward lanes stayed healthy - surfacing as
ready-handshake timeouts and 502s on an already-admitted split. Split
admission still gates eligibility via gossiped, hysteresis-smoothed RTT plus
re-election; operational streams now warn and proceed.
* skippy: raise return-sink ready timeout 5s->20s for cold WAN bridge setup
Ported from the WAN lab branch (46108cfc). Over a WAN mesh the return sink
connects to a local bridge alias, but the remote ready byte only arrives after
the bridge cold-establishes a fresh stage QUIC connection (~10s budget) and the
remote handler dials its local server. 5s timed out during that cold setup on
a healthy ~26ms split; forward lanes already use a 20s budget. Match it.
* runtime: relaunch withdrawn splits when peers return instead of ending the model task
Observed live on a real WAN split (Sydney M5 <-> AU 4090): one transient
direct-return 502 led periodic_check to mark the remote stage unavailable;
after the 75s grace the coordinator withdrew the topology. The Withdraw event
returned StartupLoopControl::Break, so startup_local_model_loop tore down and
the task ended permanently - while the remote worker sat healthy, logging
'standing by for stage assignment' forever. Only recovery was manually
restarting both nodes with a fresh token.
Make withdraw non-terminal: a new RelaunchSplit control/outcome runs the full
existing teardown, then loops back to the launch phase and re-enters
wait_for_split_participants, relaunching the split when an eligible peer
returns. The stop channel is checked before relaunch so explicit shutdown
still wins. LocalFallback (model fits locally) is unchanged.
The participant-wait loop's 30s cadence and stable-participant gating act as
the natural retry throttle; no extra backoff added.
---------
Co-authored-by: Michael Neale <14976+michaelneale@users.noreply.github.com>
Co-authored-by: Michael Neale <michael.neale@gmail.com>
* Replace Skippy verify span with verify windows
* Add verify window reply metadata
* Pipeline direct-return n-gram verify windows
* Pipeline MTP-anchored n-gram verify windows
* Support static release builds without features
* Fix split MTP activation-frame serving
* Replace native MTP batched verifier with verify windows
* Restore native MTP verify window batching
* Replace native MTP anchor extension with composite proposals
* Keep composite decode branch on development version
* Expose decode timings for all generation modes
* Retry transient staged lane readiness
* Bound persistent lane readiness handshake
* Keep pure N-gram decode free of MTP drafts
* Report composite proposal totals in decode timings
* Gate composite decode pipeline by candidate depth
* Account direct GGUF MTP weights in split planning
* Avoid MTP cooldown after N-gram tail rejection
* Improve hybrid MTP verification telemetry
* Pipeline native MTP verification replies
* Require useful N-gram tails for hybrid MTP
* Adapt N-gram MTP extensions to tail acceptance
* Fix direct GGUF planning fallback
* Gate N-gram tails on MTP prefix agreement
* Widen initial async verify windows
* Restore anchored N-gram MTP extensions
* Retain ready stages across transient refresh failures
* Document pipelined VerifyWindow decode
* Use llama.cpp N-gram proposer for Skippy
* Add cache-based N-gram proposer
* Add declarative speculative proposer package schema
* Productize Skippy speculative decode plans
* Productize Skippy speculative decode plans
* Support direct N-gram Skippy plans
* Validate speculative package strategy plans
* Add coding agent loop benchmark corpus
* Expose Skippy speculative benchmark counters
* Validate cache N-gram proposer limits
* Document speculative decode configuration
* Fix native MTP proposals and fused restore routing
* Honor configured N-gram extension width
* Document speculative runtime overrides
* Refresh speculative config schema contracts
* Keep N-gram tail rejects from penalizing MTP
* Preserve MTP state after serial tail rejects
* Report adaptive verify width changes accurately
* Fix short simple N-gram extension budgets
* Make VerifyWindow pipelining cost-aware
* Profile prospective VerifyWindow widths
* docs: WAN split performance model + measured latency/compute decomposition
Adds docs/skippy/WAN_SPLIT_PERF.md: the single-stream per-token cost model
(TPOT ~= C_total + (S-1)*2*RTT + (S-1)*P), compute-bound vs latency-bound
criteria, when adding a stage helps (memory, concurrency/pipeline overlap,
dense compute-bound models), and speculation as the WAN amortization lever.
Backed by 2026-07-18 Sydney<->Melbourne 2-node measurements: solo 12.9 ms/tok
compute, split 57.8 ms/tok, decomposing to 12.9 compute + 40 (2xRTT) + 4.9
protocol. Workload was latency-bound (~78% network).
* docs: plan for fast-fail on new requests routed to a dead split stage
Documents the measured ~30s hang when a new request routes to a killed
split stage, the confirmed root cause (60s heartbeat / lenient failure
threshold + slow lane-open timeouts), and a two-layer fix (short
steady-state lane-open deadline; feed lane failures into target_health
cooldown) plus an explicit validation gate. Mesh-timing changes are out
of scope pending live multi-node validation.
* Fast-fail lane reconnects so new requests don't hang on a dead split stage
When a downstream split stage dies, a new request would open a fresh lane
and wait the full ~20s warmup ready-deadline before erroring (observed as a
~30s hang in the Sydney<->Melbourne kill test). The 20s deadline is only
needed during pool warmup, when the downstream may still be loading its
model.
Split the deadline: pool warmup keeps LANE_READY_READ_TIMEOUT (20s); mid-life
reconnects from checkout()/replace_lane() on an already-serving mesh use a
new LANE_STEADY_CONNECT_TIMEOUT (3s). A healthy peer answers in milliseconds,
so a dead stage now fails in ~3s instead of ~30s.
Restores receive_persistent_lane_ready as the shared bounded-handshake helper
(dropped during the main merge) and removes a now-obsolete retry test that
covered pre-#1011 retry behavior. Adds tests asserting the steady-state
deadline stays well under the warmup deadline and that the handshake read
fails fast on a silent downstream.
* docs: latency-aware placement — current behaviour and many-node gaps
Records verified planner behaviour (skippy-coordinator/topology.rs,
skippy-topology, host-runtime call site):
- latency is a placement cost (rtt_ms penalty), not just relay-only exclusion
- planner selects a node subset; does not have to use every eligible node
- stage count is gated on a decode-TPOT target (shallower-that-meets beats
deeper-that-does-not)
And the gaps that matter at many-node scale:
- no peer-to-peer RTT matrix in production (edge_signals never wired; only
coordinator-RTT is used) -> co-located nodes cannot be exploited
- network estimate is max(coordinator RTT) x node_count, a worst-case proxy
- no first-class prefer-fewer/never-place-above-Y policy beyond the TPOT gate
* docs: measured speculative recovery cost over WAN (why ngram hurts a latency-bound split)
* Discard dead pooled stage lanes before reuse (fast-fail improvement)
A pooled downstream lane whose stage died while checked in was a dead TCP
stream; reusing it blocked the next generation read forever (handshake
read-timeout is cleared for pooled lanes so long generations don't truncate).
checkout() now probes lane liveness with a nonblocking peek and discards a
dead lane so it reconnects with the short steady-state deadline instead of
hanging.
Validated on a loopback 2-node split with a mid-flight worker kill: new
request now fails faster than main (60s vs main's 90s baseline). Does not
fully solve the recovered-local routing path, tracked as follow-up.
---------
Co-authored-by: Michael Neale <14976+michaelneale@users.noreply.github.com>