* Fix Windows Vulkan runtime dependencies
* Fix Windows dependency test path
* Address Windows GPU routing review
* Normalize verifier path on Windows
* Use Git Bash in Windows verifier test
Closes#971. Replace the nonexistent Homebrew tap instructions, document public packages and images, and route release dispatches to Mesh-LLM/mesh-packaging.
* Replace Skippy verify span with verify windows
* Add verify window reply metadata
* Pipeline direct-return n-gram verify windows
* Pipeline MTP-anchored n-gram verify windows
* Support static release builds without features
* Fix split MTP activation-frame serving
* Replace native MTP batched verifier with verify windows
* Restore native MTP verify window batching
* Replace native MTP anchor extension with composite proposals
* Keep composite decode branch on development version
* Expose decode timings for all generation modes
* Retry transient staged lane readiness
* Bound persistent lane readiness handshake
* Keep pure N-gram decode free of MTP drafts
* Report composite proposal totals in decode timings
* Gate composite decode pipeline by candidate depth
* Account direct GGUF MTP weights in split planning
* Avoid MTP cooldown after N-gram tail rejection
* Improve hybrid MTP verification telemetry
* Pipeline native MTP verification replies
* Require useful N-gram tails for hybrid MTP
* Adapt N-gram MTP extensions to tail acceptance
* Fix direct GGUF planning fallback
* Gate N-gram tails on MTP prefix agreement
* Widen initial async verify windows
* Restore anchored N-gram MTP extensions
* Retain ready stages across transient refresh failures
* Document pipelined VerifyWindow decode
* Use llama.cpp N-gram proposer for Skippy
* Add cache-based N-gram proposer
* Add declarative speculative proposer package schema
* Productize Skippy speculative decode plans
* Productize Skippy speculative decode plans
* Support direct N-gram Skippy plans
* Validate speculative package strategy plans
* Add coding agent loop benchmark corpus
* Expose Skippy speculative benchmark counters
* Validate cache N-gram proposer limits
* Document speculative decode configuration
* Fix native MTP proposals and fused restore routing
* Honor configured N-gram extension width
* Document speculative runtime overrides
* Refresh speculative config schema contracts
* Keep N-gram tail rejects from penalizing MTP
* Preserve MTP state after serial tail rejects
* Report adaptive verify width changes accurately
* Fix short simple N-gram extension budgets
* Make VerifyWindow pipelining cost-aware
* Profile prospective VerifyWindow widths
* docs: WAN split performance model + measured latency/compute decomposition
Adds docs/skippy/WAN_SPLIT_PERF.md: the single-stream per-token cost model
(TPOT ~= C_total + (S-1)*2*RTT + (S-1)*P), compute-bound vs latency-bound
criteria, when adding a stage helps (memory, concurrency/pipeline overlap,
dense compute-bound models), and speculation as the WAN amortization lever.
Backed by 2026-07-18 Sydney<->Melbourne 2-node measurements: solo 12.9 ms/tok
compute, split 57.8 ms/tok, decomposing to 12.9 compute + 40 (2xRTT) + 4.9
protocol. Workload was latency-bound (~78% network).
* docs: plan for fast-fail on new requests routed to a dead split stage
Documents the measured ~30s hang when a new request routes to a killed
split stage, the confirmed root cause (60s heartbeat / lenient failure
threshold + slow lane-open timeouts), and a two-layer fix (short
steady-state lane-open deadline; feed lane failures into target_health
cooldown) plus an explicit validation gate. Mesh-timing changes are out
of scope pending live multi-node validation.
* Fast-fail lane reconnects so new requests don't hang on a dead split stage
When a downstream split stage dies, a new request would open a fresh lane
and wait the full ~20s warmup ready-deadline before erroring (observed as a
~30s hang in the Sydney<->Melbourne kill test). The 20s deadline is only
needed during pool warmup, when the downstream may still be loading its
model.
Split the deadline: pool warmup keeps LANE_READY_READ_TIMEOUT (20s); mid-life
reconnects from checkout()/replace_lane() on an already-serving mesh use a
new LANE_STEADY_CONNECT_TIMEOUT (3s). A healthy peer answers in milliseconds,
so a dead stage now fails in ~3s instead of ~30s.
Restores receive_persistent_lane_ready as the shared bounded-handshake helper
(dropped during the main merge) and removes a now-obsolete retry test that
covered pre-#1011 retry behavior. Adds tests asserting the steady-state
deadline stays well under the warmup deadline and that the handshake read
fails fast on a silent downstream.
* docs: latency-aware placement — current behaviour and many-node gaps
Records verified planner behaviour (skippy-coordinator/topology.rs,
skippy-topology, host-runtime call site):
- latency is a placement cost (rtt_ms penalty), not just relay-only exclusion
- planner selects a node subset; does not have to use every eligible node
- stage count is gated on a decode-TPOT target (shallower-that-meets beats
deeper-that-does-not)
And the gaps that matter at many-node scale:
- no peer-to-peer RTT matrix in production (edge_signals never wired; only
coordinator-RTT is used) -> co-located nodes cannot be exploited
- network estimate is max(coordinator RTT) x node_count, a worst-case proxy
- no first-class prefer-fewer/never-place-above-Y policy beyond the TPOT gate
* docs: measured speculative recovery cost over WAN (why ngram hurts a latency-bound split)
* Discard dead pooled stage lanes before reuse (fast-fail improvement)
A pooled downstream lane whose stage died while checked in was a dead TCP
stream; reusing it blocked the next generation read forever (handshake
read-timeout is cleared for pooled lanes so long generations don't truncate).
checkout() now probes lane liveness with a nonblocking peek and discards a
dead lane so it reconnects with the short steady-state deadline instead of
hanging.
Validated on a loopback 2-node split with a mid-flight worker kill: new
request now fails faster than main (60s vs main's 90s baseline). Does not
fully solve the recovered-local routing path, tracked as follow-up.
---------
Co-authored-by: Michael Neale <14976+michaelneale@users.noreply.github.com>
* Avoid calling ToString() on a null RuntimeInformation OSArchitecture value under Windows PowerShell 5.1.
* Fall back to PROCESSOR_ARCHITECTURE while retaining the x64-only installation guard, and add a regression test for the unsafe probe.
* publish install runbook at meshllm.cloud/setup-mesh
* add setup-mesh self-reference and absolute doc links in install runbook
* add goal routing, filter installed-model noise, trim issue links
- Add a goal->section routing table so single-machine goals (e.g. local coding
agent) take a short path instead of the full multi-node mesh runbook.
- Instruct assistants to filter `models installed` output down to runnable
models, excluding layer-package internals and split shards, and show at most
three candidates. Workaround for Mesh-LLM/mesh-llm#1004.
- Collapse the GitHub issue reference list into inline pointers; the operational
lessons are already folded into the relevant sections.
The catalog described sub-4B models (Llama-3.2-3B, Hermes-2-Pro) as strong
tool-calling / goose defaults, but they fall apart in real agentic harnesses
(goose, Fizz): wrong paths, repeated failing calls, hallucinated tools, and
leaked raw tool-call text. Meanwhile the Gemma 4 family survives the same
harness cleanly and stays snappy.
Update descriptions to reflect observed harness behavior: temper the sub-4B
agentic claims and mark Gemma-4-E4B as a good non-reasoning mini-class agent
default. Descriptions are mirrored in both embedded catalogs.
Co-authored-by: Michael Neale <14976+michaelneale@users.noreply.github.com>
* Complete plugin web UI host surface
* Align plugin web UI contract and docs
* Harden plugin web UI release path
* Fix plugin navigation and settings UX
* Align local and CI repo consistency gates
* Run publish consistency in test-all
* Simplify native log test module wiring
* Skip re-hashing unchanged GGUF sources with a sidecar hash cache
* Harden sidecar hash cache against malformed records and path edge cases(review addressed)
* Validate ctime in the sidecar hash cache to catch same-size replacements that restore mtime(review addressed)
* Clean up temp files when sidecar cache writes fail(coderabbit review addressed)
---------
Co-authored-by: James Dumay <jameswdumay@gmail.com>
* fix: use GGUF file stems as display names for synthetic local-gguf refs
Synthetic local-gguf/sha256-... refs leaked into user-facing model
names on the dashboard and in model target views. Preserve a
filename-derived label in the local inventory snapshot for synthetic
keys and resolve display names as catalog name, then local label, then
raw ref. The canonical model key is unchanged. Fixes#882
* test: cover synthetic model display-name wiring
* test: cover Windows synthetic GGUF labels
---------
Co-authored-by: James Dumay <jameswdumay@gmail.com>