Commit graph

38 commits

Author SHA1 Message Date
Nick DiZazzo
851888d0b0
fix(ci): unblock prerelease SDK producers (#1124)
* fix(ci): unblock prerelease SDK producers

* test(ci): pin SDK producer regression contracts
2026-07-30 17:10:22 -04:00
Nick DiZazzo
265afa90bb
fix(ci): harden cache writes and readiness cleanup (#1123)
* fix(ci): isolate high-fanout compiler cache writes

* fix(ci): use service signal for readiness cleanup
2026-07-30 11:57:18 -04:00
Nick DiZazzo
d4091455c3
fix(swift): synchronize generated binding checksums (#1122) 2026-07-30 09:51:59 -04:00
Nick DiZazzo
43874233ae
fix(ci): route central workflow changes through SDK validation (#1121) 2026-07-30 07:34:12 -04:00
Nick DiZazzo
a6f6f83c6c
fix(release): keep prereleases out of downstream publishing (#1120)
* fix(release): keep prereleases out of downstream publishing

* fix(ci): budget exhaustive Swift cold builds

* fix(ci): align Windows cache version paths
2026-07-30 06:47:59 -04:00
Nick DiZazzo
f7517f8b77
fix(ci): verify exact Windows cache publication (#1119) 2026-07-30 05:57:31 -04:00
Nick DiZazzo
3e30937ada
ci: compose reusable products and add Depot routing (#1113)
* ci: compose reusable products and add Depot routing

* ci: configure job-local sccache storage

* fix: harden Windows artifact composition

* test: assert pinned nightly artifact action

* ci: allow superseded SDK smokes to cancel

* docs: document cancellable CI fan-in gates

* ci: harden composable build graph and metrics

* fix(ci): install actionlint from verified release

* fix(installer): normalize runtime digest paths

* fix(ci): align installer contract output

* docs(ci): ground runner image migration plan

* ci: route trusted ARM lanes through Depot selector

* ci: enable remote sccache for fast lanes

* fix(ci): await remote sccache writes

* fix(ci): align sccache policy contract

* refactor(ci): reuse typed SDK and static ABI inputs

* fix(ci): reuse configured sccache server

* fix(ci): harden exact native cache reuse

* fix(ci): make PR compiler cache read-only

* fix(ci): isolate pull request compiler writes

* feat(ci): produce immutable Node addon artifacts

* fix(ci): restrict Depot canary to main

* fix(ci): isolate Depot canary cache keys
2026-07-30 04:02:53 -04:00
Nick DiZazzo
44ed9aa130
fix: make release composition self-contained (#1112) 2026-07-29 14:16:44 -04:00
Nick DiZazzo
abe6e2fdb8
ci: make composed product readiness hermetic (#1111) 2026-07-29 13:27:04 -04:00
Nick DiZazzo
cf54577738
fix: restore release host executable bit (#1110)
Restore executable permissions after Unix host artifacts are downloaded, covering every Unix release composition lane.
2026-07-29 12:55:22 -04:00
Nick DiZazzo
bd45d4a98e
fix: smoke composed release product (#1108) 2026-07-29 12:17:05 -04:00
Nick DiZazzo
1a2cd8d260
fix: invoke host dependency verifier with Python
- invoke the non-executable verifier through Python in Unix release jobs
- cover neutral x86_64, ARM64, and macOS host producers
- add regression coverage for the release workflow
2026-07-29 12:03:17 -04:00
Nick DiZazzo
ed41f366dc
feat: unify release hosts and native runtimes
- build one backend-neutral host per platform and package native runtimes separately
- compose immutable product-v2 bundles from verified host and runtime artifacts
- enforce host import policy and runtime provenance across release and SDK lanes
- align debug, release, Windows, Kotlin, and Swift validation with the composed product model
- update release documentation, CI topology, and agent guidance for the unified path
- require no-device client readiness and bounded clean shutdown for packaged runtimes
2026-07-29 10:16:23 -04:00
Nick DiZazzo
ed286b909d
feat(runtime): add daemon model lifecycle reconciliation (#1082)
* Add daemon-managed runtime model lifecycle

Introduce persistent runtime lifecycle reconciliation with authenticated owner controls, profile-aware load, unload, ensure, and drain semantics, activity-based admission and priority policy, and additive gossip/protocol support.

Expose the lifecycle through config, CLI, UI, and management APIs; consolidate shared owner-control protocol handling; and add cross-platform build and QA coverage for CUDA setup, mixed versions, process teardown, and SDK/platform paths.
2026-07-27 19:37:50 -04:00
Nick DiZazzo
2804e1f078 fix nightly stability Qwen thinking 2026-07-21 22:56:49 -04:00
Nick DiZazzo
cf2d6addad
fix: Windows Vulkan runtime dependencies (#1046)
* Fix Windows Vulkan runtime dependencies

* Fix Windows dependency test path

* Address Windows GPU routing review

* Normalize verifier path on Windows

* Use Git Bash in Windows verifier test
2026-07-21 21:21:52 -04:00
James Dumay
2c5dacf212
Pipeline MTP-anchored n-gram verify windows (#938)
* Replace Skippy verify span with verify windows

* Add verify window reply metadata

* Pipeline direct-return n-gram verify windows

* Pipeline MTP-anchored n-gram verify windows

* Support static release builds without features

* Fix split MTP activation-frame serving

* Replace native MTP batched verifier with verify windows

* Restore native MTP verify window batching

* Replace native MTP anchor extension with composite proposals

* Keep composite decode branch on development version

* Expose decode timings for all generation modes

* Retry transient staged lane readiness

* Bound persistent lane readiness handshake

* Keep pure N-gram decode free of MTP drafts

* Report composite proposal totals in decode timings

* Gate composite decode pipeline by candidate depth

* Account direct GGUF MTP weights in split planning

* Avoid MTP cooldown after N-gram tail rejection

* Improve hybrid MTP verification telemetry

* Pipeline native MTP verification replies

* Require useful N-gram tails for hybrid MTP

* Adapt N-gram MTP extensions to tail acceptance

* Fix direct GGUF planning fallback

* Gate N-gram tails on MTP prefix agreement

* Widen initial async verify windows

* Restore anchored N-gram MTP extensions

* Retain ready stages across transient refresh failures

* Document pipelined VerifyWindow decode

* Use llama.cpp N-gram proposer for Skippy

* Add cache-based N-gram proposer

* Add declarative speculative proposer package schema

* Productize Skippy speculative decode plans

* Productize Skippy speculative decode plans

* Support direct N-gram Skippy plans

* Validate speculative package strategy plans

* Add coding agent loop benchmark corpus

* Expose Skippy speculative benchmark counters

* Validate cache N-gram proposer limits

* Document speculative decode configuration

* Fix native MTP proposals and fused restore routing

* Honor configured N-gram extension width

* Document speculative runtime overrides

* Refresh speculative config schema contracts

* Keep N-gram tail rejects from penalizing MTP

* Preserve MTP state after serial tail rejects

* Report adaptive verify width changes accurately

* Fix short simple N-gram extension budgets

* Make VerifyWindow pipelining cost-aware

* Profile prospective VerifyWindow widths

* docs: WAN split performance model + measured latency/compute decomposition

Adds docs/skippy/WAN_SPLIT_PERF.md: the single-stream per-token cost model
(TPOT ~= C_total + (S-1)*2*RTT + (S-1)*P), compute-bound vs latency-bound
criteria, when adding a stage helps (memory, concurrency/pipeline overlap,
dense compute-bound models), and speculation as the WAN amortization lever.

Backed by 2026-07-18 Sydney<->Melbourne 2-node measurements: solo 12.9 ms/tok
compute, split 57.8 ms/tok, decomposing to 12.9 compute + 40 (2xRTT) + 4.9
protocol. Workload was latency-bound (~78% network).

* docs: plan for fast-fail on new requests routed to a dead split stage

Documents the measured ~30s hang when a new request routes to a killed
split stage, the confirmed root cause (60s heartbeat / lenient failure
threshold + slow lane-open timeouts), and a two-layer fix (short
steady-state lane-open deadline; feed lane failures into target_health
cooldown) plus an explicit validation gate. Mesh-timing changes are out
of scope pending live multi-node validation.

* Fast-fail lane reconnects so new requests don't hang on a dead split stage

When a downstream split stage dies, a new request would open a fresh lane
and wait the full ~20s warmup ready-deadline before erroring (observed as a
~30s hang in the Sydney<->Melbourne kill test). The 20s deadline is only
needed during pool warmup, when the downstream may still be loading its
model.

Split the deadline: pool warmup keeps LANE_READY_READ_TIMEOUT (20s); mid-life
reconnects from checkout()/replace_lane() on an already-serving mesh use a
new LANE_STEADY_CONNECT_TIMEOUT (3s). A healthy peer answers in milliseconds,
so a dead stage now fails in ~3s instead of ~30s.

Restores receive_persistent_lane_ready as the shared bounded-handshake helper
(dropped during the main merge) and removes a now-obsolete retry test that
covered pre-#1011 retry behavior. Adds tests asserting the steady-state
deadline stays well under the warmup deadline and that the handshake read
fails fast on a silent downstream.

* docs: latency-aware placement — current behaviour and many-node gaps

Records verified planner behaviour (skippy-coordinator/topology.rs,
skippy-topology, host-runtime call site):
- latency is a placement cost (rtt_ms penalty), not just relay-only exclusion
- planner selects a node subset; does not have to use every eligible node
- stage count is gated on a decode-TPOT target (shallower-that-meets beats
  deeper-that-does-not)

And the gaps that matter at many-node scale:
- no peer-to-peer RTT matrix in production (edge_signals never wired; only
  coordinator-RTT is used) -> co-located nodes cannot be exploited
- network estimate is max(coordinator RTT) x node_count, a worst-case proxy
- no first-class prefer-fewer/never-place-above-Y policy beyond the TPOT gate

* docs: measured speculative recovery cost over WAN (why ngram hurts a latency-bound split)

* Discard dead pooled stage lanes before reuse (fast-fail improvement)

A pooled downstream lane whose stage died while checked in was a dead TCP
stream; reusing it blocked the next generation read forever (handshake
read-timeout is cleared for pooled lanes so long generations don't truncate).
checkout() now probes lane liveness with a nonblocking peek and discards a
dead lane so it reconnects with the short steady-state deadline instead of
hanging.

Validated on a loopback 2-node split with a mid-flight worker kill: new
request now fails faster than main (60s vs main's 90s baseline). Does not
fully solve the recovered-local routing path, tracked as follow-up.

---------

Co-authored-by: Michael Neale <14976+michaelneale@users.noreply.github.com>
2026-07-19 20:14:59 +10:00
Eric Wendland
3567b7ea74
fix: Windows installer null architecture probe (#968)
* Avoid calling ToString() on a null RuntimeInformation OSArchitecture value under Windows PowerShell 5.1. 
* Fall back to PROCESSOR_ARCHITECTURE while retaining the x64-only installation guard, and add a regression test for the unsafe probe.
2026-07-18 13:23:39 -04:00
James Dumay
67839e770b
Validate installer bundles before replacing binaries (#957) 2026-07-11 10:11:00 +10:00
Nick DiZazzo
36a44b9188
feature: Refactor setup-first installer flow (#933)
* add uninstall command
* parse checksums without awk intervals
2026-07-10 07:13:20 -04:00
James Dumay
7f6ffef93a
Rely on cargo publish for crate status (#935) 2026-07-01 09:57:01 +10:00
Nick DiZazzo
d798eb651c
fix: RC5 release readiness corrections (#918)
* fix: RC5 testing corrections
* Add missing runtimes
* Add signing attestation key
* Correct smoke script
* Fix `doctor --json` output

* fix: sign linux arm64 cuda release bundles

* fix: address PR review comments

* fix: stop synthesizing fallback GPU ordinals
2026-06-29 05:04:41 -04:00
Nick DiZazzo
e2410a8e5b
fix: validate release native runtimes from explicit matrix (#917)
* fix: validate release native runtime targets explicitly

* fix: handle explicit native runtime target validation

* test: close release manifest fixture before validation
2026-06-28 22:01:24 -04:00
Nick DiZazzo
9968323ee0
fix(ci): nightly stability run (#914) 2026-06-28 16:07:23 -04:00
Nick DiZazzo
a3bab4fe44
fix: Release candidate fixes (#912)
* Fix runtime CLI help surfaces
* Fix hardware profile CUDA detection
* Fail closed without native runtime
* Fix CLI runtime status details
* Validate native runtime release matrix
* Fix UI pnpm workspace metadata
* Treat unsigned release footer tails as missing
* Add RC release smoke script
* Satisfy hardware profile clippy gate
* Move entrypoint tests after runtime items
* Update CLI docs for runtime help
2026-06-28 13:35:19 -04:00
Nick DiZazzo
f1c6acec84
fix(cli): fix gpu command to restore stderr output (#844)
* preserve fatal output without manager

* extract fatal output routing
2026-06-14 07:26:07 -04:00
James Dumay
4a3dce6664
fix installer fallback for runtime releases (#785) 2026-06-04 07:48:31 +10:00
James Dumay
d5612cb06f
fix installer runtime compatibility (#784) 2026-06-04 07:39:25 +10:00
Ivan Golovach
76bc52c926
Harden crates.io publish preflight (#745)
Harden crates.io publish preflight

Validation
* Validation tier: Tier 4 - release/publish workflow hardening refreshed on current main; workflow_dispatch publish preflight now prepares the requested release version before consistency and dry-run validation, with post-review clippy/Linux corrections included.
* git fetch --no-tags origin main:refs/remotes/origin/main: PASS, origin/main at 95695f9f7d.
* git rebase origin/main: PASS, no conflicts.
* git rev-list --left-right --count origin/main...HEAD: PASS, 0 behind / 2 ahead.
* git diff --check origin/main...HEAD: PASS, no output.
* git diff --check: PASS, no output.
* git diff --cached --check: PASS, no output.
* cargo fmt --all -- --check: PASS.
* cargo check -p xtask: PASS.
* cargo run -p xtask -- repo-consistency publish-crates: PASS.
* cargo run -p xtask -- repo-consistency release-targets: PASS.
* cargo run -p xtask -- repo-consistency ci-crate-lists: PASS.
* python3 -m unittest scripts.tests.test_publish_crates: PASS, 9 passed.
* cargo test -p mesh-llm-system hardware:: --lib: PASS, 67 passed.
* cargo test -p mesh-llm-gpu-bench --lib: PASS, 0 passed.
* cargo test -p mesh-llm-node --lib: PASS, 9 passed.
* cargo clippy -p mesh-llm-system -p mesh-llm-gpu-bench -p xtask --all-targets -- -D warnings: PASS.
* cargo publish --dry-run --locked --allow-dirty -p mesh-llm-node: PASS.
* scripts/publish-crates.sh --dry-run --allow-dirty --sleep-seconds 0: PASS, first five crates verified; downstream crates reported expected local dry-run dependency availability limits and script exited 0.
* Remote PR Quality Checks on f743be85b5: PASS.
* Remote PR Builds on f743be85b5: PASS.
* Ledger: not applicable - not required for selected validation tier/change family.
* Version: not applicable - release tooling/packageability guard and compile-hygiene correction only; no release/version sync required.
* Not run: full release workflow locally - not required for selected validation tier; targeted xtask/script/package dry-run checks cover the changed release path and mandatory PR CI passed on the final SHA.
* Not run: just build - not required for selected validation tier; no UI assets or release bundle changed.

Rollback
* git revert HEAD
2026-05-29 23:48:43 -07:00
Ivan Golovach
95695f9f7d
Add KV overlap tool-loop certification (#756)
Validation
* Validation tier: Tier 4 - #652 verification tooling for Goose/Pi-shaped overlapping title/tool request pressure; no runtime, protocol, workflow, or CI gate changes.
* git fetch --no-tags origin main:refs/remotes/origin/main: PASS, origin/main at 3be3869d3d.
* git rebase origin/main: PASS, no conflicts.
* git diff --check origin/main...HEAD: PASS, no output.
* git diff --check: PASS, no output.
* git diff --cached --check: PASS, no output.
* python3 -m unittest scripts.tests.test_qa_kv_tool_loop_stability: PASS, 18 passed.
* python3 -m unittest scripts.tests.test_qa_kv_tool_loop_stability scripts.tests.test_qa_agent_tool_call_reliability: PASS, 25 passed.
* python3 -m py_compile scripts/qa-kv-tool-loop-stability.py scripts/tests/test_qa_kv_tool_loop_stability.py: PASS.
* python3 scripts/qa-kv-tool-loop-stability.py --print-plan --models Qwen/Qwen2.5-3B-Instruct-GGUF:q4_k_m --attempts 1 --pressure-turns 2 --overlap-requests 3 --timeout 30 --min-cached-tokens 128 --suffix-prefill-limit 64 --native-log /tmp/skippy-native.log --output-dir target/kv-tool-loop-stability/overlap-plan | python3 -m json.tool >/tmp/mesh-kv-overlap-plan.json: PASS, emitted valid JSON plan.
* Ledger: not applicable - not required for selected validation tier/change family.
* Version: not applicable - verification tooling only; no release/version sync required.
* Not run: live #652 direct-model overlap certification - no local loaded direct-model endpoint was available; deterministic unit and plan coverage cover the new harness branches.
* Not run: cargo check/tests - not required for selected validation tier; no Rust/runtime code changed.

Rollback
* git revert HEAD
2026-05-29 23:13:02 -07:00
James Dumay
fa490cf558
bump version to 0.68.0 (#711) 2026-05-27 18:38:56 +10:00
Ivan Golovach
eb6b39a5f4
Add KV/tool-loop stability certification harness (#676)
Summary
Adds a repeatable KV/tool-loop stability certification harness for direct-model agent/tool-call pressure runs, including plan output, manifest evidence, transcript handling, native-log checkpoint scanning, docs, tests, and a repo-local agent skill.

Validation
* Validation tier: Tier 4 - verification tooling and testing documentation for KV/tool-loop stability certification; no runtime, protocol, workflow, or CI gate change.
* git fetch --no-tags origin main:refs/remotes/origin/main: PASS, origin/main at 12aa611de5.
* git rebase origin/main: PASS, rebased cleanly after #695/#696.
* git diff --check origin/main...HEAD: PASS, no output
* git diff --check: PASS, no output
* git diff --cached --check: PASS, no output
* python3 -m unittest scripts.tests.test_qa_kv_tool_loop_stability: PASS, 16 passed
* python3 -m unittest discover -s scripts/tests: PASS, 43 passed
* python3 -m py_compile scripts/qa-kv-tool-loop-stability.py scripts/tests/test_qa_kv_tool_loop_stability.py: PASS
* mkdir -p target/kv-tool-loop-stability && python3 scripts/qa-kv-tool-loop-stability.py --models auto,mesh --attempts 1 --pressure-turns 2 --timeout 30 --min-cached-tokens 128 --suffix-prefill-limit 64 --output-dir target/kv-tool-loop-stability/review-smoke --print-plan > /tmp/kv-tool-loop-plan.json && python3 -m json.tool /tmp/kv-tool-loop-plan.json >/dev/null: PASS
* Remote PR checks after rebase: PASS, changes + summary green; non-applicable build/test matrices skipped by path filters.
* Review/merge state: APPROVED, MERGEABLE, CLEAN.
* Ledger: not applicable - not required for selected validation tier/change family.
* Version: not applicable - verification tooling/docs only; no release/version sync required.
* Not run: cargo check/tests - not required for selected validation tier; no Rust/runtime code changed.
* Not run: live direct-model Skippy certification run - no local loaded direct-model endpoint was available; deterministic unit and plan-smoke coverage proves the harness behavior.

Rollback
* Revert this PR.
2026-05-26 16:38:54 -07:00
Ivan Golovach
12aa611de5
Harden crates.io publish retries (#695)
Validation
* Validation tier: Tier 4 - release publish verification tooling and crates.io partial-publish recovery for issue #691.
* git fetch --no-tags origin main:refs/remotes/origin/main: PASS, origin/main at 5c805ead34.
* git diff --check: PASS, no output
* git diff --cached --check: PASS, no output
* bash -n scripts/publish-crates.sh: PASS
* scripts/publish-crates.sh --allow-dirty: PASS, rejected because --allow-dirty requires --dry-run
* python3 -m unittest scripts.tests.test_publish_crates: PASS, 7 passed
* python3 -m unittest discover -s scripts/tests: PASS, 27 passed
* cargo run -p xtask -- repo-consistency release-targets: PASS
* CARGO_TARGET_DIR=$(mktemp -d /tmp/mesh-llm-publish-target.XXXXXX) scripts/publish-crates.sh --dry-run --allow-dirty: PASS; dry-run verified publishable crates through model-artifact and deferred downstream crates whose 0.66.0 registry dependencies are not published yet.
* Ledger: not applicable - not required for selected validation tier/change family.
* Version: not applicable - release publish tooling/docs only; no release version sync required.
* Not run: actionlint - not installed locally and workflow file was not changed.
* Not run: real cargo publish / just release - intentionally avoided for PR validation.

Rollback
* git revert HEAD
2026-05-27 09:28:49 +10:00
James Dumay
cedf9b66e2
fix Swift XCFramework release env (#678) 2026-05-26 06:07:29 +10:00
Nick DiZazzo
5bcacc197c
feature(configuration-ui): integrate configuration UI settings for the local host (#572) 2026-05-24 21:40:21 -04:00
Ivan Golovach
472105cc13
Add repeatable new-model split onboarding (#648)
Add repeatable new-model split onboarding

Validation
* Validation tier: Tier 4 - verification tooling, workflow dispatch inputs, and Skippy split-certification onboarding docs for issue #630, plus Tier 0 post-review docs-only clarification for Command A+ certification status.
* git fetch --no-tags origin main:refs/remotes/origin/main: PASS, origin/main at c7948d7181.
* git diff --check: PASS, no output
* git diff --cached --check: PASS, no output
* python3 -m unittest discover -s scripts/tests: PASS, 17 passed
* python3 scripts/skippy-llama-parity.py validate: PASS, no output
* ruby -ryaml -e 'YAML.load_file(ARGV.fetch(0)); puts "ok"' .github/workflows/queue-unsloth-layer-packages.yml: PASS, ok
* cargo fmt --all -- --check: PASS
* cargo check -p model-package --bin queue-unsloth-layer-packages: PASS
* cargo test -p model-package --lib: PASS, 8 passed
* Ledger: not applicable - not required for selected validation tier/change family.
* Version: not applicable - onboarding/tooling/docs change only; no release/version sync required.
* Not run: actionlint - not installed locally.
* Not run: real HF Jobs submission or package publication - intentionally avoided for PR validation.

Rollback
* git revert HEAD
2026-05-24 13:18:26 -07:00
Ivan Golovach
421cc08667
Add nightly mesh stability harness (#631) 2026-05-22 13:38:52 -04:00
Ivan Golovach
b5a16d68d9
Add agent tool-call reliability harness (#623)
Validation
* Validation tier: Tier 4 — verification tooling plus docs and contract tests for the OpenAI-compatible agent tool-call QA path.
* git fetch --no-tags origin main:refs/remotes/origin/main: PASS
* git diff --check: PASS, no output
* git diff --cached --check: PASS, no output
* python3 -m unittest discover -s scripts/tests -p 'test_*.py': PASS, 7 passed
* python3 -m py_compile scripts/qa-agent-tool-call-reliability.py scripts/tests/test_qa_agent_tool_call_reliability.py: PASS
* scripts/qa-agent-tool-call-reliability.py --models auto,mesh --attempts 2 --print-plan >/tmp/mesh-agent-tool-plan.json && python3 -m json.tool /tmp/mesh-agent-tool-plan.json >/dev/null: PASS
* cargo fmt --all -- --check: PASS
* LLAMA_STAGE_BUILD_DIR=/Users/Funtland/Downloads/mesh-llm/.deps/llama-build/build-stage-abi-metal cargo test -p mesh-llm --test qa_agent_tool_call_reliability: PASS, 4 passed
* Ledger: not applicable — not required for selected validation tier/change family.
* Version: not applicable — not required for selected validation tier/change family.
* Not run: live endpoint probe — not required before opening this opt-in QA harness PR; local parser, plan, and script contract tests cover the new tooling surface.
* Not run: just build — not required for selected validation tier; no shipped runtime, UI bundle, or release artifact changed.

Rollback
* git revert HEAD
2026-05-22 11:37:37 +10:00