Publish verified Windows CUDA release sidecars - #37
Draft
leehack wants to merge 4 commits into
Draft
Conversation
10 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
ggml-base.dllSHA-256assets.jsonprovenance flowThe dependent Dart consumer is leehack/llamadart#340. Runtime downloading is intentionally out of scope and tracked in leehack/llamadart#339.
Why
The source-built CUDA lane dominated Windows release time even though upstream llama.cpp already publishes matching CUDA binaries. The previous experiment proved that the exact b10453 assets can be repackaged and loader-smoked without a GPU in about four minutes; this change moves that measured path into the owning native release workflow while preserving an exact compatibility boundary with the locally built core.
Validation
actionlint -shellcheck=passedgit diff --checkpassedPublication gates
This PR remains draft. No native release has been published and no physical NVIDIA GPU was available. Before publication, validate representative CUDA 12/13 hardware and real-model inference, including kernel/PTX execution and the open MTP acceptance, CUDA lockup, DFlash concurrency, DSpark VRAM-leak, EAGLE, and failed-state-restore reports.
Tracked by #64.