One-flag setup for local Qwen3.8: builds a CUDA llama.cpp with HauhauCS FastMTP speculative decoding, downloads model/template/sidecar, serves on localhost — PowerShell & bash.
-
Updated
Aug 25, 2026 - Shell
One-flag setup for local Qwen3.8: builds a CUDA llama.cpp with HauhauCS FastMTP speculative decoding, downloads model/template/sidecar, serves on localhost — PowerShell & bash.
Benchmarks and serving notes for Qwen3.5-27B fine-tune on Apple Silicon via MLX. Companion to dgx-spark-nvfp4-serving.
To associate your repository with the fastmtp topic, visit your repo's landing page and select "manage topics."