3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
-
Updated
Sep 5, 2026 - Python
3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
Run an isolated Claude Science app copy through local or OpenAI-compatible model backends.
Beautiful, zero-dependency realtime dashboard + live activity log for a local MTPLX inference server — reads the /metrics endpoint, no build step.
Recover the native MTP predictor missing from the 8-bit MLX Qwen3.8-27B-Uncensored package, build a BF16 sidecar, and reproduce a 15.59 → 48.75 tok/s controlled M4 Max result with MTPLX.
Double Qwen 3.8 27B inference speed on Apple Silicon with a one-click local coding agent
To associate your repository with the mtplx topic, visit your repo's landing page and select "manage topics."