Skip to content
View tonyd2wild's full-sized avatar

Sponsoring

@rationalsa

Block or report tonyd2wild

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Popular repositories Loading

  1. DeepSeek-v4-Flash-DSpark-1M-NVFP4-KV-2x-DGX-Spark DeepSeek-v4-Flash-DSpark-1M-NVFP4-KV-2x-DGX-Spark Public

    DeepSeek V4 Flash DSpark 1M NVFP4 KV recipe for 2x DGX Spark

    Python 157 11

  2. GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s Public

    Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster

    Python 70 8

  3. Deepseek-v4-Flash-TP2-DGX-Spark-500k-CTX Deepseek-v4-Flash-TP2-DGX-Spark-500k-CTX Public

    Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — Docker image, launch scripts, RDMA/NCCL setup, and the gotchas.

    Shell 57 10

  4. MiniMax-M3-2x-DGX-Spark-36-tok-s MiniMax-M3-2x-DGX-Spark-36-tok-s Public

    MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three serving lanes: speed / balanced / long-context.

    40 3

  5. MiMo-V2.5-TP2-1M-NVFP4-KV-2xDGX-Spark MiMo-V2.5-TP2-1M-NVFP4-KV-2xDGX-Spark Public

    MiMo-V2.5 Omni TP=2 on 2x DGX Spark · 1M context · NVFP4 4-bit KV (~1.97M-token KV pool @ 1M, ~30 tok/s) · 69-eval: thinking-OFF 97.8 beats thinking-ON 90.6 for tool/agent work

    Python 37 4

  6. deepseek-v4-flash-2x-spark-1m deepseek-v4-flash-2x-spark-1m Public

    DeepSeek V4 Flash @ 1M token context on 2x NVIDIA DGX Spark — production-tested recipe (45 tok/s decode, real 800K prompts served)

    Shell 36 3