GPUs: 4× Tesla V100-SXM2-32GB (sm_70, 128 GB VRAM total)
CUDA: 12.9 / Driver 580.159.03
Model: DeepSeek V4 Flash q2-q4-imatrix (98 GB, 43 layers, 256 experts)
Tensor Parallel: 4-way, half-resident expert split
Canonical sweep from the speed-bench README: ctx 2048..65536, step 2048, 128 greedy decode tokens per frontier.
Commit 0a7ad77
Headlines: 122.2 t/s prefill / 30.1 t/s gen @2k, 108.1 t/s prefill / 21.0 t/s gen @65k.
CSV is in the current ds4-bench output format (includes the gen_first_ms / steady-state columns); plot_speed.py renders it as-is.
v100x4_q2q4_bench.csv
GPUs: 4× Tesla V100-SXM2-32GB (sm_70, 128 GB VRAM total)
CUDA: 12.9 / Driver 580.159.03
Model: DeepSeek V4 Flash q2-q4-imatrix (98 GB, 43 layers, 256 experts)
Tensor Parallel: 4-way, half-resident expert split
Canonical sweep from the speed-bench README: ctx 2048..65536, step 2048, 128 greedy decode tokens per frontier.
Commit 0a7ad77
Headlines: 122.2 t/s prefill / 30.1 t/s gen @2k, 108.1 t/s prefill / 21.0 t/s gen @65k.
CSV is in the current ds4-bench output format (includes the gen_first_ms / steady-state columns); plot_speed.py renders it as-is.
v100x4_q2q4_bench.csv