Skip to content

Benchmark results for 4× Tesla V100-SXM2-32GB, CUDA backend, q2-q4-imatrix DeepSeek V4 Flash GGUF #603

Description

@asasasasasbc

GPUs: 4× Tesla V100-SXM2-32GB (sm_70, 128 GB VRAM total)
CUDA: 12.9 / Driver 580.159.03
Model: DeepSeek V4 Flash q2-q4-imatrix (98 GB, 43 layers, 256 experts)
Tensor Parallel: 4-way, half-resident expert split

Canonical sweep from the speed-bench README: ctx 2048..65536, step 2048, 128 greedy decode tokens per frontier.

Commit 0a7ad77

Headlines: 122.2 t/s prefill / 30.1 t/s gen @2k, 108.1 t/s prefill / 21.0 t/s gen @65k.

CSV is in the current ds4-bench output format (includes the gen_first_ms / steady-state columns); plot_speed.py renders it as-is.

v100x4_q2q4_bench.csv

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions