deepseek-v4-cmp170hx
DeepSeek-V4-Flash-0731 on 4x CMP 170HX (sm_80): 98 tok/s decode, ~5300 tok/s prefill. DSpark speculative decoding under pipeline parallelism, which vLLM does not support upstream.
// repository documentation
Was this content helpful?
(0 ratings)
