deepseek-v4-cmp170hx

(โ˜… 89)

DeepSeek-V4-Flash-0731 on 4x CMP 170HX (sm_80): 98 tok/s decode, ~5300 tok/s prefill. DSpark speculative decoding under pipeline parallelism, which vLLM does not support upstream.

  • .gitattributes
  • .gitignore
  • LICENSE
  • README.md
  • RESULTS.md
  • SETTINGS.md
// repository documentation