Deepseek-v4-Flash-TP2-DGX-Spark-500k-CTX:Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — Docker image, launch scripts, RDMA/NCCL setup, and the gotchas.
dspark-vllm-gx10:Two-node DGX Spark/ASUS GX10 DeepSeek V4 Flash DSpark NVFP4 port for vLLM 0.25, with live dashboard and reproducible deployment.
ds4-ssd:DeepSeek V4 Flash specific inference engine. SSD MoE expert paging (slot-bank) + disk KV cache for long agent sessions. Metal-first, narrow & high-performance.
deepseek-v4-cmp170hx:DeepSeek-V4-Flash-0731 on 4x CMP 170HX (sm_80): 98 tok/s decode, ~5300 tok/s prefill. DSpark speculative decoding under pipeline parallelism, which vLLM does not support upstream.
// repository documentation
Was this content helpful?
★ 0(0 ratings)
Recent Feedback
Download README
Do you want to download the README.md file for deepseek-v4-flash-sm120?