Deepseek-v4-Flash-TP2-DGX-Spark-500k-CTX
Working recipe to serve DeepSeek-V4-Flash across two NVIDIA DGX Spark (GB10) nodes with vLLM (TP=2, FP8 KV, MTP) over a RoCE/RDMA link — Docker image, launch scripts, RDMA/NCCL setup, and the gotchas.
// repository documentation
Was this content helpful?
(0 ratings)
