deepseek-v4-cmp170hx
DeepSeek-V4-Flash-0731 on 4x CMP 170HX (sm_80): 98 tok/s decode, ~5300 tok/s prefill. DSpark speculative decoding under pipeline parallelism, which vLLM does not support upstream.
File Explorer
Download Latest Version (.zip)- analyze_accum.py
- analyze_content_vs_degen.py
- analyze_thinking_ab.py
- bench_ceiling.py
- bench_chat_accumulate.py
- bench_chunk_truth.py
- bench_conc_needle.py
- bench_concurrency.py
- bench_decode3.py
- bench_decode_ctx.py
- bench_decode_stream.py
- bench_longctx.py
- bench_needle.py
- bench_prefill.py
- bench_probe.py
- README.md
- Dockerfile.devel
- Dockerfile.fullbuild
- dockerignore.txt
- run-a100.sh
- run-pp-dspark.sh
- 0001-sparse_attn_indexer.patch
- 0002-speculative.patch
- 0003-pp_utils.patch
- 0004-model_runner.patch
- 0005-dspark-utils.patch
- 0005a-prefill-topk-torch-fallback.patch
- 0006-logits-row-chunk.patch
- README.md
- .gitattributes
- .gitignore
- LICENSE
- README.md
- RESULTS.md
- SETTINGS.md
// repository documentation
Was this content helpful?
(0 ratings)
