GLM-5.2-655K-MTP-4x-DGX-Spark---25-32tok-s

(★ 28)

Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 655,360-token context + MTP spec decode (DCP4, fp8_ds_mla) on a 4x NVIDIA DGX Spark (GB10) cluster

GLM-5.2-655K-MTP-4x-DGX-Spark---25-32tok-s Latest Version Download

Download Latest Version (.zip)
// repository documentation