decoding_attention

(★ 48)

Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.

decoding_attention 최신버젼 다운로드

최종 버전 다운로드 (.zip)
// repository documentation