decoding_attention
Decoding Attention is specially optimized for MHA, MQA, GQA and MLA using CUDA core for the decoding stage of LLM inference.
decoding_attention Latest Version Download
Download Latest Version (.zip)// repository documentation
Was this content helpful?
(0 ratings)
