Multi-Head Attention, Transformer, Perceiver, Linear Attention.
Explore Similar Repositories
DiJiang:[ICML'24 Oral] The official code of "DiJiang: Efficient Large Language Models through Compact Kernelization", a novel DCT-based linear attention mechanism.
mad:Code for "Online and Linear Time Attention by Enforcing Monotonic Alignments"
protein-localization:Using Transformer protein embeddings with a linear attention mechanism to make SOTA de-novo predictions for the subcellular location of proteins :microscope:
// repository documentation
Was this content helpful?
★ 0(0 ratings)
Recent Feedback
Download README
Do you want to download the README.md file for transformer-blocks?