LLM:LightningLM 0.1V — Reference training pipeline for the LightningLM family. 2B dense seed → 5B MoE → 9B MoE → 120B sparse MoE through TurboQuant-PreTraining on a single eight-GPU node. Companion code for *Reversible Foundations*.
BatchTopK:Implementation of the BatchTopK activation function for training sparse autoencoders (SAEs)
Sparse-vLLM:A sparse-first inference engine (sparsevllm). It also contains DeltaKV compressor training + evaluation tooling (deltakv).
2by4-pretrain:Efficient 2:4 sparse training algorithms and implementations
Sparsity-Win-Robust-Generalization:[ICLR 2022] "Sparsity Winning Twice: Better Robust Generalization from More Efficient Training" by Tianlong Chen*, Zhenyu Zhang*, Pengjun Wang*, Santosh Balachandra*, Haoyu Ma*, Zehao Wang, Zhangyang Wang
// repository documentation
Was this content helpful?
★ 0(0 ratings)
Recent Feedback
Download README
Do you want to download the README.md file for e2e_sae?