MoE-Mixture-of-Experts-in-PyTorch

(★ 83)

Implementations of a Mixture-of-Experts (MoE) architecture designed for research on large language models (LLMs) and scalable neural network designs. One implementation targets a **single-device/NPU environment** while the other is built for multi-device distributed computing. Both versions showcase the core principles.

MoE-Mixture-of-Experts-in-PyTorch Latest Version Download

Download Latest Version (.zip)
// repository documentation