Rearchitecting-LLMs
Official code for the Manning book on structural LLM optimization: depth/width pruning, knowledge distillation, and attention optimization, runnable on free Colab GPUs.
// repository documentation
Was this content helpful?
(0 ratings)
