Speculative-Decoding:Implementation of the paper Fast Inference from Transformers via Speculative Decoding, Leviathan et al. 2023.
mergenetic:Flexible library for merging large language models (LLMs) via evolutionary optimization (ACL 2025 Demo).
REASONING_COMPILER:Official implementation of "REASONING COMPILER: LLM-Guided Optimizations for Efficient Model Serving" (NeurIPS 2025)
acon:Official implementation of paper "ACON: Optimizing Context Compression for Long-horizon LLM Agents"
pytorch-memory-optim:This code repository contains the code used for my "Optimizing Memory Usage for Training LLMs and Vision Transformers in PyTorch" blog post.
// repository documentation
Was this content helpful?
★ 0(0 ratings)
Recent Feedback
Download README
Do you want to download the README.md file for llm-proxy?