flash-llm
Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
파일 탐색기
최종 버전 다운로드 (.zip)- Makefile
- SpMM_API.cuh
- AsyncCopy_PTX.cuh
- MatMulUtilities.cuh
- MMA_PTX.cuh
- Reduction_Kernel.cuh
- SpMM_API.cu
- SpMM_Kernel.cuh
- TilingConfig.h
- ExistingSpMM.png
- Inference_OPT_175B.png
- Inference_OPT_66B.png
- KernelBenchmarking.png
- MatMulsInLLMs.png
- 1_Preparations.md
- 2_KernelBenchmarking.md
- 3_LLMInferenceExample.md
- inference-test.py
- requirements.txt
- utils.py
- huggingface_opt_convert_Phase1.py
- huggingface_opt_convert_Phase2.py
- benchmark.sh
- Log2Excel.py
- Makefile
- profiling.sh
- sparTA.h
- spmm_test.cu
- spmm_test_utils.h
- sputnik_utils.h
- test_env
- FasterTransformer
- ft.patch
- sputnik
- sputnik.patch
- .clang-format
- .gitignore
- .gitmodules
- Init_FlashLLM.sh
- LICENSE
- README.md
// repository documentation
Was this content helpful?
(0 ratings)
