llama.cpp
LLM inference in C/C++
File Explorer
Download Latest Version (.zip)Showing a partial file list β download the zip above to see everything.
- apps.nix
- devshells.nix
- docker.nix
- jetson-support.nix
- nixpkgs-instances.nix
- package-gguf-py.nix
- package.nix
- python-scripts.nix
- scope.nix
- sif.nix
- cann.Dockerfile
- cpu.Dockerfile
- cuda.Dockerfile
- intel.Dockerfile
- llama-cli-cann.Dockerfile
- llama-cpp-cuda.srpm.spec
- llama-cpp.srpm.spec
- musa.Dockerfile
- openvino.Dockerfile
- rocm.Dockerfile
- s390x.Dockerfile
- tools.sh
- vulkan.Dockerfile
- zendnn.Dockerfile
- settings.json
- action.yml
- action.yml
- action.yml
- action.yml
- action.yml
- action.yml
- action.yml
- action.yml
- action.yml
- 010-bug-compilation.yml
- 011-bug-results.yml
- 019-bug-misc.yml
- 020-enhancement.yml
- 030-research.yml
- 040-refactor.yml
- config.yml
- ai-issues.yml
- bench.yml.disabled
- build-3rd-party.yml
- build-and-test-snapdragon.yml
- build-android.yml
- build-apple.yml
- build-cache.yml
- build-cann.yml
- build-cmake-pkg.yml
- build-cpu.yml
- build-cross.yml
- build-cuda-ubuntu.yml
- build-cuda-windows.yml
- build-ibm.yml
- build-msys.yml
- build-opencl.yml
- build-openvino.yml
- build-riscv.yml
- build-sanitize.yml
- build-self-hosted.yml
- build-sycl.yml
- build-virtgpu.yml
- build-vulkan.yml
- build-wasm.yml
- build-webgpu.yml
- check-vendor.yml
- close-issue.yml
- code-style.yml
- copilot-setup-steps.yml
- docker.yml
- editorconfig.yml
- gguf-publish.yml
- hip-quality-check.yml
- labeler.yml
- make-release.yml
- pr-draft-label.yml
- pre-tokenizer-hashes.yml
- python-check-requirements.yml
- python-lint.yml
- python-type-check.yml
- release.yml
- server-sanitize.yml
- server-self-hosted.yml
- server.yml
- ui-build-self-hosted.yml
- ui-build.yml
- ui-publish.yml
- ui-self-hosted.yml
- ui.yml
- update-ops-docs.yml
- winget.yml
- labeler.yml
- pull_request_template.md
- SYSTEM.md
- CMakeLists.txt
- download.cpp
- llama.cpp
- aime25_openai__gpt-oss-120b-high_temp1.0_20251109_094547.html
- aime25_openai__gpt-oss-120b-high_temp1.0_20251109_094547.json
- aime25_openai__gpt-oss-120b-high_temp1.0_20251109_094547_allresults.json
- dgx-spark.md
- run-aime-120b-t8-x8-high.log
- mac-m2-ultra.md
- nemotron-dgx-spark.md
- README-MUSA.md
- README.md
- run.sh
- arm64-apple-clang.cmake
- arm64-linux-clang.cmake
- arm64-windows-llvm.cmake
- arm64-windows-msvc-cuda.cmake
- build-info.cmake
- common.cmake
- download-models.cmake
- git-vars.cmake
- license.cmake
- llama-config.cmake.in
- llama.pc.in
- riscv64-spacemit-linux-gnu-gcc.cmake
- x64-windows-llvm.cmake
- caps.cpp
- caps.h
- lexer.cpp
- lexer.h
- parser.cpp
- parser.h
- README.md
- runtime.cpp
- runtime.h
- string.cpp
- string.h
- utils.h
- value.cpp
- value.h
- arg.cpp
- arg.h
- base64.hpp
- build-info.cpp.in
- build-info.h
- chat-auto-parser-generator.cpp
- chat-auto-parser-helpers.cpp
- chat-auto-parser-helpers.h
- chat-auto-parser.h
- chat-diff-analyzer.cpp
- chat-peg-parser.cpp
- chat-peg-parser.h
- chat.cpp
- chat.h
- CMakeLists.txt
- common.cpp
- common.h
- console.cpp
- console.h
- debug.cpp
- debug.h
- download.cpp
- download.h
- fit.cpp
- fit.h
- hf-cache.cpp
- hf-cache.h
- http.h
- imatrix-loader.cpp
- imatrix-loader.h
- json-schema-to-grammar.cpp
- json-schema-to-grammar.h
- llguidance.cpp
- log.cpp
- log.h
- ngram-cache.cpp
- ngram-cache.h
- ngram-map.cpp
- ngram-map.h
- ngram-mod.cpp
- ngram-mod.h
- peg-parser.cpp
- peg-parser.h
- preset.cpp
- preset.h
- reasoning-budget.cpp
- reasoning-budget.h
- sampling.cpp
- sampling.h
- speculative.cpp
- speculative.h
- subproc.cpp
- subproc.h
- trie.cpp
- trie.h
- unicode.cpp
- unicode.h
- __init__.py
- afmoe.py
- arctic.py
- baichuan.py
- bailingmoe.py
- bailingmoe3.py
- base.py
- bert.py
- bitnet.py
- bloom.py
- chameleon.py
- chatglm.py
- codeshell.py
- cogvlm.py
- command_r.py
- dbrx.py
- deci.py
- deepseek.py
- dots1.py
- dotsocr.py
- dream.py
- ernie.py
- exaone.py
- falcon.py
- falcon_h1.py
- gemma.py
- glm.py
- gpt2.py
- gpt_oss.py
- gptneox.py
- granite.py
- grok.py
- grovemoe.py
- hunyuan.py
- internlm.py
- internvl.py
- jais.py
- jamba.py
- januspro.py
- kimi_k3.py
- kimi_linear.py
- kimivl.py
- laguna.py
- lfm2.py
- lighton_ocr.py
- llada.py
- llama.py
- llama4.py
- llava.py
- maincoder.py
- mamba.py
- mellum.py
- mimo.py
- minicpm.py
- minimax.py
- mistral.py
- mistral3.py
- mpt.py
- muse_glimmer.py
- nanbeige.py
- nemotron.py
- olmo.py
- openelm.py
- orion.py
- pangu.py
- phi.py
- pixtral.py
- plamo.py
- plm.py
- pockettts.py
- qwen.py
- qwen3tts.py
- qwen3vl.py
- qwenvl.py
- refact.py
- rwkv.py
- sarashina2.py
- smallthinker.py
- smolvlm.py
- stablelm.py
- starcoder.py
- step3.py
- t5.py
- talkie.py
- ultravox.py
- wavtokenizer.py
- xverse.py
- youtuvl.py
- imported-into-android-studio.jpg
- CMakeUserPresets.json
- developer.md
- linux.md
- README.md
- windows.md
- configuration.md
- development.md
- BLIS.md
- CANN.md
- CUDA-FEDORA.md
- ET.md
- OPENCL.md
- OPENVINO.md
- SYCL.md
- VirtGPU.md
- zDNN.md
- ZenDNN.md
- idea-arch.key
- idea-arch.pdf
- debugging-tests.md
- HOWTO-add-model.md
- parsing.md
- token_generation_performance_tips.md
- gemma3.md
- glmedge.md
- granitevision.md
- llava.md
- minicpmo2.6.md
- minicpmo4.0.md
- minicpmv2.5.md
- minicpmv2.6.md
- minicpmv4.0.md
- minicpmv4.5.md
- minicpmv4.6.md
- MobileVLM.md
- BLAS.csv
- CANN.csv
- CPU.csv
- CUDA.csv
- ET.csv
- Metal.csv
- OpenCL.csv
- SYCL.csv
- Vulkan.csv
- WebGPU.csv
- zDNN.csv
- ZenDNN.csv
- android.md
- autoparser.md
- build-riscv64-spacemit.md
- build-s390x.md
- build.md
- completions.md
- docker.md
- function-calling.md
- install.md
- llguidance.md
- models.md
- multi-gpu.md
- multimodal.md
- ops.md
- preset.md
- release.md
- speculative.md
- xcframework.md
- batched.cpp
- CMakeLists.txt
- README.md
- main.swift
- .gitignore
- Makefile
- Package.swift
- README.md
- CMakeLists.txt
- convert-llama2c-to-ggml.cpp
- README.md
- CMakeLists.txt
- debug.cpp
- README.md
- deprecation-warning.cpp
- README.md
- CMakeLists.txt
- diffusion-cli.cpp
- diffusion.cpp
- diffusion.h
- README.md
- CMakeLists.txt
- embedding.cpp
- README.md
- CMakeLists.txt
- eval-callback.cpp
- README.md
- CMakeLists.txt
- gen-docs.cpp
- CMakeLists.txt
- gguf.cpp
- CMakeLists.txt
- gguf-hash.cpp
- README.md
- CMakeLists.txt
- idle.cpp
- README.md
- llama-eval.py
- llama-server-simulator.py
- README.md
- test-simulator.sh
- MainActivity.kt
- MessageAdapter.kt
- bg_assistant_message.xml
- bg_user_message.xml
- ic_launcher_background.xml
- ic_launcher_foreground.xml
- outline_folder_open_24.xml
- outline_send_24.xml
- activity_main.xml
- item_message_assistant.xml
- item_message_user.xml
- ic_launcher.xml
- ic_launcher_round.xml
- ic_launcher.webp
- ic_launcher_round.webp
- ic_launcher.webp
- ic_launcher_round.webp
- ic_launcher.webp
- ic_launcher_round.webp
- ic_launcher.webp
- ic_launcher_round.webp
- ic_launcher.webp
- ic_launcher_round.webp
- colors.xml
- strings.xml
- themes.xml
- backup_rules.xml
- data_extraction_rules.xml
- AndroidManifest.xml
- .gitignore
- build.gradle.kts
- proguard-rules.pro
- gradle-wrapper.jar
- gradle-wrapper.properties
- libs.versions.toml
- ExampleInstrumentedTest.kt
- ai_chat.cpp
- CMakeLists.txt
- logging.h
- FileType.kt
- GgufMetadata.kt
- GgufMetadataReader.kt
- GgufMetadataReaderImpl.kt
- InferenceEngineImpl.kt
- AiChat.kt
- InferenceEngine.kt
- AndroidManifest.xml
- ExampleUnitTest.kt
- .gitignore
- build.gradle.kts
- consumer-rules.pro
- proguard-rules.pro
- .gitignore
- build.gradle.kts
- gradle.properties
- gradlew
- settings.gradle.kts
- LibLlama.swift
- Contents.json
- Contents.json
- LlamaState.swift
- .gitignore
- ContentView.swift
- DownloadButton.swift
- InputButton.swift
- LoadCustomButton.swift
- llama_swiftuiApp.swift
- IDEWorkspaceChecks.plist
- contents.xcworkspacedata
- project.pbxproj
- .gitignore
- README.md
- CMakeLists.txt
- lookahead.cpp
- README.md
- CMakeLists.txt
- lookup-create.cpp
- lookup-merge.cpp
- lookup-stats.cpp
- lookup.cpp
- README.md
- compare-embeddings-logits.sh
- compare-logits.py
- convert-model.sh
- modelcard.template
- run-casual-gen-embeddings-org.py
- run-converted-model-embeddings-logits.sh
- run-converted-model.sh
- run-org-model.py
- compare-embeddings-logits.sh
- convert-model.sh
- modelcard.template
- run-converted-model.sh
- run-original-model.py
- __init__.py
- check-nmse.py
- common.py
- compare_tokens.py
- create-collection-add-model.sh
- curl-embedding-server.sh
- hf-add-model-to-collection.py
- hf-create-collection.py
- hf-create-model.py
- hf-upload-gguf-model.py
- inspect-converted-model.sh
- inspect-org-model.py
- perplexity-gen.sh
- perplexity-run-simple.sh
- perplexity-run.sh
- quantize.sh
- run-embedding-server.sh
- semantic_check.py
- .gitignore
- Makefile
- README.md
- requirements.txt
- CMakeLists.txt
- parallel.cpp
- README.md
- CMakeLists.txt
- passkey.cpp
- README.md
- CMakeLists.txt
- README.md
- retrieval.cpp
- CMakeLists.txt
- README.md
- simple.cpp
- CMakeLists.txt
- README.md
- simple-chat.cpp
- .gitignore
- CMakeLists.txt
- README.md
- CMakeLists.txt
- README.md
- speculative.cpp
- CMakeLists.txt
- README.md
- speculative-simple.cpp
- build.sh
- CMakeLists.txt
- ls-sycl-device.cpp
- README.md
- run-llama2.sh
- start-svr.sh
- test.sh
- update-ops-doc.sh
- win-build-sycl.bat
- win-run-llama2.bat
- win-start-svr.bat
- win-test.bat
- win-update-ops-doc.bat
- .gitignore
- build-install.sh
- build.sh
- CMakeLists.txt
- README.md
- test-cmake.cpp
- CMakeLists.txt
- finetune.cpp
- README.md
- CMakeLists.txt
- convert_legacy_llama.py
- json_schema_pydantic_example.py
- json_schema_to_grammar.py
- llama.vim
- pydantic_models_to_grammar.py
- pydantic_models_to_grammar_examples.py
- reason-act.sh
- regex_to_grammar.py
- server-llama2-13B.sh
- server_embd.py
- ts-type-to-grammar.sh
- common.cmake
- FindNCCL.cmake
- ggml-config.cmake.in
- GitVars.cmake
- ggml-alloc.h
- ggml-backend.h
- ggml-blas.h
- ggml-cann.h
- ggml-cpp.h
- ggml-cpu.h
- ggml-cuda.h
- ggml-et.h
- ggml-hexagon.h
- ggml-metal.h
- ggml-opencl.h
- ggml-openvino.h
- ggml-opt.h
- ggml-rpc.h
- ggml-sycl.h
- ggml-virtgpu.h
- ggml-vulkan.h
- ggml-webgpu.h
- ggml-zdnn.h
- ggml-zendnn.h
- ggml.h
- gguf.h
- CMakeLists.txt
- ggml-blas.cpp
- acl_tensor.cpp
- acl_tensor.h
- aclnn_ops.cpp
- aclnn_ops.h
- CMakeLists.txt
- common.h
- ggml-cann.cpp
- amx.cpp
- amx.h
- common.h
- mmq.cpp
- mmq.h
- cpu-feats.cpp
- quants.c
- repack.cpp
- quants.c
- cpu-feats.cpp
- quants.c
- cpu-feats.cpp
- quants.c
- repack.cpp
- cpu-feats.cpp
- quants.c
- quants.c
- cpu-feats.cpp
- quants.c
- repack.cpp
- FindSIMD.cmake
- FindSMTIME.cmake
- kernels.cpp
- kernels.h
- kleidiai.cpp
- kleidiai.h
- sgemm.cpp
- sgemm.h
- ime.cpp
- ime.h
- ime1_kernels.cpp
- ime2_kernels.cpp
- ime_env.cpp
- ime_env.h
- ime_kernels.h
- repack.cpp
- repack.h
- rvv_kernels.cpp
- rvv_kernels.h
- spine_barrier.h
- spine_mem_pool.cpp
- spine_mem_pool.h
- spine_tcm.h
- arch-fallback.h
- binary-ops.cpp
- binary-ops.h
- CMakeLists.txt
- common.h
- ggml-cpu-impl.h
- ggml-cpu.c
- ggml-cpu.cpp
- hbm.cpp
- hbm.h
- ops.cpp
- ops.h
- quants.c
- quants.h
- repack.cpp
- repack.h
- simd-gemm.h
- simd-mappings.h
- traits.cpp
- traits.h
- unary-ops.cpp
- unary-ops.h
- vec.cpp
- vec.h
- fattn-mma-f16-instance-ncols1_1-ncols2_16.cu
- fattn-mma-f16-instance-ncols1_1-ncols2_32.cu
- fattn-mma-f16-instance-ncols1_1-ncols2_8.cu
- fattn-mma-f16-instance-ncols1_16-ncols2_1.cu
- fattn-mma-f16-instance-ncols1_16-ncols2_2.cu
- fattn-mma-f16-instance-ncols1_16-ncols2_4.cu
- fattn-mma-f16-instance-ncols1_2-ncols2_16.cu
- fattn-mma-f16-instance-ncols1_2-ncols2_32.cu
- fattn-mma-f16-instance-ncols1_2-ncols2_4.cu
- fattn-mma-f16-instance-ncols1_2-ncols2_8.cu
- fattn-mma-f16-instance-ncols1_32-ncols2_1.cu
- fattn-mma-f16-instance-ncols1_32-ncols2_2.cu
- fattn-mma-f16-instance-ncols1_4-ncols2_16.cu
- fattn-mma-f16-instance-ncols1_4-ncols2_2.cu
- fattn-mma-f16-instance-ncols1_4-ncols2_4.cu
- fattn-mma-f16-instance-ncols1_4-ncols2_8.cu
- fattn-mma-f16-instance-ncols1_64-ncols2_1.cu
- fattn-mma-f16-instance-ncols1_8-ncols2_1.cu
- fattn-mma-f16-instance-ncols1_8-ncols2_2.cu
- fattn-mma-f16-instance-ncols1_8-ncols2_4.cu
- fattn-mma-f16-instance-ncols1_8-ncols2_8.cu
- fattn-tile-instance-dkq112-dv112.cu
- fattn-tile-instance-dkq128-dv128.cu
- fattn-tile-instance-dkq192-dv128.cu
- fattn-tile-instance-dkq256-dv256.cu
- fattn-tile-instance-dkq320-dv256.cu
- fattn-tile-instance-dkq40-dv40.cu
- fattn-tile-instance-dkq512-dv512.cu
- fattn-tile-instance-dkq576-dv512.cu
- fattn-tile-instance-dkq64-dv64.cu
- fattn-tile-instance-dkq72-dv72.cu
- fattn-tile-instance-dkq80-dv80.cu
- fattn-tile-instance-dkq96-dv96.cu
- fattn-vec-instance-bf16-bf16.cu
- fattn-vec-instance-bf16-f16.cu
- fattn-vec-instance-bf16-q4_0.cu
- fattn-vec-instance-bf16-q4_1.cu
- fattn-vec-instance-bf16-q5_0.cu
- fattn-vec-instance-bf16-q5_1.cu
- fattn-vec-instance-bf16-q8_0.cu
- fattn-vec-instance-f16-bf16.cu
- fattn-vec-instance-f16-f16.cu
- fattn-vec-instance-f16-q4_0.cu
- fattn-vec-instance-f16-q4_1.cu
- fattn-vec-instance-f16-q5_0.cu
- fattn-vec-instance-f16-q5_1.cu
- fattn-vec-instance-f16-q8_0.cu
- fattn-vec-instance-q4_0-bf16.cu
- fattn-vec-instance-q4_0-f16.cu
- fattn-vec-instance-q4_0-q4_0.cu
- fattn-vec-instance-q4_0-q4_1.cu
- fattn-vec-instance-q4_0-q5_0.cu
- fattn-vec-instance-q4_0-q5_1.cu
- fattn-vec-instance-q4_0-q8_0.cu
- fattn-vec-instance-q4_1-bf16.cu
- fattn-vec-instance-q4_1-f16.cu
- fattn-vec-instance-q4_1-q4_0.cu
- fattn-vec-instance-q4_1-q4_1.cu
- fattn-vec-instance-q4_1-q5_0.cu
- fattn-vec-instance-q4_1-q5_1.cu
- fattn-vec-instance-q4_1-q8_0.cu
- fattn-vec-instance-q5_0-bf16.cu
- fattn-vec-instance-q5_0-f16.cu
- fattn-vec-instance-q5_0-q4_0.cu
- fattn-vec-instance-q5_0-q4_1.cu
- fattn-vec-instance-q5_0-q5_0.cu
- fattn-vec-instance-q5_0-q5_1.cu
- fattn-vec-instance-q5_0-q8_0.cu
- fattn-vec-instance-q5_1-bf16.cu
- fattn-vec-instance-q5_1-f16.cu
- fattn-vec-instance-q5_1-q4_0.cu
- fattn-vec-instance-q5_1-q4_1.cu
- fattn-vec-instance-q5_1-q5_0.cu
- fattn-vec-instance-q5_1-q5_1.cu
- fattn-vec-instance-q5_1-q8_0.cu
- fattn-vec-instance-q8_0-bf16.cu
- fattn-vec-instance-q8_0-f16.cu
- fattn-vec-instance-q8_0-q4_0.cu
- fattn-vec-instance-q8_0-q4_1.cu
- fattn-vec-instance-q8_0-q5_0.cu
- fattn-vec-instance-q8_0-q5_1.cu
- fattn-vec-instance-q8_0-q8_0.cu
- generate_cu_files.py
- mmf-instance-ncols_1.cu
- mmf-instance-ncols_10.cu
- mmf-instance-ncols_11.cu
- mmf-instance-ncols_12.cu
- mmf-instance-ncols_13.cu
- mmf-instance-ncols_14.cu
- mmf-instance-ncols_15.cu
- mmf-instance-ncols_16.cu
- mmf-instance-ncols_2.cu
- mmf-instance-ncols_3.cu
- mmf-instance-ncols_4.cu
- mmf-instance-ncols_5.cu
- mmf-instance-ncols_6.cu
- mmf-instance-ncols_7.cu
- mmf-instance-ncols_8.cu
- mmf-instance-ncols_9.cu
- mmq-instance-iq1_s.cu
- mmq-instance-iq2_s.cu
- mmq-instance-iq2_xs.cu
- mmq-instance-iq2_xxs.cu
- mmq-instance-iq3_s.cu
- mmq-instance-iq3_xxs.cu
- mmq-instance-iq4_nl.cu
- mmq-instance-iq4_xs.cu
- mmq-instance-mxfp4.cu
- mmq-instance-nvfp4.cu
- mmq-instance-q1_0.cu
- mmq-instance-q2_0.cu
- mmq-instance-q2_k.cu
- mmq-instance-q3_k.cu
- mmq-instance-q4_0.cu
- mmq-instance-q4_1.cu
- mmq-instance-q4_k.cu
- mmq-instance-q5_0.cu
- mmq-instance-q5_1.cu
- mmq-instance-q5_k.cu
- mmq-instance-q6_k.cu
- mmq-instance-q8_0.cu
- cuda.h
- hip.h
- musa.h
- acc.cu
- acc.cuh
- add-id.cu
- add-id.cuh
- allreduce.cu
- allreduce.cuh
- arange.cu
- arange.cuh
- argmax.cu
- argmax.cuh
- argsort.cu
- argsort.cuh
- binbcast.cu
- binbcast.cuh
- clamp.cu
- clamp.cuh
- CMakeLists.txt
- col2im-1d.cu
- col2im-1d.cuh
- common.cuh
- concat.cu
- concat.cuh
- conv-transpose-1d.cu
- conv-transpose-1d.cuh
- conv2d-dw.cu
- conv2d-dw.cuh
- conv2d-transpose.cu
- conv2d-transpose.cuh
- conv2d.cu
- conv2d.cuh
- convert.cu
- convert.cuh
- count-equal.cu
- count-equal.cuh
- cp-async.cuh
- cpy-utils.cuh
- cpy.cu
- cpy.cuh
- cross-entropy-loss.cu
- cross-entropy-loss.cuh
- cumsum.cu
- cumsum.cuh
- dequantize.cuh
- diag.cu
- diag.cuh
- diagmask.cu
- diagmask.cuh
- dsv4-hc.cu
- dsv4-hc.cuh
- fattn-common.cuh
- fattn-mma-f16.cuh
- fattn-tile.cu
- fattn-tile.cuh
- fattn-vec.cuh
- fattn.cu
- fattn.cuh
- fill.cu
- fill.cuh
- fwht.cu
- fwht.cuh
- gated_delta_net.cu
- gated_delta_net.cuh
- getrows.cu
- getrows.cuh
- ggml-cuda.cu
- gla.cu
- gla.cuh
- im2col.cu
- im2col.cuh
- lightning-indexer.cu
- lightning-indexer.cuh
- mean.cu
- mean.cuh
- mma.cuh
- mmf.cu
- mmf.cuh
- mmid.cu
- mmid.cuh
- mmq-config-ampere.cuh
- mmq-config-blackwell.cuh
- mmq-config-cdna.cuh
- mmq-config-pascal.cuh
- mmq-config-rdna2.cuh
- mmq-config-rdna3-5.cuh
- mmq-config-rdna3.cuh
- mmq-config-rdna4.cuh
- mmq-load-tiles.cuh
- mmq-vec-dot.cuh
- mmq.cu
- mmq.cuh
- mmvf.cu
- mmvf.cuh
- mmvq.cu
- mmvq.cuh
- norm.cu
- norm.cuh
- opt-step-adamw.cu
- opt-step-adamw.cuh
- opt-step-sgd.cu
- opt-step-sgd.cuh
- out-prod.cu
- out-prod.cuh
- pad.cu
- pad.cuh
- pad_reflect_1d.cu
- pad_reflect_1d.cuh
- pool2d.cu
- pool2d.cuh
- quantize.cu
- quantize.cuh
- reduce_rows.cuh
- roll.cu
- roll.cuh
- rope.cu
- rope.cuh
- scale.cu
- scale.cuh
- set-rows.cu
- set-rows.cuh
- set.cu
- set.cuh
- snake.cu
- snake.cuh
- softcap.cu
- softcap.cuh
- softmax.cu
- softmax.cuh
- solve_tri.cu
- solve_tri.cuh
- ssm-conv.cu
- ssm-conv.cuh
- ssm-scan.cu
- ssm-scan.cuh
- sum.cu
- sum.cuh
- sumrows.cu
- sumrows.cuh
- top-k.cu
- top-k.cuh
- topk-moe.cu
- topk-moe.cuh
- tri.cu
- tri.cuh
- tsembd.cu
- tsembd.cuh
- unary.cu
- unary.cuh
- upscale.cu
- upscale.cuh
- vecdotq.cuh
- wkv.cu
- wkv.cuh
- embed_one_kernel.cmake
- ggml-et-kernels-embed.cpp.in
- ggml-et-kernels-embed.hpp.in
- ggml-et-uberkernel-kernel-map.cpp.in
- ggml-et-uberkernel-kernel-map.h.in
- check_unimplemented_instructions.sh
- block_ops.h
- clamp_f32.c
- concat_f32.c
- cont_f16.c
- cont_f32.c
- conv_2d_f32_me.c
- cpy_f32_f16.c
- crt.S
- cumsum_f32.c
- diag_f32.c
- el_map_f32.c
- fill_f32.c
- flash_attn_ext_f16_me.c
- flash_attn_ext_f32.c
- gated_delta_net_f32.c
- get_rows_f32.c
- ggml_tensor.h
- glu_f32.c
- group_norm_f32.c
- im2col.c
- l2_norm_f32.c
- linker.ld
- math_fp.h
- mean_f32.c
- memops.c
- mul_mat_f16.c
- mul_mat_f16_matrix_engine.c
- mul_mat_f32.c
- mul_mat_f32_matrix_engine.c
- mul_mat_id_f32.c
- mul_mat_id_Q4_0.c
- mul_mat_id_Q8_0.c
- mul_mat_Q4_0.c
- mul_mat_Q4_0_matrix_engine.c
- mul_mat_Q8_0.c
- norm_f32.c
- pad_f32.c
- platform.h
- quants.h
- repeat_f32.c
- rms_norm_f32.c
- rms_norm_mul_f32.c
- rope_f32.c
- RunBackend.sh
- rwkv_wkv6_f32.c
- rwkv_wkv7_f32.c
- scale_f32.c
- set_f32.c
- set_rows_f32.c
- softmax_f32.c
- solve_tri_f32.c
- sqr_f32.c
- ssm_conv_f32.c
- ssm_scan_f32.c
- sum_rows_f32.c
- tensor.h
- tri_f32.c
- uberkernel.c
- unary_f32.c
- CMakeLists.txt
- CMakeLists.txt
- ggml-et-common.h
- ggml-et-cpu-compare.cpp
- ggml-et-cpu-compare.h
- ggml-et-kernels.cpp
- ggml-et-kernels.h
- ggml-et-memops.cpp
- ggml-et-memops.h
- ggml-et-ops.cpp
- ggml-et-ops.h
- ggml-et-uberkernel-common.h
- ggml-et.cpp
- act-ops.c
- argsort-ops.c
- binary-ops.c
- cmake-toolchain.cmake
- CMakeLists.txt
- concat-ops.c
- cpy-ops.c
- cumsum-ops.c
- diag-ops.c
- dma-queue.c
- dma-queue.h
- fill-ops.c
- flash-attn-ops.c
- flash-attn-ops.h
- gated-delta-net-ops.c
- get-rows-ops.c
- hex-bitmap.h
- hex-common.h
- hex-dma.h
- hex-dump.h
- hex-fastdiv.h
- hex-profile.h
- hex-utils.h
- hmx-fa-kernels.h
- hmx-mm-kernels-tiled.h
- hmx-queue.c
- hmx-queue.h
- hmx-utils.h
- htp-ctx.h
- htp-ops.h
- htp-tensor.c
- htp-tensor.h
- htp-vtcm.h
- htp_iface.idl
- hvx-arith.h
- hvx-base.h
- hvx-copy.h
- hvx-div.h
- hvx-dump.h
- hvx-exp.h
- hvx-fa-kernels.h
- hvx-flash-attn.h
- hvx-floor.h
- hvx-inverse.h
- hvx-log.h
- hvx-mm-kernels-flat.h
- hvx-mm-kernels-tiled.h
- hvx-norm.h
- hvx-pow.h
- hvx-reduce.h
- hvx-repl.h
- hvx-scale.h
- hvx-sigmoid.h
- hvx-sin-cos.h
- hvx-sqrt.h
- hvx-types.h
- hvx-utils.h
- im2col-ops.c
- main.c
- matmul-ops.c
- matmul-ops.h
- pad-ops.c
- repeat-ops.c
- rope-ops.c
- set-rows-ops.c
- softmax-ops.c
- solve-tri-ops.c
- ssm-conv.c
- sum-rows-ops.c
- unary-ops.c
- unary-ops.h
- work-queue.c
- work-queue.h
- CMakeLists.txt
- ggml-hexagon.cpp
- htp-drv.cpp
- htp-drv.h
- htp-opnode.h
- libdl.h
- libggml-htp.inf
- CMakeLists.txt
- CMakeLists.txt
- ggml-metal-common.cpp
- ggml-metal-common.h
- ggml-metal-context.h
- ggml-metal-context.m
- ggml-metal-device.cpp
- ggml-metal-device.h
- ggml-metal-device.m
- ggml-metal-impl.h
- ggml-metal-ops.cpp
- ggml-metal-ops.h
- ggml-metal.cpp
- ggml-metal.metal
- CMakeLists.txt
- mudnn.cu
- mudnn.cuh
- abs.cl
- add.cl
- add_id.cl
- argsort.cl
- clamp.cl
- concat.cl
- conv2d.cl
- conv2d_f16_f32.cl
- cpy.cl
- cumsum.cl
- cvt.cl
- diag.cl
- diag_mask_inf.cl
- div.cl
- embed_kernel.py
- exp.cl
- expm1.cl
- fill.cl
- flash_attn_f16.cl
- flash_attn_f32.cl
- flash_attn_f32_f16.cl
- flash_attn_f32_q4_0.cl
- flash_attn_f32_q8_0.cl
- flash_attn_pre_f16.cl
- gated_delta_net.cl
- gelu.cl
- gemm_moe_mxfp4_f32.cl
- gemm_moe_mxfp4_f32_ns.cl
- gemm_moe_mxfp4_q8_1_dp4a.cl
- gemm_moe_q4_0_f32_ns.cl
- gemm_moe_q4_0_q8_1_dp4a.cl
- gemm_moe_q4_1_f32_ns.cl
- gemm_moe_q4_k_f32_ns.cl
- gemm_moe_q4_k_q8_1_dp4a.cl
- gemm_moe_q5_0_f32_ns.cl
- gemm_moe_q5_1_f32_ns.cl
- gemm_moe_q5_k_f32_ns.cl
- gemm_moe_q6_k_f32_ns.cl
- gemm_moe_q6_k_q8_1_dp4a.cl
- gemm_moe_q8_0_f32_ns.cl
- gemm_moe_q8_1_dp4a.cl
- gemm_noshuffle_iq4_nl_f32.cl
- gemm_noshuffle_iq4_nl_q8_1_dp4a.cl
- gemm_noshuffle_q1_0_f32.cl
- gemm_noshuffle_q4_0_f32.cl
- gemm_noshuffle_q4_0_q8_1_dp4a.cl
- gemm_noshuffle_q4_1_f32.cl
- gemm_noshuffle_q4_k_f32.cl
- gemm_noshuffle_q4_k_q8_1_dp4a.cl
- gemm_noshuffle_q5_0_f32.cl
- gemm_noshuffle_q5_0_q8_1_dp4a.cl
- gemm_noshuffle_q5_1_f32.cl
- gemm_noshuffle_q5_k_f32.cl
- gemm_noshuffle_q5_k_q8_1_dp4a.cl
- gemm_noshuffle_q6_k_f32.cl
- gemm_noshuffle_q6_k_q8_1_dp4a.cl
- gemm_noshuffle_q8_0_f32.cl
- gemm_noshuffle_q8_0_q8_1_dp4a.cl
- gemm_xmem_f16_f32_os8.cl
- gemv_moe_mxfp4_f32.cl
- gemv_moe_mxfp4_f32_ns.cl
- gemv_moe_q4_0_f32_ns.cl
- gemv_moe_q4_1_f32_ns.cl
- gemv_moe_q4_k_f32_ns.cl
- gemv_moe_q5_0_f32_ns.cl
- gemv_moe_q5_1_f32_ns.cl
- gemv_moe_q5_k_f32_ns.cl
- gemv_moe_q6_k_f32_ns.cl
- gemv_noshuffle_iq4_nl_f32.cl
- gemv_noshuffle_q1_0_f32.cl
- gemv_noshuffle_q4_0_f32.cl
- gemv_noshuffle_q4_0_f32_spec.cl
- gemv_noshuffle_q4_1_f32.cl
- gemv_noshuffle_q4_k_f32.cl
- gemv_noshuffle_q5_0_f32.cl
- gemv_noshuffle_q5_1_f32.cl
- gemv_noshuffle_q5_k_f32.cl
- gemv_noshuffle_q6_k_f32.cl
- gemv_noshuffle_q8_0_f32.cl
- get_rows.cl
- glu.cl
- group_norm.cl
- im2col_f16.cl
- im2col_f32.cl
- l2_norm.cl
- mean.cl
- moe_combine.cl
- moe_reorder_b.cl
- moe_reorder_quant_a_q8_1.cl
- moe_sort_by_expert.cl
- mul.cl
- mul_mat_f16_f32.cl
- mul_mm_f16_f32_kq_kqv.cl
- mul_mm_f16_f32_l4_lm.cl
- mul_mm_f32_f32_l4_lm.cl
- mul_mm_iq4_nl_f32_l4_lm.cl
- mul_mm_q1_0_f32_l4_lm.cl
- mul_mm_q4_0_f32_l4_lm.cl
- mul_mm_q4_1_f32_l4_lm.cl
- mul_mm_q4_k_f32_l4_lm.cl
- mul_mm_q5_0_f32_l4_lm.cl
- mul_mm_q5_1_f32_l4_lm.cl
- mul_mm_q5_k_f32_l4_lm.cl
- mul_mm_q6_k_f32_l4_lm.cl
- mul_mm_q8_0_f32_l4_lm.cl
- mul_mv_f16_f16.cl
- mul_mv_f16_f32.cl
- mul_mv_f16_f32_1row.cl
- mul_mv_f16_f32_l4.cl
- mul_mv_f32_f32.cl
- mul_mv_id_mxfp4_f32.cl
- mul_mv_id_mxfp4_f32_flat.cl
- mul_mv_id_q4_0_f32_8x_flat.cl
- mul_mv_id_q8_0_f32.cl
- mul_mv_id_q8_0_f32_flat.cl
- mul_mv_iq4_nl_f32.cl
- mul_mv_iq4_nl_f32_flat.cl
- mul_mv_mxfp4_f32.cl
- mul_mv_mxfp4_f32_flat.cl
- mul_mv_q1_0_f32.cl
- mul_mv_q1_0_f32_flat.cl
- mul_mv_q4_0_f32.cl
- mul_mv_q4_0_f32_1d_16x_flat.cl
- mul_mv_q4_0_f32_1d_8x_flat.cl
- mul_mv_q4_0_f32_8x_flat.cl
- mul_mv_q4_0_f32_v.cl
- mul_mv_q4_1_f32.cl
- mul_mv_q4_1_f32_flat.cl
- mul_mv_q4_k_f32.cl
- mul_mv_q4_k_f32_flat.cl
- mul_mv_q5_0_f32.cl
- mul_mv_q5_0_f32_flat.cl
- mul_mv_q5_1_f32.cl
- mul_mv_q5_1_f32_flat.cl
- mul_mv_q5_k_f32.cl
- mul_mv_q5_k_f32_flat.cl
- mul_mv_q6_k_f32.cl
- mul_mv_q6_k_f32_flat.cl
- mul_mv_q8_0_f32.cl
- mul_mv_q8_0_f32_flat.cl
- neg.cl
- norm.cl
- pad.cl
- quant_a_q8_1.cl
- relu.cl
- repeat.cl
- rms_norm.cl
- rope.cl
- scale.cl
- set_rows.cl
- sigmoid.cl
- silu.cl
- softmax_4_f16.cl
- softmax_4_f32.cl
- softmax_f16.cl
- softmax_f32.cl
- softplus.cl
- solve_tri.cl
- sqr.cl
- sqrt.cl
- ssm_conv.cl
- ssm_scan.cl
- sub.cl
- sum_rows.cl
- tanh.cl
- transpose.cl
- tri.cl
- tsembd.cl
- upscale.cl
- cl-program-cache.cpp
- cl-program-cache.h
- CMakeLists.txt
- fa_tune.h
- ggml-opencl.cpp
- libdl.h
- add.cpp
- add_id.cpp
- argsort.cpp
- clamp.cpp
- concat.cpp
- cont.cpp
- cpy.cpp
- cumsum.cpp
- diag.cpp
- div.cpp
- fill.cpp
- flash_attn_ext.cpp
- gated_delta_net.cpp
- gated_delta_net.hpp
- gather_matmul.hpp
- get_rows.cpp
- glu_geglu.cpp
- glu_swiglu.cpp
- im2col.cpp
- l2_norm.cpp
- mul_mat_id.cpp
- mulmat.cpp
- norm.cpp
- pad.cpp
- permute.cpp
- repeat.cpp
- reshape.cpp
- rms_norm.cpp
- rope.cpp
- scale.cpp
- set.cpp
- set_rows.cpp
- softmax.cpp
- solve_tri.cpp
- sqr.cpp
- ssm_conv.cpp
- sum_rows.cpp
- transpose.cpp
- tri.cpp
- unary_silu.cpp
- unary_softplus.cpp
- view.cpp
- fuse_to_sdpa.cpp
- fuse_to_sdpa.h
- mark_decompression_convert_constant_folding.h
- mark_dequantization_subgraph.h
- squeeze_matmul.cpp
- squeeze_matmul.h
- weightless_caching_attributes.hpp
- decoder.h
- frontend.cpp
- frontend.h
- input_model.cpp
- input_model.h
- node_context.h
- op_table.cpp
- op_table.h
- translate_session.cpp
- translate_session.h
- utils.cpp
- utils.h
- .clang-format
- CMakeLists.txt
- ggml-decoder.cpp
- ggml-decoder.h
- ggml-openvino-extra.cpp
- ggml-openvino-extra.h
- ggml-openvino.cpp
- ggml-quants.cpp
- ggml-quants.h
- model-cache.cpp
- model-cache.h
- utils.cpp
- utils.h
- CMakeLists.txt
- ggml-rpc.cpp
- transport.cpp
- transport.h
- helper.hpp
- fattn-tile-instance-dkq112-dv112.cpp
- fattn-tile-instance-dkq128-dv128.cpp
- fattn-tile-instance-dkq256-dv256.cpp
- fattn-tile-instance-dkq40-dv40.cpp
- fattn-tile-instance-dkq512-dv512.cpp
- fattn-tile-instance-dkq576-dv512.cpp
- fattn-tile-instance-dkq64-dv64.cpp
- fattn-tile-instance-dkq72-dv72.cpp
- fattn-tile-instance-dkq80-dv80.cpp
- fattn-tile-instance-dkq96-dv96.cpp
- fattn-vec-instance-f16-f16.cpp
- fattn-vec-instance-f16-q4_0.cpp
- fattn-vec-instance-f16-q4_1.cpp
- fattn-vec-instance-f16-q5_0.cpp
- fattn-vec-instance-f16-q5_1.cpp
- fattn-vec-instance-f16-q8_0.cpp
- fattn-vec-instance-q4_0-f16.cpp
- fattn-vec-instance-q4_0-q4_0.cpp
- fattn-vec-instance-q4_0-q4_1.cpp
- fattn-vec-instance-q4_0-q5_0.cpp
- fattn-vec-instance-q4_0-q5_1.cpp
- fattn-vec-instance-q4_0-q8_0.cpp
- fattn-vec-instance-q4_1-f16.cpp
- fattn-vec-instance-q4_1-q4_0.cpp
- fattn-vec-instance-q4_1-q4_1.cpp
- fattn-vec-instance-q4_1-q5_0.cpp
- fattn-vec-instance-q4_1-q5_1.cpp
- fattn-vec-instance-q4_1-q8_0.cpp
- fattn-vec-instance-q5_0-f16.cpp
- fattn-vec-instance-q5_0-q4_0.cpp
- fattn-vec-instance-q5_0-q4_1.cpp
- fattn-vec-instance-q5_0-q5_0.cpp
- fattn-vec-instance-q5_0-q5_1.cpp
- fattn-vec-instance-q5_0-q8_0.cpp
- fattn-vec-instance-q5_1-f16.cpp
- fattn-vec-instance-q5_1-q4_0.cpp
- fattn-vec-instance-q5_1-q4_1.cpp
- fattn-vec-instance-q5_1-q5_0.cpp
- fattn-vec-instance-q5_1-q5_1.cpp
- fattn-vec-instance-q5_1-q8_0.cpp
- fattn-vec-instance-q8_0-f16.cpp
- fattn-vec-instance-q8_0-q4_0.cpp
- fattn-vec-instance-q8_0-q4_1.cpp
- fattn-vec-instance-q8_0-q5_0.cpp
- fattn-vec-instance-q8_0-q5_1.cpp
- fattn-vec-instance-q8_0-q8_0.cpp
- add-id.cpp
- add-id.hpp
- backend.hpp
- binbcast.cpp
- binbcast.hpp
- CMakeLists.txt
- col2im-1d.cpp
- col2im-1d.hpp
- common.cpp
- common.hpp
- concat.cpp
- concat.hpp
- conv.cpp
- conv.hpp
- conv2d-dw.cpp
- conv2d-dw.hpp
- conv2d-transpose.cpp
- conv2d-transpose.hpp
- conv2d.cpp
- conv2d.hpp
- conv3d.cpp
- conv3d.hpp
- convert.cpp
- convert.hpp
- count-equal.cpp
- count-equal.hpp
- cpy.cpp
- cpy.hpp
- cross_entropy_loss.cpp
- cross_entropy_loss.hpp
- cumsum.cpp
- cumsum.hpp
- dequantize.hpp
- diag.cpp
- diag.hpp
- dmmv.cpp
- dmmv.hpp
- dsv4-hc.cpp
- dsv4-hc.hpp
- element_wise.cpp
- element_wise.hpp
- esimd.hpp
- fattn-buffers.cpp
- fattn-buffers.hpp
- fattn-common.hpp
- fattn-mkl.cpp
- fattn-onednn.cpp
- fattn-onednn.hpp
- fattn-tile.cpp
- fattn-tile.hpp
- fattn-vec.hpp
- fattn.cpp
- fattn.hpp
- fill.cpp
- fill.hpp
- fusion.cpp
- fusion.hpp
- fwht.cpp
- fwht.hpp
- gated_delta_net.cpp
- gated_delta_net.hpp
- gemm.hpp
- getrows.cpp
- getrows.hpp
- ggml-sycl.cpp
- gla.cpp
- gla.hpp
- im2col.cpp
- im2col.hpp
- lightning-indexer.cpp
- lightning-indexer.hpp
- mmq.cpp
- mmq.hpp
- mmvq.cpp
- mmvq.hpp
- norm.cpp
- norm.hpp
- opt-step.cpp
- opt-step.hpp
- outprod.cpp
- outprod.hpp
- pad.cpp
- pad.hpp
- pad_reflect_1d.cpp
- pad_reflect_1d.hpp
- pool.cpp
- pool.hpp
- presets.hpp
- quantize.hpp
- quants.hpp
- repeat_back.cpp
- repeat_back.hpp
- roll.cpp
- roll.hpp
- rope.cpp
- rope.hpp
- set.cpp
- set.hpp
- set_rows.cpp
- set_rows.hpp
- softmax.cpp
- softmax.hpp
- solve_tri.cpp
- solve_tri.hpp
- ssm_conv.cpp
- ssm_conv.hpp
- ssm_scan.cpp
- ssm_scan.hpp
- sycl_hw.cpp
- sycl_hw.hpp
- topk-moe.cpp
- topk-moe.hpp
- tsembd.cpp
- tsembd.hpp
- type.hpp
- upscale.cpp
- upscale.hpp
- vecdotq.hpp
- wkv.cpp
- wkv.hpp
- api_remoting.h
- apir_backend.gen.h
- apir_backend.h
- apir_cs.h
- apir_cs_ggml.h
- apir_cs_rpc.h
- apir_cs_ggml-rpc-back.cpp
- backend-convert.h
- backend-dispatched-backend.cpp
- backend-dispatched-buffer-type.cpp
- backend-dispatched-buffer.cpp
- backend-dispatched-device.cpp
- backend-dispatched.cpp
- backend-dispatched.gen.h
- backend-dispatched.h
- backend-virgl-apir.h
- backend.cpp
- CMakeLists.txt
- apir_hw.h
- apir_cs_ggml-rpc-front.cpp
- CMakeLists.txt
- ggml-backend-buffer-type.cpp
- ggml-backend-buffer.cpp
- ggml-backend-device.cpp
- ggml-backend-reg.cpp
- ggml-backend.cpp
- ggml-remoting.h
- ggmlremoting_functions.yaml
- regenerate_remoting.py
- virtgpu-apir.h
- virtgpu-forward-backend.cpp
- virtgpu-forward-buffer-type.cpp
- virtgpu-forward-buffer.cpp
- virtgpu-forward-device.cpp
- virtgpu-forward-impl.h
- virtgpu-forward.gen.h
- virtgpu-shm.cpp
- virtgpu-shm.h
- virtgpu-utils.cpp
- virtgpu-utils.h
- virtgpu.cpp
- virtgpu.h
- host-toolchain.cmake.in
- bfloat16.comp
- coopmat.comp
- coopmat2.comp
- coopmat2_decode_vector.comp
- float_e2m1.comp
- float_e4m3.comp
- integer_dot.comp
- acc.comp
- add.comp
- add1.comp
- add_id.comp
- arange.comp
- argmax.comp
- argsort.comp
- argsort_large.comp
- CMakeLists.txt
- col2im_1d.comp
- concat.comp
- contig_copy.comp
- conv2d_dw.comp
- conv2d_mm.comp
- conv3d_mm.comp
- conv_transpose_1d.comp
- copy.comp
- copy_from_quant.comp
- copy_to_quant.comp
- copy_transpose.comp
- copy_transpose_02.comp
- count_equal.comp
- count_experts.comp
- cumsum.comp
- cumsum_multipass1.comp
- cumsum_multipass2.comp
- dequant_f32.comp
- dequant_funcs.glsl
- dequant_funcs_cm2.glsl
- dequant_head.glsl
- dequant_iq1_m.comp
- dequant_iq1_s.comp
- dequant_iq2_s.comp
- dequant_iq2_xs.comp
- dequant_iq2_xxs.comp
- dequant_iq3_s.comp
- dequant_iq3_xxs.comp
- dequant_iq4_nl.comp
- dequant_iq4_xs.comp
- dequant_mxfp4.comp
- dequant_nvfp4.comp
- dequant_q1_0.comp
- dequant_q2_0.comp
- dequant_q2_k.comp
- dequant_q3_k.comp
- dequant_q4_0.comp
- dequant_q4_1.comp
- dequant_q4_k.comp
- dequant_q5_0.comp
- dequant_q5_1.comp
- dequant_q5_k.comp
- dequant_q6_k.comp
- dequant_q8_0.comp
- dequant_tq2_0.comp
- diag.comp
- diag_mask_inf.comp
- div.comp
- dot_product_funcs.glsl
- fill.comp
- flash_attn.comp
- flash_attn_base.glsl
- flash_attn_cm1.comp
- flash_attn_cm2.comp
- flash_attn_dequant.glsl
- flash_attn_mask_opt.comp
- flash_attn_mmq_funcs.glsl
- flash_attn_split_k_reduce.comp
- fwht.comp
- gated_delta_net.comp
- geglu.comp
- geglu_erf.comp
- geglu_quick.comp
- generic_binary_head.glsl
- generic_head.glsl
- generic_unary_head.glsl
- get_rows.comp
- get_rows_back.comp
- get_rows_quant.comp
- gla.comp
- glu_head.glsl
- glu_main.glsl
- group_norm.comp
- im2col.comp
- im2col_3d.comp
- l2_norm.comp
- log.comp
- mul.comp
- mul_mat_split_k_reduce.comp
- mul_mat_vec.comp
- mul_mat_vec_base.glsl
- mul_mat_vec_iface.glsl
- mul_mat_vec_iq1_m.comp
- mul_mat_vec_iq1_s.comp
- mul_mat_vec_iq2_s.comp
- mul_mat_vec_iq2_xs.comp
- mul_mat_vec_iq2_xxs.comp
- mul_mat_vec_iq3_s.comp
- mul_mat_vec_iq3_xxs.comp
- mul_mat_vec_nc.comp
- mul_mat_vec_p021.comp
- mul_mat_vec_q2_k.comp
- mul_mat_vec_q3_k.comp
- mul_mat_vec_q4_k.comp
- mul_mat_vec_q5_k.comp
- mul_mat_vec_q6_k.comp
- mul_mat_vec_tq2_0.comp
- mul_mat_vecq.comp
- mul_mat_vecq_funcs.glsl
- mul_mm.comp
- mul_mm_cm2.comp
- mul_mm_funcs.glsl
- mul_mm_id_funcs.glsl
- mul_mmq.comp
- mul_mmq_funcs.glsl
- mul_mmq_shmem_types.glsl
- multi_add.comp
- norm.comp
- opt_step_adamw.comp
- opt_step_sgd.comp
- out_prod.comp
- pad.comp
- pool1d.comp
- pool2d.comp
- quantize_q8_1.comp
- reglu.comp
- repeat.comp
- repeat_back.comp
- rms_norm.comp
- rms_norm_back.comp
- rms_norm_partials.comp
- roll.comp
- rope_funcs.glsl
- rope_head.glsl
- rope_multi.comp
- rope_neox.comp
- rope_norm.comp
- rope_params.glsl
- rope_vision.comp
- scale.comp
- silu_back.comp
- snake.comp
- soft_max.comp
- soft_max_back.comp
- soft_max_large1.comp
- soft_max_large2.comp
- soft_max_large3.comp
- soft_max_large_common.glsl
- solve_tri.comp
- ssm_conv.comp
- ssm_scan.comp
- sub.comp
- sum_rows.comp
- sum_rows.glsl
- swiglu.comp
- swiglu_oai.comp
- timestep_embedding.comp
- topk_argsort.comp
- topk_moe.comp
- topk_nary_search.comp
- tri.comp
- types.glsl
- unary.comp
- upscale.comp
- utils.glsl
- vulkan-shaders-gen.cpp
- wkv6.comp
- wkv7.comp
- CMakeLists.txt
- ggml-vulkan.cpp
- add_id.wgsl
- argmax.wgsl
- argsort.wgsl
- argsort_merge.wgsl
- binary.wgsl
- common_decls.tmpl
- concat.wgsl
- conv2d.wgsl
- conv2d_dw.wgsl
- cpy.wgsl
- cumsum.wgsl
- embed_wgsl.py
- flash_attn.wgsl
- flash_attn_decls.tmpl
- flash_attn_staging.tmpl
- flash_attn_tile.wgsl
- flash_attn_vec_blk.wgsl
- flash_attn_vec_reduce.wgsl
- flash_attn_vec_split.wgsl
- gated_delta_net.wgsl
- get_rows.wgsl
- glu.wgsl
- im2col.wgsl
- memset.wgsl
- mul_mat_decls.tmpl
- mul_mat_id.wgsl
- mul_mat_id_gather.wgsl
- mul_mat_id_vec.wgsl
- mul_mat_reg_tile.wgsl
- mul_mat_subgroup_matrix.wgsl
- mul_mat_vec.wgsl
- mul_mat_vec_acc.tmpl
- mul_mat_vec_q_acc.tmpl
- pad.wgsl
- quant_inner_loops.tmpl
- quantize_q8.wgsl
- repeat.wgsl
- rms_norm_mul.wgsl
- rope.wgsl
- row_norm.wgsl
- scale.wgsl
- set.wgsl
- set_rows.wgsl
- set_rows_quant.wgsl
- soft_max.wgsl
- solve_tri.wgsl
- ssm_conv.wgsl
- ssm_scan.wgsl
- sum_rows.wgsl
- unary.wgsl
- upscale.wgsl
- CMakeLists.txt
- ggml-webgpu-shader-lib.hpp
- ggml-webgpu.cpp
- pre_wgsl.hpp
- .gitignore
- CMakeLists.txt
- common.hpp
- ggml-zdnn.cpp
- mmf.cpp
- mmf.hpp
- utils.cpp
- utils.hpp
- CMakeLists.txt
- ggml-zendnn.cpp
- CMakeLists.txt
- ggml-alloc.c
- ggml-backend-dl.cpp
- ggml-backend-dl.h
- ggml-backend-impl.h
- ggml-backend-meta.cpp
- ggml-backend-reg.cpp
- ggml-backend.cpp
- ggml-common.h
- ggml-feats.h
- ggml-impl.h
- ggml-opt.cpp
- ggml-quants.c
- ggml-quants.h
- ggml-threading.cpp
- ggml-threading.h
- ggml.c
- ggml.cpp
- gguf.cpp
- .gitignore
- CMakeLists.txt
- reader.py
- writer.py
- gguf_convert_endian.py
- gguf_dump.py
- gguf_editor_gui.py
- gguf_hash.py
- gguf_new_metadata.py
- gguf_set_metadata.py
- __init__.py
- constants.py
- gguf.py
- gguf_reader.py
- gguf_writer.py
- lazy.py
- metadata.py
- py.typed
- quants.py
- tensor_mapping.py
- utility.py
- vocab.py
- __init__.py
- test_gguf_reader_validation.py
- test_metadata.py
- test_quants.py
- LICENSE
- pyproject.toml
- README.md
- arithmetic.gbnf
- c.gbnf
- chess.gbnf
- english.gbnf
- japanese.gbnf
- json.gbnf
- json_arr.gbnf
- list.gbnf
- README.md
- llama-cpp.h
- llama.h
- LICENSE-jsonhpp
- llama0-banner.png
- llama0-logo.png
- llama1-banner.png
- llama1-icon-transparent.png
- llama1-icon-transparent.svg
- llama1-icon.png
- llama1-icon.svg
- llama1-logo.png
- llama1-logo.svg
- matmul.png
- matmul.svg
- Apertus-8B-Instruct.jinja
- Apriel-1.6-15b-Thinker-fixed.jinja
- Bielik-11B-v3.0-Instruct.jinja
- ByteDance-Seed-OSS.jinja
- Cohere2MoE.jinja
- CohereForAI-c4ai-command-r-plus-tool_use.jinja
- CohereForAI-c4ai-command-r7b-12-2024-tool_use.jinja
- deepseek-ai-DeepSeek-R1-Distill-Llama-8B.jinja
- deepseek-ai-DeepSeek-R1-Distill-Qwen-32B.jinja
- deepseek-ai-DeepSeek-V3.1.jinja
- deepseek-ai-DeepSeek-V3.2.jinja
- deepseek-ai-DeepSeek-V4-Flash-0731.jinja
- deepseek-ai-DeepSeek-V4.jinja
- fireworks-ai-llama-3-firefunction-v2.jinja
- GigaChat3-10B-A1.8B.jinja
- GigaChat3.1-10B-A1.8B.jinja
- GLM-4.6.jinja
- GLM-4.7-Flash.jinja
- google-gemma-2-2b-it.jinja
- google-gemma-4-31B-it-interleaved.jinja
- google-gemma-4-31B-it.jinja
- HuggingFaceTB-SmolLM3-3B.jinja
- ibm-granite-granite-3.3-2B-Instruct.jinja
- ibm-granite-granite-4.0.jinja
- ibm-granite-granite-4.1.jinja
- Kimi-K2-Instruct.jinja
- Kimi-K2-Thinking.jinja
- Kimi-K3.jinja
- LFM2-8B-A1B.jinja
- LFM2.5-8B-A1B.jinja
- LFM2.5-Instruct.jinja
- llama-cpp-deepseek-r1.jinja
- llama-cpp-rwkv-world.jinja
- meetkai-functionary-medium-v3.1.jinja
- meetkai-functionary-medium-v3.2.jinja
- meta-llama-Llama-3.1-8B-Instruct.jinja
- meta-llama-Llama-3.2-3B-Instruct.jinja
- meta-llama-Llama-3.3-70B-Instruct.jinja
- microsoft-Phi-3.5-mini-instruct.jinja
- MiMo-VL.jinja
- MiniMax-M1.jinja
- MiniMax-M2.jinja
- MiniMax-M3.jinja
- Mistral-Small-3.2-24B-Instruct-2506.jinja
- mistralai-Ministral-3-14B-Reasoning-2512.jinja
- mistralai-Mistral-Nemo-Instruct-2407.jinja
- moonshotai-Kimi-K2.jinja
- muse-glimmer.jinja
- NousResearch-Hermes-2-Pro-Llama-3-8B-tool_use.jinja
- NousResearch-Hermes-3-Llama-3.1-8B-tool_use.jinja
- NVIDIA-Nemotron-3-Nano-30B-A3B-BF16.jinja
- NVIDIA-Nemotron-Nano-v2.jinja
- openai-gpt-oss-120b.jinja
- openbmb-MiniCPM5-1B.jinja
- poolside-Laguna-S-2.1.jinja
- poolside-Laguna-XS-2.1.jinja
- poolside-Laguna-XS.2.jinja
- Qwen-Qwen2.5-7B-Instruct.jinja
- Qwen-Qwen3-0.6B.jinja
- Qwen-QwQ-32B.jinja
- Qwen3-Coder.jinja
- Qwen3.5-4B.jinja
- README.md
- Reka-Edge.jinja
- StepFun3.5-Flash.jinja
- tencent-Hy3.jinja
- unsloth-Apriel-1.5.jinja
- unsloth-mistral-Devstral-Small-2507.jinja
- upstage-Solar-Open-100B.jinja
- .editorconfig
- ggml-vocab-aquila.gguf
- ggml-vocab-baichuan.gguf
- ggml-vocab-bert-bge.gguf
- ggml-vocab-bert-bge.gguf.inp
- ggml-vocab-bert-bge.gguf.out
- ggml-vocab-command-r.gguf
- ggml-vocab-command-r.gguf.inp
- ggml-vocab-command-r.gguf.out
- ggml-vocab-deepseek-coder.gguf
- ggml-vocab-deepseek-coder.gguf.inp
- ggml-vocab-deepseek-coder.gguf.out
- ggml-vocab-deepseek-llm.gguf
- ggml-vocab-deepseek-llm.gguf.inp
- ggml-vocab-deepseek-llm.gguf.out
- ggml-vocab-falcon.gguf
- ggml-vocab-falcon.gguf.inp
- ggml-vocab-falcon.gguf.out
- ggml-vocab-gemma-4.gguf
- ggml-vocab-gemma-4.gguf.inp
- ggml-vocab-gemma-4.gguf.out
- ggml-vocab-gpt-2.gguf
- ggml-vocab-gpt-2.gguf.inp
- ggml-vocab-gpt-2.gguf.out
- ggml-vocab-gpt-neox.gguf
- ggml-vocab-llama-bpe.gguf
- ggml-vocab-llama-bpe.gguf.inp
- ggml-vocab-llama-bpe.gguf.out
- ggml-vocab-llama-spm.gguf
- ggml-vocab-llama-spm.gguf.inp
- ggml-vocab-llama-spm.gguf.out
- ggml-vocab-mpt.gguf
- ggml-vocab-mpt.gguf.inp
- ggml-vocab-mpt.gguf.out
- ggml-vocab-nomic-bert-moe.gguf
- ggml-vocab-phi-3.gguf
- ggml-vocab-phi-3.gguf.inp
- ggml-vocab-phi-3.gguf.out
- ggml-vocab-qwen2.gguf
- ggml-vocab-qwen2.gguf.inp
- ggml-vocab-qwen2.gguf.out
- ggml-vocab-qwen35.gguf
- ggml-vocab-qwen35.gguf.inp
- ggml-vocab-qwen35.gguf.out
- ggml-vocab-refact.gguf
- ggml-vocab-refact.gguf.inp
- ggml-vocab-refact.gguf.out
- ggml-vocab-starcoder.gguf
- ggml-vocab-starcoder.gguf.inp
- ggml-vocab-starcoder.gguf.out
- CMakeLists.txt
- q8dot.cpp
- vdot.cpp
- CMakeLists.txt
- requirements-all.txt
- requirements-compare-llama-bench.txt
- requirements-convert_hf_to_gguf.txt
- requirements-convert_hf_to_gguf_update.txt
- requirements-convert_legacy_llama.txt
- requirements-convert_llama_ggml_to_gguf.txt
- requirements-convert_lora_to_gguf.txt
- requirements-gguf_editor_gui.txt
- requirements-pydantic.txt
- requirements-server-bench.txt
- requirements-test-tokenizer-random.txt
- requirements-tool_bench.txt
- validate-apps.sh
- validate-ios.sh
- validate-macos.sh
- validate-tvos.sh
- validate-visionos.sh
- gcn-cdna-vgpr-check.py
- jinja-tester.py
- requirements.txt
- llama-cli.farf
- run-bench.sh
- run-cli.sh
- run-completion.sh
- run-mtmd.sh
- run-tool.sh
- run_linux.sh
- conftest.py
- run_backend_ops_posix.py
- run_bench_tests_posix.py
- utils.py
- requirements.txt
- run_qdc_jobs.py
- run-bench.ps1
- run-cli.ps1
- run-completion.ps1
- run-mtmd.ps1
- run-tool.ps1
- setup-build.ps1
- ggml-hexagon-profile.py
- ggml-hexagon-trace.py
- bench-models.sh
- build-info.sh
- check-requirements.sh
- compare-commits.sh
- compare-llama-bench.py
- compare-logprobs.py
- create_ops_docs.py
- debug-test.sh
- gen-authors.sh
- gen-unicode-data.py
- get-flags.mk
- get-hellaswag.sh
- get-pg.sh
- get-wikitext-2.sh
- get-winogrande.sh
- get_chat_template.py
- git-bisect-run.sh
- git-bisect.sh
- hf.sh
- install-oneapi.bat
- make-release-checks.sh
- make-release-desc.sh
- pr2wt.sh
- release.sh
- serve-static.js
- server-bench.py
- server-test-function-call.py
- server-test-model.py
- server-test-parallel-tc.py
- server-test-structured.py
- sync-ggml-am.sh
- sync-ggml.last
- sync-ggml.sh
- sync_vendor.py
- tool_bench.py
- tool_bench.sh
- ui-assets.cmake
- verify-checksum-models.py
- wc2wt.sh
- SKILL.md
- SKILL.md
- afmoe.cpp
- apertus.cpp
- arcee.cpp
- arctic.cpp
- arwkv7.cpp
- baichuan.cpp
- bailingmoe.cpp
- bailingmoe2.cpp
- bailingmoe3.cpp
- bert.cpp
- bitnet.cpp
- bloom.cpp
- chameleon.cpp
- chatglm.cpp
- clip.cpp
- codeshell.cpp
- cogvlm.cpp
- cohere2.cpp
- cohere2moe.cpp
- command-r.cpp
- dbrx.cpp
- deci.cpp
- deepseek.cpp
- deepseek2.cpp
- deepseek2ocr.cpp
- deepseek32.cpp
- deepseek4.cpp
- delta-net-base.cpp
- dflash.cpp
- dots1.cpp
- dream.cpp
- eagle3.cpp
- ernie4-5-moe.cpp
- ernie4-5.cpp
- eurobert.cpp
- exaone-moe.cpp
- exaone.cpp
- exaone4.cpp
- falcon-h1.cpp
- falcon.cpp
- gemma-embedding.cpp
- gemma.cpp
- gemma2.cpp
- gemma3.cpp
- gemma3n.cpp
- gemma4-assistant.cpp
- gemma4.cpp
- glm-dsa.cpp
- glm4-moe.cpp
- glm4.cpp
- gpt2.cpp
- gptneox.cpp
- granite-hybrid.cpp
- granite-moe.cpp
- granite-swa.cpp
- granite-switch.cpp
- granite.cpp
- grok.cpp
- grovemoe.cpp
- hunyuan-dense.cpp
- hunyuan-moe.cpp
- hunyuan-vl.cpp
- hy-v3.cpp
- internlm2.cpp
- jais.cpp
- jais2.cpp
- jamba.cpp
- jina-bert-v2.cpp
- jina-bert-v3.cpp
- kimi-k3.cpp
- kimi-linear.cpp
- laguna.cpp
- lfm2.cpp
- lfm2moe.cpp
- llada-moe.cpp
- llada.cpp
- llama-embed.cpp
- llama.cpp
- llama4.cpp
- maincoder.cpp
- mamba-base.cpp
- mamba.cpp
- mamba2.cpp
- mellum.cpp
- mimo2.cpp
- minicpm.cpp
- minicpm3.cpp
- minimax-01.cpp
- minimax-m2.cpp
- minimax-m3.cpp
- mistral3.cpp
- mistral4.cpp
- models.h
- modern-bert.cpp
- mpt.cpp
- muse-glimmer.cpp
- nanbeige.cpp
- nemotron-h-moe.cpp
- nemotron-h.cpp
- nemotron.cpp
- neo-bert.cpp
- nomic-bert-moe.cpp
- nomic-bert.cpp
- olmo.cpp
- olmo2.cpp
- olmoe.cpp
- openai-moe.cpp
- openelm.cpp
- orion.cpp
- paddleocr.cpp
- pangu-embed.cpp
- phi2.cpp
- phi3.cpp
- phimoe.cpp
- plamo.cpp
- plamo2.cpp
- plamo3.cpp
- plm.cpp
- pockettts.cpp
- qwen.cpp
- qwen2.cpp
- qwen2moe.cpp
- qwen2vl.cpp
- qwen3.cpp
- qwen35.cpp
- qwen35moe.cpp
- qwen3moe.cpp
- qwen3next.cpp
- qwen3tts.cpp
- qwen3vl.cpp
- qwen3vlmoe.cpp
- refact.cpp
- rnd1.cpp
- rwkv6-base.cpp
- rwkv6.cpp
- rwkv6qwen2.cpp
- rwkv7-base.cpp
- rwkv7.cpp
- seed-oss.cpp
- smallthinker.cpp
- smollm3.cpp
- stablelm.cpp
- starcoder.cpp
- starcoder2.cpp
- step35.cpp
- t5.cpp
- t5encoder.cpp
- talkie.cpp
- wavtokenizer-dec.cpp
- xverse.cpp
- CMakeLists.txt
- llama-adapter.cpp
- llama-adapter.h
- llama-arch.cpp
- llama-arch.h
- llama-batch.cpp
- llama-batch.h
- llama-chat.cpp
- llama-chat.h
- llama-context.cpp
- llama-context.h
- llama-cparams.cpp
- llama-cparams.h
- llama-ext.h
- llama-grammar.cpp
- llama-grammar.h
- llama-graph.cpp
- llama-graph.h
- llama-hparams.cpp
- llama-hparams.h
- llama-impl.cpp
- llama-impl.h
- llama-io.cpp
- llama-io.h
- llama-kv-cache-dsa.cpp
- llama-kv-cache-dsa.h
- llama-kv-cache-dsv4.cpp
- llama-kv-cache-dsv4.h
- llama-kv-cache-iswa.cpp
- llama-kv-cache-iswa.h
- llama-kv-cache-msa.cpp
- llama-kv-cache-msa.h
- llama-kv-cache.cpp
- llama-kv-cache.h
- llama-kv-cells.h
- llama-memory-hybrid-iswa.cpp
- llama-memory-hybrid-iswa.h
- llama-memory-hybrid.cpp
- llama-memory-hybrid.h
- llama-memory-recurrent.cpp
- llama-memory-recurrent.h
- llama-memory.cpp
- llama-memory.h
- llama-mmap.cpp
- llama-mmap.h
- llama-model-loader.cpp
- llama-model-loader.h
- llama-model-saver.cpp
- llama-model-saver.h
- llama-model.cpp
- llama-model.h
- llama-quant.cpp
- llama-quant.h
- llama-sampler.cpp
- llama-sampler.h
- llama-vocab.cpp
- llama-vocab.h
- llama.cpp
- unicode-data.cpp
- unicode-data.h
- unicode.cpp
- unicode.h
- simple-tokenize.cpp
- simple-tokenize.h
- test-basic.cpp
- test-gbnf-generation.cpp
- test-json-parser.cpp
- test-json-serialization.cpp
- test-python-dict-parser.cpp
- test-unicode.cpp
- tests.h
- deepseek-v3.1.schema
- gemma-3-4b-it.schema
- glm-4.6v.schema
- gpt-oss-120b.schema
- meta-llama-3.1-70b-instruct.schema
- nemotron-nano-3-30b-a3b.schema
- qwen3-0.6b.schema
- qwen3-14b.schema
- qwen3-coder-next.schema
- qwen3.5-397b-a17b.schema
- qwen3.6-27b.schema
- step-3.5-flash.schema
- .gitignore
- CMakeLists.txt
- gguf-model-data.cpp
- gguf-model-data.h
- test-alloc.cpp
- test-arg-parser.cpp
- test-autorelease.cpp
- test-backend-ops.cpp
- test-backend-sampler.cpp
- test-barrier.cpp
- test-batch-alloc.cpp
- test-c.c
- test-chat-auto-parser.cpp
- test-chat-peg-parser.cpp
- test-chat-template.cpp
- test-chat.cpp
- test-col2im-1d.cpp
- test-double-float.cpp
- test-export-graph-ops.cpp
- test-gbnf-validator.cpp
- test-gguf-model-data.cpp
- test-gguf.cpp
- test-grammar-integration.cpp
- test-grammar-llguidance.cpp
- test-grammar-parser.cpp
- test-jinja.cpp
- test-json-schema-to-grammar.cpp
- test-llama-archs.cpp
- test-llama-grammar.cpp
- test-log.cpp
- test-lora-conversion-inference.sh
- test-model-load-cancel.cpp
- test-model-resolution.cpp
- test-mtmd-c-api.c
- test-mtmd-impl.cpp
- test-opt.cpp
- test-peg-parser.cpp
- test-quant-type-selection.cpp
- test-quantize-fns.cpp
- test-quantize-perf.cpp
- test-quantize-stats.cpp
- test-reasoning-budget.cpp
- test-recurrent-state-rollback.cpp
- test-rope.cpp
- test-rset-release.cpp
- test-sampling.cpp
- test-save-load-state.cpp
- test-state-restore-fragmented.cpp
- test-thread-safety.cpp
- test-tokenizer-0.cpp
- test-tokenizer-0.py
- test-tokenizer-0.sh
- test-tokenizer-1-bpe.cpp
- test-tokenizer-1-spm.cpp
- test-tokenizer-random.py
- test-tokenizers-repo.sh
- test-unicode.cpp
- testing.h
- batched-bench.cpp
- CMakeLists.txt
- main.cpp
- README.md
- cli-client.cpp
- cli-client.h
- cli-context.cpp
- cli-context.h
- cli-server.h
- cli-ui.h
- cli.cpp
- CMakeLists.txt
- main.cpp
- README.md
- CMakeLists.txt
- completion.cpp
- main.cpp
- README.md
- CMakeLists.txt
- completions.txt
- cvector-generator.cpp
- mean.hpp
- negative.txt
- pca.hpp
- positive.txt
- README.md
- CMakeLists.txt
- export-lora.cpp
- README.md
- CMakeLists.txt
- fit-params.cpp
- main.cpp
- README.md
- CMakeLists.txt
- gguf-split.cpp
- README.md
- tests.sh
- CMakeLists.txt
- imatrix.cpp
- README.md
- CMakeLists.txt
- llama-bench.cpp
- main.cpp
- README.md
- mtmd-debug.cpp
- mtmd-debug.h
- mtmd-debug.md
- convert_image_encoder_to_gguf.py
- glmedge-convert-image-encoder-to-gguf.py
- glmedge-surgery.py
- llava_surgery.py
- llava_surgery_v2.py
- minicpmv-convert-image-encoder-to-gguf.py
- minicpmv-surgery.py
- cogvlm.cpp
- conformer.cpp
- deepseekocr.cpp
- deepseekocr2.cpp
- dotsocr.cpp
- exaone4_5.cpp
- gemma4a.cpp
- gemma4ua.cpp
- gemma4uv.cpp
- gemma4v.cpp
- glm4v.cpp
- granite-speech.cpp
- granite4-vision.cpp
- hunyuanvl.cpp
- internvl.cpp
- kimik25.cpp
- kimivl.cpp
- llama4.cpp
- llava.cpp
- mimo-audio.cpp
- mimovl.cpp
- minicpmv.cpp
- minimax-m3.cpp
- mobilenetv5.cpp
- models.h
- muse-glimmer.cpp
- nemotron-v2-vl.cpp
- paddleocr.cpp
- parakeet.cpp
- pixtral.cpp
- pockettts-gen.cpp
- pockettts-seanet.cpp
- pockettts-spkenc.cpp
- qwen2vl.cpp
- qwen3a.cpp
- qwen3tts-gen.cpp
- qwen3tts-spkenc.cpp
- qwen3vl.cpp
- siglip.cpp
- step3vl.cpp
- whisper-enc.cpp
- yasa2.cpp
- youtuvl.cpp
- test-1-ground-truth.txt
- test-1-positive.png
- test-deepseek-ocr.py
- tests-requirements.txt
- clip-graph.h
- clip-impl.h
- clip-model.h
- clip.cpp
- clip.h
- CMakeLists.txt
- deprecation-warning.cpp
- mtmd-audio.cpp
- mtmd-audio.h
- mtmd-cli.cpp
- mtmd-helper-common.h
- mtmd-helper-gen.cpp
- mtmd-helper.cpp
- mtmd-helper.h
- mtmd-image.cpp
- mtmd-image.h
- mtmd-internal.h
- mtmd.cpp
- mtmd.h
- README-dev.md
- README.md
- requirements.txt
- test-1.jpeg
- test-2.mp3
- test-3.mp4
- tests.sh
- CMakeLists.txt
- debug-template-parser.cpp
- template-analysis.cpp
- CMakeLists.txt
- main.cpp
- perplexity.cpp
- README.md
- CMakeLists.txt
- main.cpp
- quantize.cpp
- README.md
- tests.sh
- CMakeLists.txt
- README.md
- results.cpp
- CMakeLists.txt
- README.md
- rpc-server.cpp
- README.md
- requirements.txt
- speed_bench.py
- speed_bench_compare.py
- bench.py
- prometheus.yml
- README.md
- requirements.txt
- script.js
- mcp_burst_server.py
- mcp_crash_server.py
- mcp_echo_server.py
- mcp_grandchild_server.py
- mcp_malformed_server.py
- mcp_slow_server.py
- test_basic.py
- test_chat_completion.py
- test_compat_anthropic.py
- test_compat_gcp.py
- test_compat_oai_responses.py
- test_completion.py
- test_ctx_shift.py
- test_embedding.py
- test_ignore_eos.py
- test_infill.py
- test_kv_keep_only_active.py
- test_lora.py
- test_mcp_servers.py
- test_metrics.py
- test_proxy.py
- test_rerank.py
- test_router.py
- test_security.py
- test_sleep.py
- test_slot_save.py
- test_speculative.py
- test_stream.py
- test_template.py
- test_tokenize.py
- test_tool_call.py
- test_tools_builtin.py
- test_vision_api.py
- .gitignore
- conftest.py
- pytest.ini
- README.md
- requirements.txt
- tests.sh
- utils.py
- CMakeLists.txt
- main.cpp
- README-dev.md
- README.md
- server-chat.cpp
- server-chat.h
- server-common.cpp
- server-common.h
- server-context.cpp
- server-context.h
- server-cors-proxy.h
- server-http.cpp
- server-http.h
- server-mcp.cpp
- server-mcp.h
- server-models.cpp
- server-models.h
- server-queue.cpp
- server-queue.h
- server-schema.cpp
- server-schema.h
- server-stream.cpp
- server-stream.h
- server-task.cpp
- server-task.h
- server-tools.cpp
- server-tools.h
- server.cpp
- CMakeLists.txt
- tokenize.cpp
- CMakeLists.txt
- README.md
- tts.cpp
- ModeWatcherDecorator.svelte
- TooltipProviderDecorator.svelte
- main.ts
- preview.ts
- install.sh
- pre-commit.sh
- pre-push.sh
- dev.sh
- favicon-colorize.ts
- make-icons-circular.js
- vite-plugin-build-info.ts
- vite-plugin-nerdamer.ts
- vite-plugin-relativize-base.ts
- vite-plugin-splash-screen.ts
- logo.svg
- ActionIcon.svelte
- ActionIconCopyToClipboard.svelte
- index.ts
- BadgeInfo.svelte
- BadgesModality.svelte
- index.ts
- ChatAttachmentsListItem.svelte
- ChatAttachmentsListItemMcpPrompt.svelte
- ChatAttachmentsListItemMcpResource.svelte
- ChatAttachmentsListItemThumbnailFile.svelte
- ChatAttachmentsListItemThumbnailImage.svelte
- ChatAttachmentsList.svelte
- ChatAttachmentsPreviewCurrentItem.svelte
- ChatAttachmentsPreviewCurrentItemAudio.svelte
- ChatAttachmentsPreviewCurrentItemImage.svelte
- ChatAttachmentsPreviewCurrentItemPdf.svelte
- ChatAttachmentsPreviewCurrentItemText.svelte
- ChatAttachmentsPreviewCurrentItemUnavailable.svelte
- ChatAttachmentsPreviewCurrentItemVideo.svelte
- ChatAttachmentsPreview.svelte
- ChatAttachmentsPreviewFileInfo.svelte
- ChatAttachmentsPreviewNavButtons.svelte
- ChatAttachmentsPreviewThumbnailStrip.svelte
- ChatFormActionAddButton.svelte
- ChatFormActionAddDropdown.svelte
- ChatFormActionAddMcpServersSubmenu.svelte
- ChatFormActionAddReasoningSubmenu.svelte
- ChatFormActionAddSheet.svelte
- ChatFormActionAddToolsSubmenu.svelte
- ChatFormActionsAdd.svelte
- ChatFormActionModels.svelte
- ChatFormActionRecord.svelte
- ChatFormActions.svelte
- ChatFormActionSubmit.svelte
- ChatFormContextGauge.svelte
- context-gauge.ts
- ContextGaugeDetailRow.svelte
- ContextGaugeDetails.svelte
- ContextGaugeDial.svelte
- ContextGaugeLoadModel.svelte
- ContextGaugePopup.svelte
- gauge-popup.svelte.ts
- ChatFormCurrentWorkingDirectory.svelte
- ChatFormCurrentWorkingDirectoryChip.svelte
- ChatFormCurrentWorkingDirectoryResultsList.svelte
- ChatFormInput.svelte
- ChatFormInputBasic.svelte
- ChatFormInputFileInputInvisible.svelte
- ChatFormInputRich.svelte
- ChatFormPickerItemHeader.svelte
- ChatFormPickerList.svelte
- ChatFormPickerListItem.svelte
- ChatFormPickerListItemSkeleton.svelte
- ChatForm.svelte
- ChatFormMcpResourcesList.svelte
- SKILL.md
- .gitignore
- app.css
- app.d.ts
- app.html
- .env.example
- .gitignore
- .npmrc
- .prettierignore
- .prettierrc
- CMakeLists.txt
- components.json
- embed.cpp
- eslint.config.js
- package-lock.json
- package.json
- playwright.config.ts
- pwa-assets-dark.config.ts
- pwa-assets.config.ts
- README.md
- sources.cmake
- CMakeLists.txt
- .clang-format
- .clang-tidy
- .dockerignore
- .ecrc
- .editorconfig
- .flake8
- .gitignore
- .gitmodules
- .pre-commit-config.yaml
- AGENTS.md
- AUTHORS
- build-xcframework.sh
- CLAUDE.md
- CMakeLists.txt
- CMakePresets.json
- CODEOWNERS
- CONTRIBUTING.md
- convert_hf_to_gguf.py
- convert_hf_to_gguf_update.py
- convert_llama_ggml_to_gguf.py
- convert_lora_to_gguf.py
- flake.nix
- LICENSE
- Makefile
- mypy.ini
- pyproject.toml
- pyrightconfig.json
- README.md
- requirements.txt
- SECURITY.md
# Installation Guide
git clone https://github.com/ggml-org/llama.cpp
Downloads the entire project code from GitHub to your computer.
cd llama.cpp
Moves into the project folder you just downloaded.
2. CMake
Medium Recommendedmkdir build && cd build
Creates a folder to hold the build output and moves into it.
cmake ..
Analyzes the source code and generates build configuration files (must be run inside the build folder).
make
Compiles the code based on the generated build configuration to produce an executable.
3. Node.js
Easycd tools/ui
This project's files live in a subfolder, so move into it first.
npm install
Downloads and installs the libraries listed in package.json.
npm start
Starts the development/run server.
4. Python
Easypip install -r requirements.txt
Installs the Python libraries listed in requirements.txt (or similar).
python <μ€νν νμΌλͺ
>.py # READMEμμ μ νν μ€ν νμΌλͺ
μ νμΈνμΈμ
Runs the Python script (or module).
5. Make
Medium- Git Needed to download the project code from GitHub.
- Make Usually pre-installed on Linux/macOS. On Windows, install separately (e.g. via MSYS2 or WSL).
make
Compiles the code based on the generated build configuration to produce an executable.
