llama_cpp_ex
Elixir bindings for llama.cpp — run LLMs locally with Metal, CUDA, Vulkan, or CPU. Streaming, chat templates, embeddings, structured output, and concurrent batched inference.
// repository documentation
Was this content helpful?
(0 ratings)