memra

(β˜… 319)

Rust + CUDA inference engine for NVIDIA RTX PRO 6000 Blackwell and RTX 5090. Serves safetensors and GGUF over an OpenAI-compatible API, with per-device tuned defaults and speculative decode gated byte-identical to plain decode. Hosted instance: inference.tiyuvta.ai

Showing a partial file list β€” download the zip above to see everything.

  • .gitattributes
  • .gitignore
  • .ignore
  • AGENTS.md
  • ARCHITECTURE-H100.md
  • ARCHITECTURE.md
  • bench_vllm.py
  • Cargo.lock
  • Cargo.toml
  • CITATION.cff
  • CLAUDE.md
  • CONTRIBUTING.md
  • LICENSE
  • README.md
  • SECURITY.md
// repository documentation