ray_vllm_inference

A simple service that integrates vLLM with Ray Serve for fast and scalable LLM serving.

// repository documentation