ComfyUI_VLM_nodes

(β˜… 585)

ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.

  • .gitignore
  • __init__.py
  • CHANGELOG.md
  • CITATION.cff
  • COMPATIBILITY.md
  • CONTRIBUTING.md
  • LICENSE
  • MODEL_VALIDATION.md
  • pyproject.toml
  • README.md
  • requirements-dev.txt
  • requirements-llama-cpp.txt
  • requirements-moondream31.txt
  • requirements-quantization.txt
  • requirements-robotics-client.txt
  • requirements.txt
  • SECURITY.md
  • vlmnodes.default.json
  • vlmnodes.json
// repository documentation