Groma

★ 586 Open GitHub ↗

[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization

// repository documentation