Groma

โ˜… 586 Open GitHub โ†—

[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization

// repository documentation