Groma
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
// repository documentation
Was this content helpful?
(0 ratings)
