Video-LLaVA

โ˜… 3,498 Open GitHub โ†—

ใ€EMNLP 2024๐Ÿ”ฅใ€‘Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

// repository documentation