Tempo

Tempo: Small Vision-Language Models are Smart Compressors for Long Video Understanding, ECCV 2026

// repository documentation