arXiv:2504.06835cs.CV2025-04被引 7

轻量压缩框架LVC让视觉语言模型更好理解长视频。

LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding

  • 用查询注意力压缩机制解决模型采样稀疏问题
  • 仅用1万对短视频数据提升时序推理能力
  • 适合想低成本增强视频理解的开发者

长视频理解需兼顾空间细节与时间感知。尽管视觉语言模型(VLMs)通过多帧输入获得帧级理解能力,但稀疏采样导致信息丢失。相比之下,视频大语言模型(Video-LLMs)虽能捕捉视觉特征中的时间关系,却受限于高质量视频-文本数据集稀缺。为以极低数据与计算成本将长视频理解能力迁移至VLMs,我们提出轻量级视频压缩框架(LVC),其核心是查询注意力视频压缩机制,有效缓解VLMs的稀疏采样问题。仅用1万对短视频-文本对训练对齐层,LVC显著增强VLMs的时序推理能力。大量实验表明,LVC在多个模型上均实现稳定提升,包括InternVL2系列和Phi-3.5-Vision。其中,InternVL2-40B-LVC在长视频理解基准MLVU和Video-MME上分别取得68.2和65.9分,相对提升达14.6%和7.7%。模型与代码即将开源。

原文摘要 · Abstract (English)

Long video understanding is a complex task that requires both spatial detail and temporal awareness. While Vision-Language Models (VLMs) obtain frame-level understanding capabilities through multi-frame input, they suffer from information loss due to the sparse sampling strategy. In contrast, Video Large Language Models (Video-LLMs) capture temporal relationships within visual features but are limited by the scarcity of high-quality video-text datasets. To transfer long video understanding capabilities to VLMs with minimal data and computational cost, we propose Lightweight Video Compression (LVC), a novel method featuring the Query-Attention Video Compression mechanism, which effectively tackles the sparse sampling problem in VLMs. By training only the alignment layer with 10k short video-text pairs, LVC significantly enhances the temporal reasoning abilities of VLMs. Extensive experiments show that LVC provides consistent performance improvements across various models, including the InternVL2 series and Phi-3.5-Vision. Notably, the InternVL2-40B-LVC achieves scores of 68.2 and 65.9 on the long video understanding benchmarks MLVU and Video-MME, respectively, with relative improvements of 14.6% and 7.7%. The enhanced models and code will be publicly available soon.

视频理解轻量压缩VLMs时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。