arXiv:2506.04953cs.CV2025-06AAAI被引 8

提出自适应视觉信息检索框架,让大模型读懂小时级长视频。

APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval

  • 分层检索关键帧和视觉令牌,动态保留重要信息
  • 在三个数据集上提升9.7%、9.5%、4.6%准确率
  • 无需训练,适合资源受限的长视频理解任务

当前多模态大模型在处理小时级视频时面临信息量大、内存与计算资源瓶颈等问题。尽管近期无训练方法通过压缩视觉特征降低资源消耗,但依赖不完整视觉信息限制了性能。为此,我们提出无训练框架APVR,通过分层检索保留充分且关键的视觉信息。其核心由两个互补组件构成:基于查询扩展与迭代时空语义置信度评分的枢纽帧检索,以及在最多1024个枢纽帧内进行查询感知注意力驱动的令牌选择。该双粒度策略使模型能处理小时级视频并保持语义保真。在三种不同基线MLLM上的实验验证表明,APVR在LongVideoBench、VideoMME和MLVU上分别取得最高9.5%、4.6%和9.7%的性能提升,达到无训练与有训练方法的最优水平。

原文摘要 · Abstract (English)

Current multimodal large language models (MLLMs) struggle with hour-level video understanding, facing significant challenges not only in modeling the substantial information volume of long videos but also in overcoming the memory wall and resource constraints during both training and inference. Although recent training-free approaches have alleviated resource demands by compressing visual features, their reliance on incomplete visual information limits the performance potential. To address these limitations, we propose Adaptive Pivot Visual information Retrieval (APVR), a training-free framework that hierarchically retrieves and retains sufficient and important visual information. It breakthroughs the memory wall limitation via two complementary components: Pivot Frame Retrieval employs query expansion and iterative spatio-semantic confidence scoring to identify relevant video frames, and Pivot Token Retrieval performs query-aware attention-driven token selection within up to 1024 pivot frames. This dual granularity approach enables the processing of hour-long videos while maintaining semantic fidelity. Experimental validations on three different baseline MLLMs demonstrate significant performance improvements up to 9.5\%, 4.6\% and 9.7\% on LongVideoBench, VideoMME and MLVU, respectively. APVR achieves state-of-the-art results for both training-free and training-based approaches.

长视频理解视觉检索多模态模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。