arXiv:2504.01407cs.CVcs.AI2025-04被引 10

通过动态聚焦关键片段,提升长视频理解效率与精度

ZoomV: Temporal Zoom-in for Efficient Long Video Understanding

  • 基于查询动态筛选视频关键时段,避免全帧处理
  • 在Charades-STA上实现11.8%的mIoU提升,在LVBench上提升9.7%
  • 适合需要高效处理长视频的应用场景

长视频理解对大视频-语言模型(LVLMs)构成根本挑战,因其帧数庞大且简单下采样易丢失关键上下文。受人类在手机上观看视频时不断缩放关注点的启发,我们提出ZoomV,一种查询感知的时序局部放大框架,用于高效准确的长视频理解。具体分为三阶段:(1) 时序兴趣定位:由查询引导,检索相关事件及其时间窗口作为候选;(2) 事件兴趣聚焦:在候选窗口池中,通过模型自身反思评分并筛选,高置信度窗口更具代表性;(3) 紧凑表征:对选定事件编码并时序下采样,保留关键语义同时大幅减少冗余。大量实验表明,ZoomV显著优于以往视频智能体方法。在时序定位任务上,其释放了LVLM的潜在能力,在Charades-STA上实现11.8%的mIoU提升。更显著的是,在LVBench上准确率提升9.7%,凸显其在长视频基准上的有效性。

原文摘要 · Abstract (English)

Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to the overwhelming number of frames and the risk of losing essential context through naive downsampling. Inspired by the way humans watch videos on mobile phones, constantly zooming in on frames of interest, we propose ZoomV, a query-aware temporal zoom-in framework designed for efficient and accurate long video understanding. Specifically, ZoomV operates in three stages: (1) Temporal interests grounding: guided by the query, ZoomV retrieves relevant events and their associated temporal windows as candidates. (2) Event interests spotlighting: within pools of candidate windows, each window is scored through the model itself reflection and filtered accordingly, where higher-confidence windows are more representative. (3) Compact representation: the selected events are encoded and temporally downsampled to preserve critical semantics while significantly reducing redundancy. Extensive experiments demonstrate that ZoomV substantially outperforms prior video agent approaches. On temporal grounding, ZoomV unlocks the latent capability of LVLMs, achieving an 11.8% mIoU gain on Charades-STA. Remarkably, ZoomV further boosts accuracy on LVBench by 9.7%, underscoring its effectiveness on long-video benchmarks.

视频理解长视频高效推理时序聚焦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。