无需训练即可理解超长视频,通过持续记忆机制动态聚焦关键片段。
$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation
- 用连续时间记忆机制让模型处理无限制长度视频
- 在Video-LLaMA和VideoChat2上问答准确率显著提升
- 适合需要高效处理长视频的场景,如视频检索与分析
现有视频-语言模型受限于上下文长度和稀疏采帧,难以有效理解长视频,常导致信息丢失。本文提出∞-Video,通过连续时间长期记忆(LTM)凝聚机制,使模型能无训练地处理任意长度视频。该框架增强视频Q-former,实现对无限视频上下文的高效处理。通过连续注意力,模型动态分配更高粒度到最相关片段,形成随时间演化的“粘性”记忆。在Video-LLaMA和VideoChat2上的实验表明,该方法在视频问答任务中表现更优,展示了连续时间LTM机制在实现可扩展、免训练长视频理解方面的潜力。
原文摘要 · Abstract (English)
Current video-language models struggle with long-video understanding due to limited context lengths and reliance on sparse frame subsampling, often leading to information loss. This paper introduces $\infty$-Video, which can process arbitrarily long videos through a continuous-time long-term memory (LTM) consolidation mechanism. Our framework augments video Q-formers by allowing them to process unbounded video contexts efficiently and without requiring additional training. Through continuous attention, our approach dynamically allocates higher granularity to the most relevant video segments, forming "sticky" memories that evolve over time. Experiments with Video-LLaMA and VideoChat2 demonstrate improved performance in video question-answering tasks, showcasing the potential of continuous-time LTM mechanisms to enable scalable and training-free comprehension of long videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。