SALOVA通过分段检索提升长视频理解,解决大模型信息丢失问题。
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
- 构建场景连续的分段标注数据集SceneWalk,支持精准定位视频片段
- 设计动态路由与时空投影器,按查询高效检索相关视频段
- 在87.8万长视频上验证,显著提升长序列上下文一致性
尽管大型多模态模型取得进展,将其应用于长且未剪辑的视频内容仍面临上下文长度限制和高内存开销的挑战,导致信息丢失和响应相关性下降。随着网络平台视频数据指数级增长,理解长视频对推动通用智能至关重要。本文提出SALOVA:一种基于分段增强的长视频助手框架,通过目标检索机制提升长视频理解能力。针对两大挑战:(i) 构建高质量数据集SceneWalk,包含87.8万条长视频,每段密集标注,以捕捉场景连续性和丰富描述上下文;(ii) 设计集成动态路由机制与时空投影器的鲁棒架构,根据用户查询高效检索并处理相关视频片段。SALOVA通过精准识别与检索,显著提升生成回复的上下文相关性。大量实验表明,该框架在处理复杂长视频时具备更强能力,能有效维持长序列中的上下文完整性。
原文摘要 · Abstract (English)
Despite advances in Large Multi-modal Models, applying them to long and untrimmed video content remains challenging due to limitations in context length and substantial memory overhead. These constraints often lead to significant information loss and reduced relevance in the model responses. With the exponential growth of video data across web platforms, understanding long-form video is crucial for advancing generalized intelligence. In this paper, we introduce SALOVA: Segment-Augmented LOng Video Assistant, a novel video-LLM framework designed to enhance the comprehension of lengthy video content through targeted retrieval process. We address two main challenges to achieve it: (i) We present the SceneWalk dataset, a high-quality collection of 87.8K long videos, each densely captioned at the segment level to enable models to capture scene continuity and maintain rich descriptive context. (ii) We develop robust architectural designs integrating dynamic routing mechanism and spatio-temporal projector to efficiently retrieve and process relevant video segments based on user queries. Our framework mitigates the limitations of current video-LMMs by allowing for precise identification and retrieval of relevant video segments in response to queries, thereby improving the contextual relevance of the generated responses. Through extensive experiments, SALOVA demonstrates enhanced capability in processing complex long-form videos, showing significant capability to maintain contextual integrity across extended sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。