轻量级动态融合提升长视频理解效率与精度
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
- 动态加权多帧融合,自适应保留关键信息
- 关键帧选择策略提升压缩效果,降低冗余
- 无需大量人工标注,训练数据生成更高效
视觉大语言模型(VLLMs)在视频理解方面取得显著进展,但视频数据复杂性与上下文处理限制仍制约长视频理解。现有方法常通过视频特征压缩减少输入到大语言模型的令牌数,但许多方法未能有效聚焦关键特征,导致帧间冗余信息残留,或引入高计算开销模块。为此,我们提出FiLA-Video框架,采用轻量级动态权重多帧融合策略,自适应将多帧整合为单一表示,同时保留关键视频信息并降低计算成本。为优化融合帧的选择,引入关键帧选择机制,从更大帧池中识别出有信息量的帧以提升摘要质量。此外,提出一种简单有效的长视频训练数据生成策略,在无需大量人工标注的情况下提升模型性能。实验表明,相较于现有方法,FiLA-Video在长视频理解任务中实现了更高的效率与准确率。
原文摘要 · Abstract (English)
Recent advancements in video understanding within visual large language models (VLLMs) have led to notable progress. However, the complexity of video data and contextual processing limitations still hinder long-video comprehension. A common approach is video feature compression to reduce token input to large language models, yet many methods either fail to prioritize essential features, leading to redundant inter-frame information, or introduce computationally expensive modules.To address these issues, we propose FiLA(Fine-grained Vision Language Model)-Video, a novel framework that leverages a lightweight dynamic-weight multi-frame fusion strategy, which adaptively integrates multiple frames into a single representation while preserving key video information and reducing computational costs. To enhance frame selection for fusion, we introduce a keyframe selection strategy, effectively identifying informative frames from a larger pool for improved summarization. Additionally, we present a simple yet effective long-video training data generation strategy, boosting model performance without extensive manual annotation. Experimental results demonstrate that FiLA-Video achieves superior efficiency and accuracy in long-video comprehension compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。