通过语义与视觉共识选择关键帧,提升长视频理解准确率
SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
- 用语义与视觉双分支协同筛选关键帧,兼顾时间依赖性
- 在多个长视频基准上超越现有方法,准确率提升显著
- 无需训练、适配任意模型,适合长视频分析任务
长视频理解因内容复杂、多样且时间分布分散而困难。尽管视频大语言模型(Video-LLMs)可处理长达数十分钟的视频,但应用于真正长序列时计算成本过高,常导致推理不聚焦或不一致。现有方法虽通过选取最信息量帧缓解问题,但通常忽略时间依赖性或仅依赖单模态证据,难以提供完整且查询相关的上下文。我们提出无训练、模型无关的语义-视觉共识证据选择框架SeViCES。其核心包含两个组件:(1)时序感知的语义分支利用大语言模型对字幕进行推理,(2)聚类引导的视觉分支通过互信息对齐嵌入与语义得分。答案共识优化模块进一步融合多模态证据,约束答案空间以消除预测不一致。在多个长视频理解基准上的实验表明,SeViCES在准确率和鲁棒性上持续优于当前最优方法,验证了共识驱动证据选择对视频大模型的重要性。
原文摘要 · Abstract (English)
Long video understanding remains challenging due to its complex, diverse, and temporally scattered content. Although video large language models (Video-LLMs) can process videos lasting tens of minutes, applying them to truly long sequences is computationally prohibitive and often leads to unfocused or inconsistent reasoning. A promising solution is to select only the most informative frames, yet existing approaches typically ignore temporal dependencies or rely on unimodal evidence, limiting their ability to provide complete and query-relevant context. We propose a Semantic-Visual Consensus Evidence Selection (SeViCES) framework for effective and reliable long video understanding. SeViCES is training-free and model-agnostic, and introduces two key components. The Semantic-Visual Consensus Frame Selection (SVCFS) module selects frames through (1) a temporal-aware semantic branch that leverages LLM reasoning over captions, and (2) a cluster-guided visual branch that aligns embeddings with semantic scores via mutual information. The Answer Consensus Refinement (ACR) module further resolves inconsistencies between semantic- and visual-based predictions by fusing evidence and constraining the answer space. Extensive experiments on long video understanding benchmarks show that SeViCES consistently outperforms state-of-the-art methods in both accuracy and robustness, demonstrating the importance of consensus-driven evidence selection for Video-LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。