让视频大模型看更多、想更深,提升长视频理解能力
See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding

- 通过动态扩展查询获取更全面的视觉证据
- 用答案特异性视觉反馈验证推理过程,提升准确性
- 适合需要深度视频理解的科研与工业场景
视频大语言模型在长视频理解任务上取得进展,但现有方法仍存在两大局限:证据获取依赖单一搜索意图,答案生成缺乏有效的视觉反馈机制。为此,我们提出CoVER框架,使视频大模型能够‘看得更全’——通过动态收集查询扩展的视觉证据;‘想得更深’——利用答案特异性视觉反馈验证草稿答案。这两种机制共同推动长视频理解从以答案为中心转向以证据为中心且可视觉验证的推理模式。实验表明,CoVER-7B在相同参数规模下显著优于其他模型,并在部分指标上超越当前最先进的闭源模型。
原文摘要 · Abstract (English)
Recent advances in Video Large Language Models (Video-LLMs) have enabled performance on long-video understanding tasks. However, existing methods still face two key limitations: evidence acquisition often relies on a single search intent, and answer generation lacks an effective visual feedback mechanism. To address these limitations, we propose \textbf{CoVER}, a Comprehensive Visual Evidence and Reflection framework for long-video understanding. CoVER enables Video-LLMs to \textbf{See More} by dynamically gathering query-expanded visual evidence, and \textbf{Think Deeper} by verifying draft answers with effective answer-specific visual feedback. Together, these mechanisms shift long-video understanding from answer-centric generation to evidence-centric and visually verifiable reasoning. Experimental results show that CoVER-7B substantially outperforms models with the same parameter scale and even surpasses state-of-the-art closed-source models on certain metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。