用证据链框架提升视频理解的准确率与效率,防止模型胡说八道。
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
- 引入证据链框架,分步提取关键视觉证据并锚定推理过程。
- 在5个基准上达到新最好性能,准确率显著超越现有方法。
- 适合需要高可信度视频分析的研究者与开发者使用。
大型视觉语言模型在视频推理中面临根本矛盾:详尽推理计算成本过高,而快速推理又易产生幻觉。为此,我们提出链式证据(CoE)框架,从架构上解耦感知定位与推理效率,并协同优化。CoE包含两项核心创新:(1) 轻量级证据定位模块(EGM),作为查询引导过滤器,动态识别并提取高保真、紧凑的视觉证据;(2) 基于强化学习优化的证据锚定协议。我们设计复合奖励机制,强制模型在推断过程中严格参考已识别的时间锚点,有效抑制幻觉。为支持该方法,我们构建了大规模数据集CoE-Instruct(16.4万样本),采用新型双标注范式分别监督感知与推理。在五个基准(包括Video-MME、MVBench和VSI-Bench)上的大量实验表明,增强后的模型达到新的最先进水平,在准确率上显著优于现有方法,证明CoE是一种强大且实用的可靠视频理解范式。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches. To resolve this, we introduce the Chain of Evidence (CoE), a novel framework that architecturally decouples and co-optimizes perceptual grounding and reasoning efficiency. CoE incorporates two core innovations: (1) A lightweight Evidence Grounding Module (EGM) that acts as a query-guided filter, dynamically identifying and extracting a compact set of high-fidelity visual evidence; and (2) An Evidence-Anchoring Protocol optimized via Reinforcement Learning. Crucially, we design a composite reward mechanism that enforces process alignment, compelling the model to strictly reference identified temporal anchors during deduction, thereby mitigating hallucinations. To enable this, we construct CoE-Instruct, a large-scale dataset (164k samples) featuring a novel dual-annotation schema for separate perception and reasoning supervision. Extensive experiments on five benchmarks, including Video-MME, MVBench, and VSI-Bench, demonstrate that CoE-enhanced models establish a new state-of-the-art. They significantly outperform existing methods in accuracy, proving CoE to be a powerful and practical paradigm for reliable video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。