arXiv:2510.15440cs.CVcs.AI2025-10被引 4

让视频模型少选帧、多推理,提升关键证据的准确性。

Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning

  • 用强化学习动态选关键帧,并局部重采样获取细节。
  • 7B模型在多个基准上达59.8%~69.0%准确率,超越开源模型。
  • 适合需要高精度视频理解的科研与工业应用。

长视频推理仍是视频大模型的主要挑战,静态均匀采帧导致信息稀释,掩盖关键证据。现有像素空间视频推理代理因缺乏严格奖励机制保障证据纯度,且无法在预采帧外补充时序信息,表现不佳。为此,我们提出基于“少选多思”理念的证据优先自适应框架,核心为证据感知强化学习(EARL)。EARL使模型成为主动证据探查者,动态选择最相关帧,并在关键帧附近进行局部重采样以获取精细时序细节。在五个高难度视频推理基准上的实验证明,经EARL训练的模型达到开源视频大模型新最优水平,同时学习到高效且高纯度的视觉证据选择策略。令人印象深刻的是,7B模型在LongVideoBench上达59.8%,MVBench上达69.0%,VideoMME上达64.9%。结果凸显了证据纯度优先的重要性及框架的有效性。

原文摘要 · Abstract (English)

Long-form video reasoning remains a major challenge for Video Large Language Models (Video LLMs), as static uniform frame sampling leads to information dilution and obscures critical evidence. Furthermore, existing pixel-space video reasoning agents, which are designed to actively interact with the video to acquire new visual information, remain suboptimal due to their lack of rigorous reward mechanisms to enforce evidence purity and their inability to perform temporal information supplementation beyond pre-sampled frames. To address this critical gap, we propose a novel evidence-prioritized adaptive framework built upon our core philosophy: "Select Less, Reason More." Our core contribution is the evidence-aware reinforcement learning (EARL) framework, which transforms the model into an active interrogator of evidence. EARL is precisely engineered to dynamically select the most relevant frames and, crucially, to perform localized re-sampling around the selected key frames to access fine-grained temporal detail. Extensive experiments on five demanding video reasoning benchmarks demonstrate that our EARL-trained model achieves new state-of-the-art among open-source Video LLMs, simultaneously learning an effective and high-purity visual evidence selection policy. Impressively, our 7B model achieves 59.8% on LongVideoBench, 69.0% on MVBench and 64.9% on VideoMME. These results highlight the importance of prioritizing evidence purity and the effectiveness of our framework.

视频推理强化学习证据选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。