提出视觉证据奖励机制,让视频推理模型真正看懂画面再思考。
When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
- 用强化学习奖励与视觉证据一致的推理过程。
- 在10个视频基准上表现领先,显著减少错误联想。
- 适合关注多模态推理可靠性的研究者和开发者。
视频推理旨在让机器通过多步逻辑从动态视觉内容中推断信息,对先进人工智能至关重要。尽管思维链(CoT)机制在文本任务中提升了推理能力,但其在视频理解中的应用仍不充分。本文系统分析发现,CoT常导致性能下降,产生冗长却误导的内部独白,引发幻觉性视觉细节并掩盖正确直觉——我们称之为“视觉思维漂移”。从贝叶斯视角解释,该漂移源于思维链偏离真实视觉证据,反而放大内部偏见或语言先验,使模型更像讲故事而非基于证据推理。为此,我们提出视觉证据奖励(VER),一种新型强化学习框架,明确奖励与视觉证据可验证一致的推理轨迹。在10个多样化视频理解基准上的全面评估表明,Video-VER持续取得最佳性能。本工作揭示了以视频为中心推理的独特挑战,推动构建能将推断牢固锚定于视觉证据的AI——让大型多模态模型不仅‘思而后答’,更能‘边思边看’。
原文摘要 · Abstract (English)
Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism has enhanced reasoning in text-based tasks, its application to video understanding remains underexplored. This paper presents a systematic analysis revealing that CoT often degrades performance in video reasoning, generating verbose but misleading internal monologues, and leading to hallucinated visual details and overridden correct intuitions - a phenomenon we term "visual thinking drift". We explain this drift through a Bayesian lens, positing that CoT traces often diverge from actual visual evidence, instead amplifying internal biases or language priors, causing models to storytell rather than engage in grounded reasoning. To counteract this, we introduce Visual Evidence Reward (VER), a novel reinforcement learning framework that explicitly rewards the generation of reasoning traces that are verifiably grounded in visual evidence. Comprehensive evaluation across 10 diverse video understanding benchmarks demonstrates that our Video-VER consistently achieves top performance. Our work sheds light on the distinct challenges of video-centric reasoning and encourages the development of AI that robustly grounds its inferences in visual evidence - for large multimodal models that not only "think before answering", but also "see while thinking".
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。