arXiv:2506.22385cs.CVcs.AI2025-06被引 3

让视频模型学会根据新信息推翻或坚持结论,提升动态推理能力。

Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment

  • 设计反事实思维链框架,结合语音转录与逻辑修正减少偏见。
  • 在新数据上,模型对强/弱化判断准确率提升18.3个百分点。
  • 适合研究多模态推理、可信AI的学者和开发者参考。

视频大模型虽在理解视频内容上取得显著进展,但常难以进行抽象与动态推理——即在新信息出现时修正原有判断。现实中的结论很少一成不变,新上下文可能强化或削弱初始推断。为此,我们提出可撤销视频蕴含(DVidE)任务,要求模型像怀疑者一样,依据不断变化的证据更新推理。在该任务中,给定视频前提和文本假设,模型需判断新信息是加强还是削弱假设(分类版),或生成能改变蕴含关系的连贯更新(生成版)。针对分类任务,提出反事实思维链框架,融合反事实推理、增强型语音转录与理由优化以降低偏见;针对生成任务,结合语音转录与大语言模型生成符合目标的强/弱化更新。此外,构建首个含强/弱化标注的新基准数据集,并设计基于LLM的评估指标以衡量生成性能。实验表明,所提方法显著提升模型动态推理能力。

原文摘要 · Abstract (English)

Video Large Multimodal Models (VLMMs) have made impressive strides in understanding video content, but they often struggle with abstract and adaptive reasoning-the ability to revise their interpretations when new information emerges. In reality, conclusions are rarely set in stone; additional context can strengthen or weaken an initial inference. To address this, we introduce Defeasible Video Entailment (DVidE), a new task that challenges models to think like doubters, constantly updating their reasoning based on evolving evidence. In DVidE, given a video premise and a textual hypothesis, models must determine whether a new update strengthens or weakens the hypothesis (classification version) or generate a coherent update that modifies the entailment relationship (generation version). For solving the classification task, we propose the Chain of Counterfactual Thought framework, utilizing counterfactual reasoning, ASR-enhanced video content, and rationale refinement to reduce inference bias. For the generation task, we develop a framework that combines ASR output with a Large Language Model (LLM) to produce coherent, contextually relevant updates aligned with the intended strengthener or weakener goals. Additionally, we introduce a novel benchmark dataset, with strengthener/weakener annotations and an LLM-based evaluation metric specifically designed for assessing generative performance. Experimental results demonstrate significant improvements, highlighting our proposed method in enhancing dynamic reasoning capabilities of VLMMs.

视频理解动态推理多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。