arXiv:2603.20194cs.CV2026-03被引 5

评测视频生成模型的因果连贯性,发现提示类型影响推理一致性。

MME-CoF-Pro: Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints

  • 构建包含16类303个样本的基准,评估模型推理过程连贯性。
  • 文本提示提升表面正确率但引发幻觉和不一致,视觉提示对细节感知无效。
  • 适合关注视频生成可靠性与提示设计的研究者使用。

视频生成模型展现出初步的推理能力,确保生成事件在帧间保持因果一致性对可靠部署至关重要,我们将其定义为推理连贯性。为弥补现有文献中缺乏推理连贯性评估的空白,我们提出MME-CoF-Pro,一个全面的视频推理评估基准。该基准包含16个类别共303个样本,涵盖从视觉逻辑到科学推理的任务。引入推理得分(Reasoning Score)作为评估指标,衡量必要中间推理步骤的完整性,并设置三种评估场景:(a) 无提示,(b) 文本提示,(c) 视觉提示,以可控方式探究提示引导的内在机制。对7个开源与闭源视频模型的评估揭示:(1) 视频生成模型推理连贯性弱,且与生成质量无关;(2) 文本提示虽提高表面正确性,但常导致不一致和幻觉推理;(3) 视觉提示在结构化感知任务中有益,但在细粒度感知上表现不佳。

原文摘要 · Abstract (English)

Video generative models show emerging reasoning behaviors. It is essential to ensure that generated events remain causally consistent across frames for reliable deployment, a property we define as reasoning coherence. To bridge the gap in literature for missing reasoning coherence evaluation, we propose MME-CoF-Pro, a comprehensive video reasoning benchmark to assess reasoning coherence in video models. Specifically, MME-CoF-Pro contains 303 samples across 16 categories, ranging from visual logical to scientific reasoning. It introduces Reasoning Score as evaluation metric for assessing process-level necessary intermediate reasoning steps, and includes three evaluation settings, (a) no hint (b) text hint and (c) visual hint, enabling a controlled investigation into the underlying mechanisms of reasoning hint guidance. Evaluation results in 7 open and closed-source video models reveals insights including: (1) Video generative models exhibit weak reasoning coherence, decoupled from generation quality. (2) Text hints boost apparent correctness but often cause inconsistency and hallucinated reasoning (3) Visual hints benefit structured perceptual tasks but struggle with fine-grained perception. Website: https://video-reasoning-coherence.github.io/

视频生成推理连贯性提示设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。