测试多模态模型视觉推理是否真实可信
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
- 设计新评测集,要求模型选出符合视觉证据和逻辑的唯一推理链
- 顶尖模型在该任务上表现不佳,生成流畅但不真实
- 适合研究视觉推理真实性、模型可信度的学者使用
能够进行链式思维(CoT)推理是多模态模型的重要里程碑,使其能解决复杂的视觉推理问题。然而关键问题仍存:这种推理是否真正基于视觉证据且逻辑自洽?现有评测侧重生成能力,忽视验证能力,即判断推理链是否视觉一致且逻辑有效。为此,我们提出MM-CoT,一个专门用于探测多模态模型在链式思维中视觉锚定性和逻辑连贯性的诊断性基准。模型无需生成自由文本解释,而是必须从多个选项中选出唯一同时满足两个正交约束的事件链:(i) 视觉一致性,确保每一步均有可观察证据支持;(ii) 逻辑连贯性,确保因果关系和常识合理。通过设计对抗性干扰项,使其中一条件被违反,从而暴露不同类型的推理错误。我们在主流视觉-语言模型上评估发现,即使最先进的系统也表现欠佳,揭示生成流畅性与真实推理一致性之间存在显著差距。MM-CoT与现有基准相关性低,证实其测量的是视觉锚定与逻辑推理的独特组合。该基准为未来开发在视觉世界中既合理又忠实推理的模型奠定了基础。
原文摘要 · Abstract (English)
The ability to perform Chain-of-Thought (CoT) reasoning marks a major milestone for multimodal models (MMs), enabling them to solve complex visual reasoning problems. Yet a critical question remains: is such reasoning genuinely grounded in visual evidence and logically coherent? Existing benchmarks emphasize generation but neglect verification, i.e., the capacity to assess whether a reasoning chain is both visually consistent and logically valid. To fill this gap, we introduce MM-CoT, a diagnostic benchmark specifically designed to probe the visual grounding and logical coherence of CoT reasoning in MMs. Instead of generating free-form explanations, models must select the sole event chain that satisfies two orthogonal constraints: (i) visual consistency, ensuring all steps are anchored in observable evidence, and (ii) logical coherence, ensuring causal and commonsense validity. Adversarial distractors are engineered to violate one of these constraints, exposing distinct reasoning failures. We evaluate leading vision-language models on MM-CoT and find that even the most advanced systems struggle, revealing a sharp discrepancy between generative fluency and true reasoning fidelity. MM-CoT shows low correlation with existing benchmarks, confirming that it measures a unique combination of visual grounding and logical reasoning. This benchmark provides a foundation for developing future models that reason not just plausibly, but faithfully and coherently within the visual world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。