arXiv:2505.24182cs.CVcs.AI2025-05被引 9

评测大模型在复杂物理场景下的视觉推理能力,发现顶尖模型仍严重不足。

Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT

  • 构建多图输入的视觉链式推理基准,要求模型逐步分析动态视觉线索。
  • 顶尖模型在物理推理任务上准确率低,且推理路径与图像不一致。
  • 揭示强化学习微调反而损害空间推理,挑战现有训练范式。

理解由运动定律、空间关系和因果性支配的物理世界,是多模态大语言模型(MLLM)面临的核心挑战。尽管OpenAI o3和GPT-4o等模型展现出出色的感知与推理能力,我们的研究发现它们在复杂场景中对基本物理规律、空间互动和因果效应的理解依然薄弱,尤其在需要多步推理时难以维持基于视觉证据的一致性推理链。为此,我们提出MVPBench,一个精心设计的基准,通过视觉链式思维(CoT)视角评估视觉物理推理能力。每个样本包含交错的多图像输入,不仅要求正确答案,还需基于不断演化的视觉线索生成连贯的逐步推理路径,模拟人类对真实物理过程的推演。为实现细粒度评估,我们引入基于图的CoT一致性度量,验证推理路径是否符合有效物理逻辑。同时,通过最小化文本先验带来的捷径,促使模型依赖视觉理解。实验结果揭示:即使是最先进的MLLMs,在物理领域也表现出较差的推理准确率和弱的图像-文本对齐。令人意外的是,常被认为能提升视觉推理性能的强化学习后训练,反而常损害空间推理,提示需重新审视当前微调策略。

原文摘要 · Abstract (English)

Understanding the physical world - governed by laws of motion, spatial relations, and causality - poses a fundamental challenge for multimodal large language models (MLLMs). While recent advances such as OpenAI o3 and GPT-4o demonstrate impressive perceptual and reasoning capabilities, our investigation reveals these models struggle profoundly with visual physical reasoning, failing to grasp basic physical laws, spatial interactions, and causal effects in complex scenes. More importantly, they often fail to follow coherent reasoning chains grounded in visual evidence, especially when multiple steps are needed to arrive at the correct answer. To rigorously evaluate this capability, we introduce MVPBench, a curated benchmark designed to rigorously evaluate visual physical reasoning through the lens of visual chain-of-thought (CoT). Each example features interleaved multi-image inputs and demands not only the correct final answer but also a coherent, step-by-step reasoning path grounded in evolving visual cues. This setup mirrors how humans reason through real-world physical processes over time. To ensure fine-grained evaluation, we introduce a graph-based CoT consistency metric that verifies whether the reasoning path of model adheres to valid physical logic. Additionally, we minimize shortcut exploitation from text priors, encouraging models to rely on visual understanding. Experimental results reveal a concerning trend: even cutting-edge MLLMs exhibit poor visual reasoning accuracy and weak image-text alignment in physical domains. Surprisingly, RL-based post-training alignment - commonly believed to improve visual reasoning performance - often harms spatial reasoning, suggesting a need to rethink current fine-tuning practices.

视觉推理多模态链式思维评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。