arXiv:2509.12263cs.AIcs.LG2025-09中稿 · TMLR

提出首个评估大模型归纳物理推理能力的基准,发现其在新物理场景下表现差。

InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning

  • 设计合成视频基准,测试模型对未见物理规律的推理能力
  • 13个主流模型均在新物理场景中表现不佳,准确率显著下降
  • 模型易受语言干扰,忽视视觉信息,可靠性存疑

大型多模态模型(LMMs)将训练中观察到的物理规律(如动量守恒)编码为参数化知识,可基于视觉输入回答物理推理问题。然而,由于参数化知识仅包含训练中见过的物理规律,无法应对训练中未见的新物理环境。人类可通过演示适应新物理规律进行归纳推理,这一能力对安全关键应用中的模型替代至关重要。现有视觉评测仅关注参数化知识,忽略归纳推理。为此,我们提出InPhyRe——首个评估LMM归纳物理推理能力的视觉问答基准。InPhyRe通过算法生成的合成视频,评估模型对碰撞事件结果的预测能力。检测超过13个开源与专有LMM后发现:(1)LMMs难以将有限的通用物理规律知识应用于推理;(2)当推理场景涉及训练中未见的物理规律时,归纳推理能力极弱;(3)模型存在语言偏差,可能忽略视觉输入,质疑其对视觉信息的信任度。

原文摘要 · Abstract (English)

Large multimodal models (LMMs) encode physical laws observed during training, such as momentum conservation, as parametric knowledge. It allows LMMs to answer physical reasoning queries, such as the outcome of a potential collision event from visual input. However, since parametric knowledge includes only the physical laws seen during training, it is insufficient for reasoning in inference scenarios that follow physical laws unseen during training. In such novel physical environments, humans could adapt their physical reasoning based on provided demonstrations. This inductive physical reasoning ability is indispensable for LMMs if they are to replace human agents in safety-critical applications. Despite its importance, existing visual benchmarks do not evaluate inductive physical reasoning and only consider the parametric knowledge in LMMs. To this end, we propose InPhyRe, the first visual question answering benchmark to measure inductive physical reasoning in LMMs. InPhyRe evaluates LMMs' ability to predict the outcome of collision events in algorithmically generated synthetic videos. By inspecting over 13 open-source and proprietary LMMs, InPhyRe informs us that (1) LMMs struggle to apply their limited parametric knowledge about universal physical laws to reasoning, (2) inductive physical reasoning in LMMs is weak when the physical laws underlying inference scenarios were unseen during training, and (3) inductive physical reasoning in LMMs suffers from language bias and may ignore the visual inputs, questioning the trustworthiness of LMMs regarding visual inputs.

物理推理多模态评估基准模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。