arXiv:2507.16572cs.CL2025-07被引 2

测试大模型对物理常识的理解能力,发现视觉与语言信息融合是短板。

Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models

  • 用GRASP和IntPhys 2数据集评估多模态大模型的物理推理能力
  • 模型在判断场景是否合理时准确率不高,尤其复杂任务更差
  • 关键问题是视觉与语言模块间信息未有效对齐,适合关注多模态融合的研究者

本文系统评估了当前最先进的多模态大语言模型(MLLMs)在直观物理任务上的表现,使用GRASP和IntPhys 2数据集。我们测试了开源模型InternVL 2.5、Qwen 2.5 VL、LLaVA-OneVision以及专有模型Gemini 2.0 Flash Thinking,发现即使最新模型也难以可靠区分物理上合理与不合理的情境。为进一步超越性能指标,我们对模型嵌入进行探针分析,提取关键处理阶段的中间表示,考察任务相关信息的保留程度。结果表明,根据任务难度,会出现显著的视觉-语言错位:视觉编码器能有效捕捉物理合理性线索,但语言模型未能充分利用这些信息,导致推理失败。这一错位表明,MLLM在直观物理任务中的主要瓶颈并非视觉部分,而是视觉与语言信息的无效整合。研究强调了视觉-语言对齐的重要性,为未来多模态模型的发展提供了重要启示。

原文摘要 · Abstract (English)

This paper presents a systematic evaluation of state-of-the-art multimodal large language models (MLLMs) on intuitive physics tasks using the GRASP and IntPhys 2 datasets. We assess the open-source models InternVL 2.5, Qwen 2.5 VL, LLaVA-OneVision, and the proprietary Gemini 2.0 Flash Thinking, finding that even the latest models struggle to reliably distinguish physically plausible from implausible scenarios. To go beyond performance metrics, we conduct a probing analysis of model embeddings, extracting intermediate representations at key processing stages to examine how well task-relevant information is preserved. Our results show that, depending on task difficulty, a critical vision-language misalignment can emerge: vision encoders successfully capture physical plausibility cues, but this information is not effectively utilized by the language model, leading to failures in reasoning. This misalignment suggests that the primary limitation of MLLMs in intuitive physics tasks is not the vision component but the ineffective integration of visual and linguistic information. Our findings highlight vision-language alignment as a key area for improvement, offering insights for future MLLMs development.

多模态物理理解模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。