测试大模型能否像人一样理解世界变化,发现它们在运动预测上几乎全靠猜。
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation
- 用6个模拟环境设计23项原子级测试,评估视觉、空间、时间等感知与预测能力
- 15个主流模型在运动轨迹判断中接近随机准确率,平均仅48%正确
- 模型存在错误关联(如误认蓝色物体更快),缺乏对世界的独立理解
内部世界模型(WM)使智能体能够理解世界状态并预测变化,是高级推理的基础。近期大型视觉-语言模型(如OpenAI o3、GPT-4o、Gemini)展现出作为通用世界模型的潜力。尽管已有研究评估其在特定能力上的表现,但对其基础世界建模能力的系统性测评仍为空白。借鉴比较心理学与认知科学,我们提出两阶段评估框架,涵盖感知(视觉、空间、时间、数量、运动)与预测(机制模拟、传递推理、组合推理),实现对视觉-语言模型作为世界模型的原子级评测。基于此框架,我们构建了包含23个细粒度维度的大型基准WM-ABench,覆盖6个多样化模拟环境,并设置受控反事实场景。在15个最新商业与开源模型上进行660次实验,结果显示这些模型在基本世界建模能力上存在显著缺陷:几乎所有模型在区分运动轨迹时表现接近随机水平(准确率约48%);同时缺乏解耦理解——部分模型错误认为蓝色物体比绿色物体运动更快。更丰富的分析揭示了模型与人类世界建模能力间的巨大差距。
原文摘要 · Abstract (English)
Internal world models (WMs) enable agents to understand the world's state and predict transitions, serving as the basis for advanced deliberative reasoning. Recent large Vision-Language Models (VLMs), such as OpenAI o3, GPT-4o and Gemini, exhibit potential as general-purpose WMs. While the latest studies have evaluated and shown limitations in specific capabilities such as visual understanding, a systematic evaluation of VLMs' fundamental WM abilities remains absent. Drawing on comparative psychology and cognitive science, we propose a two-stage framework that assesses Perception (visual, spatial, temporal, quantitative, and motion) and Prediction (mechanistic simulation, transitive inference, compositional inference) to provide an atomic evaluation of VLMs as WMs. Guided by this framework, we introduce WM-ABench, a large-scale benchmark comprising 23 fine-grained evaluation dimensions across 6 diverse simulated environments with controlled counterfactual simulations. Through 660 experiments on 15 latest commercial and open-source VLMs, we find that these models exhibit striking limitations in basic world modeling abilities. For instance, almost all models perform at near-random accuracy when distinguishing motion trajectories. Additionally, they lack disentangled understanding -- e.g., some models tend to believe blue objects move faster than green ones. More rich results and analyses reveal significant gaps between VLMs and human-level world modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。