arXiv:2507.15521cs.AI2025-07

用物理题测试大模型是否真会模拟世界,发现它靠数滑轮猜结果,但分不清力怎么传。

LLM world models are mental: Output layer evidence of brittle world model use in LLM mechanical reasoning

  • 用滑轮图测大模型能否理解机械原理,看它是否真在脑中模拟系统。
  • 模型能大致猜出机械优势,但只靠数滑轮数量,不真懂力的传递路径。
  • 适合研究大模型认知机制的人看,尤其关注它像不像真在‘想’问题。

大型语言模型(LLMs)是构建并操作内部世界模型,还是仅依赖输出层词元概率表示的统计关联?我们借鉴人类心智模型研究的认知科学方法,使用TikZ渲染的滑轮系统图像测试LLMs。实验1考察模型估算机械优势(MA)的能力,当前最佳模型表现略高于随机水平且与真实MA显著相关;其估计值与滑轮数量的相关性表明,模型采用了滑轮计数启发法,而非真正模拟系统以得出精确值。实验2通过对比功能连通的滑轮系统与组件随机排列的伪系统,检验模型对关键全局特征的表征能力:无明确提示下,模型以F1=0.8识别出功能系统具有更高机械优势,说明其具备一定空间关系表征能力。实验3进一步要求模型区分功能系统与结构连接但无力传递至重物的匹配系统,模型仅以F1=0.46识别正确,接近随机水平。这些发现表明,大模型可能具备足够能力来利用滑轮数量与机械优势间的统计关联(实验1),以及近似表征组件空间关系(实验2),但在处理精细结构连通性推理方面存在局限(实验3)。研究主张采用认知科学方法评估人工智能系统的世界建模能力。

原文摘要 · Abstract (English)

Do large language models (LLMs) construct and manipulate internal world models, or do they rely solely on statistical associations represented as output layer token probabilities? We adapt cognitive science methodologies from human mental models research to test LLMs on pulley system problems using TikZ-rendered stimuli. Study 1 examines whether LLMs can estimate mechanical advantage (MA). State-of-the-art models performed marginally but significantly above chance, and their estimates correlated significantly with ground-truth MA. Significant correlations between number of pulleys and model estimates suggest that models employed a pulley counting heuristic, without necessarily simulating pulley systems to derive precise values. Study 2 tested this by probing whether LLMs represent global features crucial to MA estimation. Models evaluated a functionally connected pulley system against a fake system with randomly placed components. Without explicit cues, models identified the functional system as having greater MA with F1=0.8, suggesting LLMs could represent systems well enough to differentiate jumbled from functional systems. Study 3 built on this by asking LLMs to compare functional systems with matched systems which were connected up but which transferred no force to the weight; LLMs identified the functional system with F1=0.46, suggesting random guessing. Insofar as they may generalize, these findings are compatible with the notion that LLMs manipulate internal world models, sufficient to exploit statistical associations between pulley count and MA (Study 1), and to approximately represent system components' spatial relations (Study 2). However, they may lack the facility to reason over nuanced structural connectivity (Study 3). We conclude by advocating the utility of cognitive scientific methods to evaluate the world-modeling capacities of artificial intelligence systems.

大模型世界模型机械推理认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。