测试大模型能否从文本构建空间世界模型,发现跨语言推理存在普遍瓶颈。
Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning

- 设计多语言诊断基准MentalMap,分六级评估空间推理能力
- 发现所有模型在视角推理上出现性能断崖式下跌,准确率超40%后骤降
- 人类与模型表现一致,说明瓶颈源于纯文本记忆限制
大语言模型是否能仅从纯文本描述构建内部空间世界模型仍存争议,且跨语言迁移能力尚未系统研究。我们提出MentalMap,一个包含六级能力层级(L0-L5)的多语言诊断基准,涵盖从原子空间事实到生成式世界图谱的完整链条,并设置四个诊断维度:参考框架、阅读方向偏差、推理资源分配与幻觉。该基准基于100个ProcTHOR家庭场景,覆盖八种语言及结构化文本对照组,包含39类任务共1,950个评估单元。对十三个不同规模和架构的LLM进行评测,发现普遍存在L3推理断崖:当基础原子准确率超过40%后,视角推理性能降至不足一半。该现象在跨语言、跨规模与跨提示策略下持续存在,而结构化输出失败与推理模式则因模型而异。人类在相同纯文本协议下重现相同失败模式,表明瓶颈源于纯文本工作记忆限制,而非当前架构特有。研究将纯文本空间推理重新定义为多轴世界建模问题,推动未来采用多模态与草稿垫增强推理。
原文摘要 · Abstract (English)
Whether large language models (LLMs) construct internal spatial world models from pure-text descriptions remains contested, and whether such capabilities transfer across languages has not been systematically studied. We introduce MentalMap, a multilingual diagnostic benchmark with a six-level capability hierarchy (L0-L5) spanning atomic spatial facts to generative world-graph construction, together with four diagnostic axes probing frame of reference, reading-direction bias, reasoning-effort allocation, and hallucination. MentalMap is built from 100 ProcTHOR household scenes, covers eight typologically diverse languages plus a structured-text control, and contains 39 task families across 1,950 evaluation cells. Evaluating thirteen LLMs across scales and model families, we identify a universal L3 reasoning cliff: no model retains even half of its L0 performance on viewpoint reasoning once baseline atomic accuracy exceeds 40%. The cliff persists across languages, scales, and prompting strategies, while structured-output failures and reasoning patterns vary substantially across models. Human evaluation under the identical pure-text protocol reproduces the same failure pattern, suggesting that the bottleneck arises from text-only working memory constraints rather than being specific to current LLM architectures. Our findings reframe pure-text spatial reasoning as a multi-axis world-modeling problem and motivate multimodal and scratchpad-augmented reasoning as future directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。