arXiv:2603.03002cs.AI2026-03

用纯文本测试大模型的空间认知能力,发现其依赖语言习惯而非真实空间推理。

SpatialText: A Pure-Text Cognitive Benchmark for Spatial Understanding in Large Language Models

  • 通过真人描述与代码生成场景结合,构建纯文本空间推理评测框架。
  • 模型在全局坐标推理上表现尚可,但在视角转换和局部参考系上严重出错。
  • 适合研究大模型认知机制的学者,揭示其空间理解的本质局限。

真正的空间推理依赖于构建并操作连贯的内部空间表征,常被视作心智模型,而非仅处理表面语言关联。尽管大语言模型在多个领域表现出色,现有评测难以将这种内在空间认知与统计语言启发式区分开来。多模态评估也常混淆空间推理与视觉感知。为系统探究模型是否能构建灵活的空间心智模型,我们提出SpatialText——一个理论驱动的诊断框架。它不只作为数据集,而是通过双源方法分离文本空间推理:整合人类标注的真实三维室内环境描述(含自然模糊性、视角变化和功能关系),以及代码生成的逻辑精确场景,用于探测形式化空间推断与认识论边界。对主流模型的系统评估揭示了根本性表征局限:模型虽能检索显式空间事实并在全局非自我中心坐标系中运作,但在自我中心视角变换和局部参考系推理上出现关键错误。这些系统性失误强烈表明,当前模型高度依赖语言共现启发式,而非构建连贯、可验证的内部空间表征。因此,SpatialText成为诊断人工空间智能认知边界的严格工具。

原文摘要 · Abstract (English)

Genuine spatial reasoning relies on the capacity to construct and manipulate coherent internal spatial representations, often conceptualized as mental models, rather than merely processing surface linguistic associations. While large language models exhibit advanced capabilities across various domains, existing benchmarks fail to isolate this intrinsic spatial cognition from statistical language heuristics. Furthermore, multimodal evaluations frequently conflate genuine spatial reasoning with visual perception. To systematically investigate whether models construct flexible spatial mental models, we introduce SpatialText, a theory-driven diagnostic framework. Rather than functioning simply as a dataset, SpatialText isolates text-based spatial reasoning through a dual-source methodology. It integrates human-annotated descriptions of real 3D indoor environments, which capture natural ambiguities, perspective shifts, and functional relations, with code-generated, logically precise scenes designed to probe formal spatial deduction and epistemic boundaries. Systematic evaluation across state-of-the-art models reveals fundamental representational limitations. Although models demonstrate proficiency in retrieving explicit spatial facts and operating within global, allocentric coordinate systems, they exhibit critical failures in egocentric perspective transformation and local reference frame reasoning. These systematic errors provide strong evidence that current models rely heavily on linguistic co-occurrence heuristics rather than constructing coherent, verifiable internal spatial representations. SpatialText thus serves as a rigorous instrument for diagnosing the cognitive boundaries of artificial spatial intelligence.

空间认知大模型评测心智模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。