测试大模型在网格空间中的推理能力,发现其表现随复杂度上升急剧下降。
Stuck in the Matrix: Probing Spatial Reasoning in Large Language Models
- 设计五类网格任务,从基础识别到多步推理解析空间关系。
- 小规模任务准确率超50%,复杂度提升后平均下降42.7%,最高达84%。
- 揭示大模型缺乏稳健的空间表征,适合关注认知局限的研究者。
本文通过一组五项任务系统探查大语言模型(LLMs)在文本输入下的空间推理能力,涵盖象限识别、几何变换、距离评估、词搜索和方块滑动等,在结构化网格环境中测试其基础空间理解和多步问题求解能力。任务复杂度通过增加网格尺寸逐步提升,要求模型超越简单模式识别,实现抽象空间推理。结果显示,尽管在小规模任务中模型表现中等,但随着复杂度增加,准确率迅速下降:平均损失达42.7%,最高达84%;所有初始准确率超过50%的任务均出现至少48%的降幅,表明性能退化具有普遍性。这暗示模型底层架构缺乏稳健的空间表示。论文揭示了语言与空间推理之间的差距,为未来语言与几何交叉领域的整合性评测提供了基础。
原文摘要 · Abstract (English)
This paper explores the spatial reasoning capability of large language models (LLMs) over textual input through a suite of five tasks aimed at probing their spatial understanding and computational abilities. The models were tested on both fundamental spatial reasoning and multi-step problem-solving within structured grid-based environments using tasks such as quadrant identification, geometric transformations, distance evaluation, word searches, and tile sliding. Each task was scaled in complexity through increasing grid dimensions, requiring models to extend beyond simple pattern recognition into abstract spatial reasoning. Our results reveal that while LLMs demonstrate moderate success in all tasks with small complexity and size, performance drops off rapidly as scale increases, with an average loss in accuracy of 42.7%, and reaching as high as 84%. Every test that began with over 50% accuracy showed a loss of at least 48%, illustrating the consistent nature of the deterioration. Furthermore, their struggles with scaling complexity hint at a lack of robust spatial representations in their underlying architectures. This paper underscores the gap between linguistic and spatial reasoning in LLMs, offering insights into their current limitations, and laying the groundwork for future integrative benchmarks at the intersection of language and geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。