arXiv:2512.24504cs.AI2025-12被引 1

测试大模型在地图上的探索、记忆与推理能力,发现记忆结构比模型大小更重要。

Thinking on Maps: How Foundation Model Agents Explore, Remember, and Reason Map Environments

  • 设计交互式框架,让模型逐步探索局部地图并积累空间经验。
  • 图结构记忆显著提升路径规划等复杂任务表现,优于普通记忆方式。
  • 模型能力达阈值后继续扩容无效,需专门优化空间推理机制。

地图环境是表达空间结构的基础媒介。理解基础模型(FM)在其中的感知与行为,对实现可靠的基于地图的推理至关重要。然而,现有评估多依赖静态地图或文本查询,忽略了空间理解的交互性与经验积累过程。本文提出一种交互式评估框架,分析FM代理在符号化地图环境中的探索、记忆与推理能力。代理在部分可观测的网格地图中逐步探索,仅接收局部观测信息。通过六类空间任务评估空间理解能力。系统地改变探索策略、记忆表示和推理方法,在多个基础模型上进行测试,揭示了各组件的功能差异:探索主要影响经验获取,对最终推理准确率影响有限;而记忆表示起核心作用,特别是序列与图结构的记忆形式,显著提升路径规划等结构密集型任务表现;推理方式进一步影响知识使用效率,高级提示支持更有效的多步推理。此外,观察到空间推理性能在模型规模达到一定阈值后趋于饱和,表明提升地图理解需针对性设计空间表征与推理机制,而非单纯扩大模型规模。

原文摘要 · Abstract (English)

Map environments provide a fundamental medium for representing spatial structure. Understanding how foundation model (FM) agents understand and act in such environments is therefore critical for enabling reliable map-based reasoning and applications. However, most existing evaluations of spatial ability in FMs rely on static map inputs or text-based queries, overlooking the interactive and experience-driven nature of spatial understanding.In this paper, we propose an interactive evaluation framework to analyze how FM agents explore, remember, and reason in symbolic map environments. Agents incrementally explore partially observable grid-based maps consisting of roads, intersections, and points of interest (POIs), receiving only local observations at each step. Spatial understanding is then evaluated using six kinds of spatial tasks. By systematically varying exploration strategies, memory representations, and reasoning schemes across multiple foundation models, we reveal distinct functional roles of these components. Exploration primarily affects experience acquisition but has a limited impact on final reasoning accuracy. In contrast, memory representation plays a central role in consolidating spatial experience, with structured memories particularly sequential and graph-based representations, substantially improving performance on structure-intensive tasks such as path planning. Reasoning schemes further shape how stored spatial knowledge is used, with advanced prompts supporting more effective multi-step inference. We further observe that spatial reasoning performance saturates across model versions and scales beyond a certain capability threshold, indicating that improvements in map-based spatial understanding require mechanisms tailored to spatial representation and reasoning rather than scaling alone.

空间推理大模型地图理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。