arXiv:2502.16690cs.AI2025-02被引 11

研究大模型如何用文本表示空间信息,提升导航能力

From Text to Space: Mapping Abstract Spatial Models in LLMs during a Grid-World Navigation Task

  • 用笛卡尔坐标等不同文本格式编码空间,对比导航表现
  • 笛卡尔表示成功率更高,模型越大效果越显著
  • 发现中间层神经元稳定响应空间特征,可跨任务复用

理解大语言模型(LLMs)如何表征和推理空间信息,对构建能导航真实与模拟环境的鲁棒智能体至关重要。本文研究不同基于文本的空间表征对模型在网格世界导航任务中性能及内部激活的影响。通过在需向目标前进的任务上评估多种规模的模型,考察空间信息编码方式对决策的影响。实验表明,笛卡尔坐标表示始终带来更高成功率与路径效率,且性能随模型规模显著提升。此外,对LLaMA-3.1-8B的探测显示,中间层存在一组神经元,无论空间信息如何表示,均稳定关联于代理位置、动作正确性等空间特征,并在无关空间推理任务中也被激活。本工作深化了对LLM处理空间信息的理解,为开发更可解释、更鲁棒的智能体AI系统提供了关键洞见。

原文摘要 · Abstract (English)

Understanding how large language models (LLMs) represent and reason about spatial information is crucial for building robust agentic systems that can navigate real and simulated environments. In this work, we investigate the influence of different text-based spatial representations on LLM performance and internal activations in a grid-world navigation task. By evaluating models of various sizes on a task that requires navigating toward a goal, we examine how the format used to encode spatial information impacts decision-making. Our experiments reveal that cartesian representations of space consistently yield higher success rates and path efficiency, with performance scaling markedly with model size. Moreover, probing LLaMA-3.1-8B revealed subsets of internal units, primarily located in intermediate layers, that robustly correlate with spatial features, such as the position of the agent in the grid or action correctness, regardless of how that information is represented, and are also activated by unrelated spatial reasoning tasks. This work advances our understanding of how LLMs process spatial information and provides valuable insights for developing more interpretable and robust agentic AI systems.

空间推理大模型机制智能体导航

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。