测试大模型在网格上的空间推理能力,发现其对视角和三维形状识别仍有短板。
Grid Spatial Understanding: A Dataset for Textual Spatial Reasoning over Grids, Embodied Settings, and Coordinate Structures
- 仅用文本构建网格空间任务,剥离视觉干扰,专注评估空间推理。
- 多数模型难以理解以代理为中心的参考系和坐标列表中的3D形状。
- 小模型微调即可逼近前沿模型表现,适合构建专用智能体。
我们提出GSU,一个仅含文本的网格空间数据集,用于评估大语言模型在导航、物体定位和结构组合三大核心任务上的空间推理能力。通过去除视觉输入,将空间推理与感知分离,实验表明:尽管多数模型能掌握基本网格概念,但在基于具身代理的参考系理解及从坐标列表识别三维形状方面仍存在明显困难。此外,接触过视觉模态的多模态模型并未展现出可泛化的三维空间理解能力。最后,我们发现最新前沿模型虽能解决大部分任务(但复杂变体仍具挑战),而微调小型语言模型或使用LoRA微调小型LLM已可达到接近前沿模型的表现,为构建专用具身智能体提供了可行路径。
原文摘要 · Abstract (English)
We introduce GSU, a text-only grid dataset to evaluate the spatial reasoning capabilities of LLMs over 3 core tasks: navigation, object localization, and structure composition. By forgoing visual inputs, isolating spatial reasoning from perception, we show that while most models grasp basic grid concepts, they struggle with frames of reference relative to an embodied agent and identifying 3D shapes from coordinate lists. We also find that exposure to a visual modality does not provide a generalizable understanding of 3D space that VLMs are able to utilize for these tasks. Finally, we show that while the very latest frontier models can solve the provided tasks (though harder variants may still stump them), fully fine-tuning a small LM or LORA fine-tuning a small LLM show potential to match frontier model performance, suggesting an avenue for specialized embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。