用坐标对齐+思维链提升视觉语言模型的空间推理能力
SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
- 通过双向坐标对齐让视觉与语言信息匹配空间位置
- 在仿真和真实场景中导航与操作任务上超越现有最佳方法
- 适合需要精准空间规划的机器人任务研究者
空间推理是具身智能研究中的核心问题。现有通过附加空间数据或微调增强空间推理的方法,在处理复杂具身任务时效果有限,主要因其依赖语言输出。尽管部分方法引入基于点的动作空间以缓解该问题,但在复杂环境中的复杂任务仍表现不足,根源在于未能充分发挥视觉语言模型(VLMs)固有的思维与推理能力。为此,我们提出SpatialCoT,一种专为增强VLM空间推理能力设计的新方法,包含两个阶段:空间坐标双向对齐,将视觉-语言输入与空间坐标对齐;以及思维链空间定位,利用语言模型的推理能力实现高级空间推理。我们在具有挑战性的导航与操作任务上评估了SpatialCoT,涵盖仿真与真实场景。实验结果表明,该方法在两类任务上均显著优于先前最先进方法。
原文摘要 · Abstract (English)
Spatial reasoning is an essential problem in embodied AI research. Efforts to enhance spatial reasoning abilities through supplementary spatial data and fine-tuning have proven limited and ineffective when addressing complex embodied tasks, largely due to their dependence on language-based outputs. While some approaches have introduced a point-based action space to mitigate this issue, they fall short in managing more intricate tasks within complex environments. This deficiency arises from their failure to fully exploit the inherent thinking and reasoning capabilities that are fundamental strengths of Vision-Language Models (VLMs). To address these limitations, we propose a novel approach named SpatialCoT, specifically designed to bolster the spatial reasoning capabilities of VLMs. Our approach comprises two stages: spatial coordinate bi-directional alignment, which aligns vision-language inputs with spatial coordinates, and chain-of-thought spatial grounding, which harnesses the reasoning capabilities of language models for advanced spatial reasoning. We evaluate SpatialCoT on challenging navigation and manipulation tasks, both in simulation and real-world settings. Experimental results demonstrate that our method significantly outperforms previous state-of-the-art approaches in both tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。