arXiv:2503.18769cs.CLcs.RO2025-03

用结构化符号让大模型精准操控3D空间物体

AlphaSpace: Enabling Robotic Actions through Semantic Tokenization and Symbolic Reasoning

  • 将物体属性与坐标编码为结构化符号,实现细粒度空间推理
  • 在3D任务中达66.67%准确率,远超GPT-4o和Claude 3.5 Sonnet
  • 适合需要精确空间操作的机器人控制场景

本文提出AlphaSpace,一种增强语言模型在三维笛卡尔空间中进行机器人操作时空间推理能力的新方法。该方法采用分层语义标记策略,在粗粒度和细粒度层面编码空间信息。通过结构化标记表示物体的属性、位置及高度信息,使大模型无需依赖传统视觉嵌入即可实现精确的空间推理,从而准确地将物体定位在特定的 (x, y, z) 坐标上。实验结果表明,AlphaSpace在操作任务中达到66.67%的总准确率,显著优于GPT-4o的37.5%和Claude 3.5 Sonnet的29.17%。这些结果证明了结构化空间编码在操纵任务中的潜力,值得进一步探索。

原文摘要 · Abstract (English)

This paper presents AlphaSpace, a novel methodology designed to enhance the spatial reasoning capabilities of language models for robotic manipulation in 3D Cartesian space. AlphaSpace employs a hierarchical semantics-based tokenization strategy that encodes spatial information at both coarse and fine-grained levels. Our approach represents objects with their attributes, positions, and height information through structured tokens, enabling precise spatial reasoning without relying on traditional vision-based embeddings. This approach enables LLMs to accurately manipulate objects by positioning them at specific (x, y, z) coordinates. Experimental results suggest that AlphaSpace demonstrates promising potential for improving manipulation tasks, achieving a total accuracy of 66.67%, compared to 37.5% for GPT-4o and 29.17% for Claude 3.5 Sonnet. These results demonstrate the potential of structured spatial encoding for manipulation tasks and warrant further exploration.

机器人操控空间推理符号系统大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。