用自然语言定义物体朝向,让机器人更精准地操作物品。
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
- 用语言描述物体朝向(如‘插口方向’),摆脱固定坐标系限制。
- 零样本下在Open6DOR任务中成功率达48.7%,SIMPLER-Env达74.9%。
- 适用于需要精细操控的机器人任务,尤其适合无标注场景。
尽管空间推理在物体定位关系上取得进展,但常忽略物体朝向——这在6-DoF精细化操作中至关重要。传统姿态表示依赖预设坐标系或模板,限制了泛化性和语义关联。本文提出语义朝向概念,以自然语言在无参考系条件下定义物体朝向(如USB的‘插入方向’或杯子的‘把手方向’)。为此,我们构建了大规模数据集OrienText300K,包含30万条3D物体的语义朝向标注,并开发了PointSO模型,实现零样本语义朝向预测。将语义朝向融入视觉-语言模型(VLM)代理后,提出的SoFar框架支持6-DoF空间推理并生成机器人动作。大量实验表明,该方法有效且具备强泛化能力:在Open6DOR任务中零样本成功率48.7%,在SIMPLER-Env中达74.9%。
原文摘要 · Abstract (English)
While spatial reasoning has made progress in object localization relationships, it often overlooks object orientation-a key factor in 6-DoF fine-grained manipulation. Traditional pose representations rely on pre-defined frames or templates, limiting generalization and semantic grounding. In this paper, we introduce the concept of semantic orientation, which defines object orientations using natural language in a reference-frame-free manner (e.g., the "plug-in" direction of a USB or the "handle" direction of a cup). To support this, we construct OrienText300K, a large-scale dataset of 3D objects annotated with semantic orientations, and develop PointSO, a general model for zero-shot semantic orientation prediction. By integrating semantic orientation into VLM agents, our SoFar framework enables 6-DoF spatial reasoning and generates robotic actions. Extensive experiments demonstrated the effectiveness and generalization of our SoFar, e.g., zero-shot 48.7% successful rate on Open6DOR and zero-shot 74.9% successful rate on SIMPLER-Env.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。