arXiv:2606.30632cs.ROcs.AI2026-06

让机器人用盘子切蛋糕,靠语言理解+3D定位实现创意工具使用。

GROW$^2$: Grounding Which and Where for Robot Tool Use

论文配图:GROW$^2$: Grounding Which and Where for Robot Tool Use
图 1 · 摘自论文原文
  • 用视觉语言模型解析指令,选出可用物体并定位关键部位。
  • 在单张RGB-D图像上精准定位工具与目标的3D作用区域。
  • 零样本泛化到新物品,在真实和模拟环境中表现更优。

机器人能否在无刀情况下用盘子切蛋糕?工具使用极大拓展了机器人的能力,但要超越工具的原始功能进行创造性使用,机器人面临开放世界用途定位挑战:从开放类别中选择一个物体作为工具,并精确定位其作用区域。为此,我们提出GROW²(GROunding Which and Where),利用物体部件作为自然抽象,将定位过程分层为语义与几何两个层次,避免依赖数据密集型端到端训练。语义层面,GROW²借助视觉语言模型的常识推理能力,解析自然语言任务指令,选择合适工具,并识别工具与目标对象中与任务相关的部分。几何层面,视觉基础模型从单张RGB-D图像中将选定部分精确定位为3D区域。在基准测试中,GROW²在用途预测任务上优于现有最优方法。进一步实验表明,其在开放类别物体上实现了零样本泛化,并在仿真与真实机器人工具使用实验中均显著超越基线方法。

原文摘要 · Abstract (English)

Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to use tools creatively beyond their intended functions, the robot faces the challenge of $\textit{open-world affordance grounding}$: select an open-category object to act as a tool and localize its specific region of action. To this end, we introduce GROW$^2$ (GROunding Which and Where), which leverages object parts as a natural abstraction to split the grounding process hierarchically into semantic and geometric levels, thus bypassing the need for data-heavy, end-to-end training. Semantically, GROW$^2$ harnesses the commonsense reasoning of Vision-Language Models (VLMs) to parse a natural-language task instruction, select a suitable object as the tool, and identify task-relevant parts on the tool and the target object. Geometrically, vision foundation models then ground the selected parts into precise 3D regions from a single RGB-D image. Experiments on established benchmarks show that GROW$^2$ outperforms state-of-the-art baselines on affordance prediction benchmarks. Further, it achieves zero-shot generalization over open-category objects and outperforms baselines in both simulated and real-world robot tool use experiments.

机器人工具使用视觉语言模型3D定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。