arXiv:2509.08126cs.RO2025-09中稿 · 2025 IEEE-RAS 24th…

让机器人通过自然语言精准抓取物体,支持有重复物品的复杂场景。

Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning

  • 基于属性的语言理解与空间推理结合,定位目标物体并预测抓取姿态。
  • 在真实和仿真环境中均超越基线,弱监督下仍保持高抓取成功率。
  • 仅需单像素标注,适合实际部署,实时运行达17.59帧/秒。

让机器人根据自然语言指令抓取物体是人机交互的关键挑战。现有方法常受限于开放性语言表达,且假设目标唯一、无重复实例,同时依赖昂贵的密集像素级标注。本文提出属性驱动的目标定位与机器人抓取框架(OGRG),可解析开放语言指令,在存在重复对象的场景中进行空间推理,实现目标定位与平面抓取姿态预测。研究涵盖两种设置:(1)全监督的指代抓取合成(RGS),(2)仅需单像素抓取标注的弱监督指代抓取可用性(RGA)。核心贡献包括双向视觉-语言融合模块及深度信息融合,增强几何推理能力。实验表明,OGRG在桌面场景中优于所有基线模型,尤其在多样化空间语言指令下表现突出。在RGS设置下,单块NVIDIA RTX 2080 Ti GPU上达到17.59 FPS,支持闭环或多物体序列抓取;在弱监督的RGA设置中,仿真与真实机器人测试均实现更高抓取成功率,验证了其空间推理设计的有效性。

原文摘要 · Abstract (English)

Enabling robots to grasp objects specified through natural language is essential for effective human-robot interaction, yet it remains a significant challenge. Existing approaches often struggle with open-form language expressions and typically assume unambiguous target objects without duplicates. Moreover, they frequently rely on costly, dense pixel-wise annotations for both object grounding and grasp configuration. We present Attribute-based Object Grounding and Robotic Grasping (OGRG), a novel framework that interprets open-form language expressions and performs spatial reasoning to ground target objects and predict planar grasp poses, even in scenes containing duplicated object instances. We investigate OGRG in two settings: (1) Referring Grasp Synthesis (RGS) under pixel-wise full supervision, and (2) Referring Grasp Affordance (RGA) using weakly supervised learning with only single-pixel grasp annotations. Key contributions include a bi-directional vision-language fusion module and the integration of depth information to enhance geometric reasoning, improving both grounding and grasping performance. Experiment results show that OGRG outperforms strong baselines in tabletop scenes with diverse spatial language instructions. In RGS, it operates at 17.59 FPS on a single NVIDIA RTX 2080 Ti GPU, enabling potential use in closed-loop or multi-object sequential grasping, while delivering superior grounding and grasp prediction accuracy compared to all the baselines considered. Under the weakly supervised RGA setting, OGRG also surpasses baseline grasp-success rates in both simulation and real-robot trials, underscoring the effectiveness of its spatial reasoning design. Project page: https://z.umn.edu/ogrg

机器人抓取语言理解空间推理弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。