arXiv:2503.00778cs.RO2025-03被引 41

让机器人通过理解用户指令,自动找出物品的可抓握位置并完成任务。

AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter

  • 利用视觉语言模型从用户指令中推理任务和抓取点
  • 在杂乱场景中实现对新物品的任务导向抓取,成功率超90%
  • 适合日常环境下的人机协作,无需预先训练特定物体

理解物体的抓取属性并以任务为导向进行抓取,是机器人成功完成操作任务的关键。抓取属性指根据物体功能确定其可抓位置与方式,是高效任务导向抓取的基础。然而,现有方法通常依赖于特定任务和物体的大量训练数据,难以泛化到新物体和复杂场景。本文提出AffordGrasp,一种新型开放词汇抓取框架,利用视觉语言模型(VLMs)进行上下文感知的抓取属性推理。不同于依赖显式任务和物体定义的方法,本方法从隐式用户指令中直接推断任务,实现更自然流畅的人机交互。基于推理结果,框架识别任务相关物体,并通过视觉定位模块将其部分级抓取属性锚定。由此生成精确位于物体抓取区域内的任务导向抓取姿态,确保操作的功能性与情境适应性。大量实验表明,AffordGrasp在仿真与真实场景中均达到领先性能,验证了方法的有效性。我们相信该方法推动了机器人操作技术发展,为具身智能领域作出贡献。

原文摘要 · Abstract (English)

Inferring the affordance of an object and grasping it in a task-oriented manner is crucial for robots to successfully complete manipulation tasks. Affordance indicates where and how to grasp an object by taking its functionality into account, serving as the foundation for effective task-oriented grasping. However, current task-oriented methods often depend on extensive training data that is confined to specific tasks and objects, making it difficult to generalize to novel objects and complex scenes. In this paper, we introduce AffordGrasp, a novel open-vocabulary grasping framework that leverages the reasoning capabilities of vision-language models (VLMs) for in-context affordance reasoning. Unlike existing methods that rely on explicit task and object specifications, our approach infers tasks directly from implicit user instructions, enabling more intuitive and seamless human-robot interaction in everyday scenarios. Building on the reasoning outcomes, our framework identifies task-relevant objects and grounds their part-level affordances using a visual grounding module. This allows us to generate task-oriented grasp poses precisely within the affordance regions of the object, ensuring both functional and context-aware robotic manipulation. Extensive experiments demonstrate that AffordGrasp achieves state-of-the-art performance in both simulation and real-world scenarios, highlighting the effectiveness of our method. We believe our approach advances robotic manipulation techniques and contributes to the broader field of embodied AI. Project website: https://eqcy.github.io/affordgrasp/.

机器人抓取视觉语言模型开放词汇具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。