用视觉语言模型零样本预测物体任务相关部位,提升机器人抓取理解能力。
GauTOAO: Gaussian-based Task-Oriented Affordance of Objects
- 基于高斯表示,结合视觉语言模型实现零样本任务导向区域识别。
- 在真实场景中验证,对多种任务泛化能力强,定位更准确。
- 适合需理解物体功能的智能机器人应用,如家庭服务或工业操作。
当机器人使用灵巧手或夹爪抓取物体时,应理解该物体的任务导向属性(TOAO),因为不同任务通常关注物体的不同部分。为解决此问题,我们提出 GauTOAO,一种基于高斯的物体任务导向属性框架,通过零样本方式利用视觉语言模型,根据自然语言查询预测物体的属性相关区域。本方法引入新范式:‘静态相机,移动物体’,使机器人在操作过程中能更好观察和理解手中物体。GauTOAO克服了现有方法在空间分组上的不足,通过 DINO 特征提取全面的 3D 物体掩码,并以此条件化高斯点,生成针对特定任务的精细语义分布。该方法显著提升 TOAO 提取精度,增强机器人对物体的理解并改善任务表现。我们在真实世界实验中验证了 GauTOAO 的有效性,证明其具备跨多种任务的泛化能力。
原文摘要 · Abstract (English)
When your robot grasps an object using dexterous hands or grippers, it should understand the Task-Oriented Affordances of the Object(TOAO), as different tasks often require attention to specific parts of the object. To address this challenge, we propose GauTOAO, a Gaussian-based framework for Task-Oriented Affordance of Objects, which leverages vision-language models in a zero-shot manner to predict affordance-relevant regions of an object, given a natural language query. Our approach introduces a new paradigm: "static camera, moving object," allowing the robot to better observe and understand the object in hand during manipulation. GauTOAO addresses the limitations of existing methods, which often lack effective spatial grouping, by extracting a comprehensive 3D object mask using DINO features. This mask is then used to conditionally query gaussians, producing a refined semantic distribution over the object for the specified task. This approach results in more accurate TOAO extraction, enhancing the robot's understanding of the object and improving task performance. We validate the effectiveness of GauTOAO through real-world experiments, demonstrating its capability to generalize across various tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。