arXiv:2411.19626cs.CVcs.AI2024-11CVPR被引 28

让机器人理解3D物体能做什么,基于几何与意图协同推理

GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding

  • 通过几何与交互意图的协同推理,挖掘物体不变属性
  • 在新数据集PIADv2上实现领先性能,显著提升泛化能力
  • 适合需要理解复杂交互场景的机器人视觉研究者

开放词汇3D物体功能定位旨在根据任意指令预测3D物体上的可操作区域,对机器人通用感知真实场景和应对操作变化至关重要。现有方法依赖图像或语言描述中的交互信息来引入外部交互先验,但受限于语义空间,未能充分利用物体的隐含不变几何特征及潜在交互意图。人类常通过多步推理和联想类比解决复杂任务。为此,我们提出GREAT(Geometry-Intention Collaborative Inference)框架,通过挖掘物体不变几何属性并进行潜在交互场景的类比推理,构建功能知识,并融合几何与视觉信息实现3D物体功能定位。此外,我们构建了当前最大的3D物体功能数据集Point Image Affordance Dataset v2(PIADv2),以支持该任务。大量实验验证了GREAT的有效性与优越性。代码与数据集已公开。

原文摘要 · Abstract (English)

Open-Vocabulary 3D object affordance grounding aims to anticipate ``action possibilities'' regions on 3D objects with arbitrary instructions, which is crucial for robots to generically perceive real scenarios and respond to operational changes. Existing methods focus on combining images or languages that depict interactions with 3D geometries to introduce external interaction priors. However, they are still vulnerable to a limited semantic space by failing to leverage implied invariant geometries and potential interaction intentions. Normally, humans address complex tasks through multi-step reasoning and respond to diverse situations by leveraging associative and analogical thinking. In light of this, we propose GREAT (GeometRy-intEntion collAboraTive inference) for Open-Vocabulary 3D Object Affordance Grounding, a novel framework that mines the object invariant geometry attributes and performs analogically reason in potential interaction scenarios to form affordance knowledge, fully combining the knowledge with both geometries and visual contents to ground 3D object affordance. Besides, we introduce the Point Image Affordance Dataset v2 (PIADv2), the largest 3D object affordance dataset at present to support the task. Extensive experiments demonstrate the effectiveness and superiority of GREAT. The code and dataset are available at https://yawen-shao.github.io/GREAT/.

3D理解功能定位机器人视觉几何推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。