让机器人根据语言指令和观察,精准找到可操作的3D物体位置。
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
- 用视觉、语言和交互信息融合建模3D物体可操作区域。
- 在部分视角下仍能准确预测物体可操作位置,效果优于现有方法。
- 适合研究具身智能、人机交互与多模态感知的开发者。
3D物体可操作性定位是将感知与动作关联的关键任务,对具身智能系统如机器人至关重要。本文提出一种新任务:基于语言指令、视觉观察与交互信息,实现3D物体可操作性的精准定位,受认知科学启发。为此,我们构建了包含点云、图像与语言指令的AGPIL数据集,涵盖全视角、部分视角与旋转视角下的可操作性标注。由于实际场景中存在视角限制、物体旋转或遮挡,仅能获取物体部分观测信息。为此,模型需在不完整信息下完成推理。本文提出LMAffordance3D,首个多模态、语言引导的3D可操作性定位网络,利用视觉-语言模型融合2D与3D空间特征及语义信息。在AGPIL数据集上的实验表明,该方法在多种未见设置下均表现优异,具备强泛化能力。项目主页见 https://sites.google.com/view/lmaffordance3d。
原文摘要 · Abstract (English)
Grounding 3D object affordance is a task that locates objects in 3D space where they can be manipulated, which links perception and action for embodied intelligence. For example, for an intelligent robot, it is necessary to accurately ground the affordance of an object and grasp it according to human instructions. In this paper, we introduce a novel task that grounds 3D object affordance based on language instructions, visual observations and interactions, which is inspired by cognitive science. We collect an Affordance Grounding dataset with Points, Images and Language instructions (AGPIL) to support the proposed task. In the 3D physical world, due to observation orientation, object rotation, or spatial occlusion, we can only get a partial observation of the object. So this dataset includes affordance estimations of objects from full-view, partial-view, and rotation-view perspectives. To accomplish this task, we propose LMAffordance3D, the first multi-modal, language-guided 3D affordance grounding network, which applies a vision-language model to fuse 2D and 3D spatial features with semantic features. Comprehensive experiments on AGPIL demonstrate the effectiveness and superiority of our method on this task, even in unseen experimental settings. Our project is available at https://sites.google.com/view/lmaffordance3d.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。