用视频教AI识别物体可操作区域,提升机器人抓取能力
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
- 用多模态大模型结合动作视频,实现动态交互下的精准操作区域定位
- 在38K视频数据上表现超越现有方法,对未见物体也有良好泛化能力
- 适合研究机器人视觉、人机交互和三维场景理解的学者与开发者
3D可操作性定位旨在识别物体上可被操作的区域,对机器人抓取任务至关重要。以往研究多依赖静态线索(如文本和图像),难以捕捉动态交互中的时序与因果信息。为此,我们构建了首个基于视频的3D可操作性数据集VIDA,包含38,000段人-物交互视频,覆盖16类可操作性、38种物体类别及22,000个点云。基于VIDA,我们提出VideoAfford模型,通过激活多模态大语言模型并引入可操作性分割能力,在统一框架中实现世界知识推理与细粒度可操作性定位。为增强动作理解,我们采用隐式动作编码器从人-物交互视频中提取动态交互先验,并设计空间感知损失函数,使模型获得完整的3D空间知识。大量实验表明,该模型显著优于现有方法,在开放世界下具备出色的可操作性推理能力。所有数据集与代码将公开发布,以推动该领域发展。
原文摘要 · Abstract (English)
3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for robotic manipulation. Previous research primarily focused on learning affordance knowledge from static cues such as language and images, which struggle to provide sufficient dynamic interaction context that can reveal temporal and causal cues. To alleviate this predicament, we collect a comprehensive video-based 3D affordance dataset, \textit{VIDA}, which contains 38K human-object-interaction videos covering 16 affordance types, 38 object categories, and 22K point clouds. Based on \textit{VIDA}, we propose a strong baseline: VideoAfford, which activates multimodal large language models with additional affordance segmentation capabilities, enabling both world knowledge reasoning and fine-grained affordance grounding within a unified framework. To enhance action understanding capability, we leverage a latent action encoder to extract dynamic interaction priors from HOI videos. Moreover, we introduce a \textit{spatial-aware} loss function to enable VideoAfford to obtain comprehensive 3D spatial knowledge. Extensive experimental evaluations demonstrate that our model significantly outperforms well-established methods and exhibits strong open-world generalization with affordance reasoning abilities. All datasets and code will be publicly released to advance research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。