用视频动态动作指导3D物体功能区域定位,更准确。
VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos
- 通过视频中的交互动作提供功能监督,替代静态图像
- 在新数据集PVAD上显著优于传统静态方法
- 适合研究具身智能与人机交互的学者
3D物体功能区域定位旨在识别支持人-物交互(HOI)的3D物体区域,对具身视觉推理至关重要。然而,现有方法多依赖静态视觉或文本线索,忽视了功能本质上由动态行为定义的事实,导致难以准确定位真实交互中的接触区域。我们提出视频引导的3D功能区域定位,利用动态交互序列提供功能监督。为此,我们设计VAGNet框架,将视频中提取的交互线索与3D结构对齐,解决静态线索无法处理的模糊性问题。为支持该设定,我们构建了首个HOI视频-3D配对的数据集PVAD,提供先前工作缺乏的功能监督。在PVAD上的大量实验表明,VAGNet达到当前最优性能,显著超越基于静态线索的基线模型。代码与数据集将公开。
原文摘要 · Abstract (English)
3D object affordance grounding aims to identify regions on 3D objects that support human-object interaction (HOI), a capability essential to embodied visual reasoning. However, most existing approaches rely on static visual or textual cues, neglecting that affordances are inherently defined by dynamic actions. As a result, they often struggle to localize the true contact regions involved in real interactions. We take a different perspective. Humans learn how to use objects by observing and imitating actions, not just by examining shapes. Motivated by this intuition, we introduce video-guided 3D affordance grounding, which leverages dynamic interaction sequences to provide functional supervision. To achieve this, we propose VAGNet, a framework that aligns video-derived interaction cues with 3D structure to resolve ambiguities that static cues cannot address. To support this new setting, we introduce PVAD, the first HOI video-3D pairing affordance dataset, providing functional supervision unavailable in prior works. Extensive experiments on PVAD show that VAGNet achieves state-of-the-art performance, significantly outperforming static-based baselines. The code and dataset will be open publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。