从第一视角交互视频中定位3D场景的可操作区域。
Grounding 3D Scene Affordance From Egocentric Interactions

- 用交互意图引导模型关注关键区域,提升定位精度。
- 通过双向查询解码器对齐多源特征,解决空间与语义对齐难题。
- 适用于具身智能体自主学习交互技能,推动机器人场景理解发展。
3D场景可操作性定位旨在识别3D环境中的可交互区域,这对具身智能体与周围环境进行智能交互至关重要。现有方法通常基于静态几何结构和视觉外观将语义映射到3D实例,这种被动策略限制了智能体主动感知与参与环境的能力,使其依赖预定义的语义指令。相比之下,人类通过观察并模仿他人与环境的互动来发展复杂交互技能。为赋予模型此类能力,我们提出新任务:从第一视角交互视频中接地3D场景可操作性,目标是根据一段第一视角交互视频,在3D场景中识别对应的可操作区域。该任务面临空间复杂性与多源对齐复杂性的挑战。为此,我们提出Ego-SAG框架,利用交互意图引导模型聚焦于相关子区域,并通过双向查询解码器机制对齐不同来源的可操作性特征。此外,我们构建了首个第一视角视频-3D场景可操作性数据集VSAD,涵盖多种常见交互类型与多样化3D环境以支持此任务。在VSAD上的大量实验验证了该任务的可行性及所提方法的有效性。
原文摘要 · Abstract (English)
Grounding 3D scene affordance aims to locate interactive regions in 3D environments, which is crucial for embodied agents to interact intelligently with their surroundings. Most existing approaches achieve this by mapping semantics to 3D instances based on static geometric structure and visual appearance. This passive strategy limits the agent's ability to actively perceive and engage with the environment, making it reliant on predefined semantic instructions. In contrast, humans develop complex interaction skills by observing and imitating how others interact with their surroundings. To empower the model with such abilities, we introduce a novel task: grounding 3D scene affordance from egocentric interactions, where the goal is to identify the corresponding affordance regions in a 3D scene based on an egocentric video of an interaction. This task faces the challenges of spatial complexity and alignment complexity across multiple sources. To address these challenges, we propose the Egocentric Interaction-driven 3D Scene Affordance Grounding (Ego-SAG) framework, which utilizes interaction intent to guide the model in focusing on interaction-relevant sub-regions and aligns affordance features from different sources through a bidirectional query decoder mechanism. Furthermore, we introduce the Egocentric Video-3D Scene Affordance Dataset (VSAD), covering a wide range of common interaction types and diverse 3D environments to support this task. Extensive experiments on VSAD validate both the feasibility of the proposed task and the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。