arXiv:2603.27970cs.CV2026-03中稿 · CVPR被引 2

通过视觉线索实现3D场景中交互区域的精准定位。

AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers

  • 利用图像与点云间的语义对应,匹配关键点以定位可交互区域。
  • 在685个高分辨率场景中构建了29万+的交互标注数据集。
  • 适合做3D场景理解、机器人交互与具身智能的研究者。

在诸多应用中,可操作性学习是一项复杂挑战。现有方法主要依赖物体的几何结构、视觉知识和可操作性标签来确定可交互区域,但将其扩展到场景级别则更为困难,因物体与场景层面的语义融合并不直接。本文提出AffordBridge,一个大规模数据集,包含685个高分辨率室内场景的点云数据,共291,637条功能交互标注。标注数据配套有与场景中同一实例对应的RGB图像。基于该数据集,我们提出AffordMatcher方法,通过建立基于图像与点云实例间的连贯语义对应关系,实现关键点匹配,从而依据所谓的视觉线索(visual signifiers)更精确地识别可操作区域。在本数据集上的实验结果证明了该方法的有效性。

原文摘要 · Abstract (English)

Affordance learning is a complex challenge in many applications, where existing approaches primarily focus on the geometric structures, visual knowledge, and affordance labels of objects to determine interactable regions. However, extending this learning capability to a scene is significantly more complicated, as incorporating object- and scene-level semantics is not straightforward. In this work, we introduce AffordBridge, a large-scale dataset with 291,637 functional interaction annotations across 685 high-resolution indoor scenes in the form of point clouds. Our affordance annotations are complemented by RGB images that are linked to the same instances within the scenes. Building upon our dataset, we propose AffordMatcher, an affordance learning method that establishes coherent semantic correspondences between image-based and point cloud-based instances for keypoint matching, enabling a more precise identification of affordance regions based on cues, so-called visual signifiers. Experimental results on our dataset demonstrate the effectiveness of our approach compared to other methods.

3D场景理解可操作性学习视觉线索点云匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。