arXiv:2505.01809cs.CV2025-05ICRA被引 3

通过类别与实例对齐,实现点云中文定位的弱监督精准识别

3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment

  • 分双分支设计:类别分支利用预训练检测器增强类别感知
  • 实例分支结合空间关系描述,区分同类别多个目标
  • 在Nr3D、Sr3D、ScanRef上均达顶尖性能,适合点云语义理解场景

3D弱监督视觉定位任务旨在基于自然语言描述,在无标注引导的情况下定位点云中的定向3D框。该设定面临两大挑战:类别级模糊性与实例级复杂性。类别级模糊性源于细粒度类别在稀疏点云中表示困难,导致类别区分困难;实例级复杂性则因同一类别多个实例共存,造成定位干扰。为此,我们提出一种新型弱监督定位方法,显式区分类别与实例。类别分支利用预训练外部检测器的丰富类别知识,将物体提议特征与句子级类别特征对齐,提升类别感知能力;实例分支则利用语言查询中的空间关系描述,精炼物体提议特征,确保同类目标间的清晰区分。该设计使模型能准确识别目标类别对象,并区分同类别内的多个实例。相比以往方法,本方案在三个主流基准(Nr3D、Sr3D、ScanRef)上均达到当前最优性能。

原文摘要 · Abstract (English)

The 3D weakly-supervised visual grounding task aims to localize oriented 3D boxes in point clouds based on natural language descriptions without requiring annotations to guide model learning. This setting presents two primary challenges: category-level ambiguity and instance-level complexity. Category-level ambiguity arises from representing objects of fine-grained categories in a highly sparse point cloud format, making category distinction challenging. Instance-level complexity stems from multiple instances of the same category coexisting in a scene, leading to distractions during grounding. To address these challenges, we propose a novel weakly-supervised grounding approach that explicitly differentiates between categories and instances. In the category-level branch, we utilize extensive category knowledge from a pre-trained external detector to align object proposal features with sentence-level category features, thereby enhancing category awareness. In the instance-level branch, we utilize spatial relationship descriptions from language queries to refine object proposal features, ensuring clear differentiation among objects. These designs enable our model to accurately identify target-category objects while distinguishing instances within the same category. Compared to previous methods, our approach achieves state-of-the-art performance on three widely used benchmarks: Nr3D, Sr3D, and ScanRef.

3D定位弱监督点云理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。