通过语言引导的场景图提升3D视觉定位精度
LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
- 构建语言引导的场景图,增强对相似物体的区分能力
- 利用2D预训练语义辅助3D多模态编码,提升特征质量
- 适合处理复杂场景中多个相似目标的定位问题
3D视觉定位旨在从3D场景中精确定位自然语言描述的目标。由于3D与语言模态间存在显著差异,仅靠空间关系难以区分多个相似物体。现有方法采用以目标为中心的学习机制,忽略对被指代物体的建模。本文提出一种新型3D视觉定位框架,通过构建语言引导的场景图并增强被指代物体的辨识能力,提升关系感知效果。该框架采用双分支视觉编码器,利用预训练的2D语义信息来增强和监督多模态3D编码。同时,引入图注意力机制,促进跨模态交互中的关系导向信息融合。学习到的物体表示与场景图结构实现了3D视觉内容与文本描述的有效对齐。在主流基准上的实验结果表明,本方法优于当前最优模型,尤其在应对多个相似干扰物时表现更优。
原文摘要 · Abstract (English)
3D visual grounding aims to localize the unique target described by natural languages in 3D scenes. The significant gap between 3D and language modalities makes it a notable challenge to distinguish multiple similar objects through the described spatial relationships. Current methods attempt to achieve cross-modal understanding in complex scenes via a target-centered learning mechanism, ignoring the modeling of referred objects. We propose a novel 3D visual grounding framework that constructs language-guided scene graphs with referred object discrimination to improve relational perception. The framework incorporates a dual-branch visual encoder that leverages pre-trained 2D semantics to enhance and supervise the multi-modal 3D encoding. Furthermore, we employ graph attention to promote relationship-oriented information fusion in cross-modal interaction. The learned object representations and scene graph structure enable effective alignment between 3D visual content and textual descriptions. Experimental results on popular benchmarks demonstrate our superior performance compared to state-of-the-art methods, especially in handling the challenges of multiple similar distractors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。