arXiv:2511.06908cs.CVcs.MM2025-11被引 1

提升单目3D视觉定位精度,强化空间描述理解与跨维度特征解耦。

Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding

  • 通过动态掩码高置信度关键词,保留隐含空间描述以增强空间理解。
  • 解耦2D/3D文本特征,避免跨维度干扰,提升视觉特征精炼效果。
  • 在Mono3DRefer数据集上取得最佳表现,远距离定位准确率提升13.54%。

单目3D视觉定位(Mono3DVG)是一项新兴任务,旨在通过带几何线索的文本描述在RGB图像中定位3D物体。现有方法存在两大局限:其一,过度依赖明确标识目标的高置信度关键词,忽视关键的空间描述;其二,通用文本特征同时包含2D与3D信息,相较于单一2D或3D视觉特征更具细节,导致在文本引导下优化视觉特征时产生跨维度干扰。为此,本文提出Mono3DVG-EnSD框架,包含两个核心组件:基于CLIP的词汇置信度适配器(CLIP-LCA)和维度解耦模块(D2M)。CLIP-LCA动态掩码高置信度关键词,保留低置信度的隐含空间描述,促使模型深入理解描述中的空间关系以实现精准定位。D2M将通用文本特征中的2D/3D分量解耦,分别引导同维度的视觉特征,从而确保跨模态交互的维度一致性,减轻干扰。在Mono3DRefer数据集上的全面对比与消融实验表明,本方法在所有指标上均达当前最优(SOTA),尤其在挑战性的远距离(Far, [email protected])场景中提升显著,达+13.54%。

原文摘要 · Abstract (English)

Monocular 3D Visual Grounding (Mono3DVG) is an emerging task that locates 3D objects in RGB images using text descriptions with geometric cues. However, existing methods face two key limitations. Firstly, they often over-rely on high-certainty keywords that explicitly identify the target object while neglecting critical spatial descriptions. Secondly, generalized textual features contain both 2D and 3D descriptive information, thereby capturing an additional dimension of details compared to singular 2D or 3D visual features. This characteristic leads to cross-dimensional interference when refining visual features under text guidance. To overcome these challenges, we propose Mono3DVG-EnSD, a novel framework that integrates two key components: the CLIP-Guided Lexical Certainty Adapter (CLIP-LCA) and the Dimension-Decoupled Module (D2M). The CLIP-LCA dynamically masks high-certainty keywords while retaining low-certainty implicit spatial descriptions, thereby forcing the model to develop a deeper understanding of spatial relationships in captions for object localization. Meanwhile, the D2M decouples dimension-specific (2D/3D) textual features from generalized textual features to guide corresponding visual features at same dimension, which mitigates cross-dimensional interference by ensuring dimensionally-consistent cross-modal interactions. Through comprehensive comparisons and ablation studies on the Mono3DRefer dataset, our method achieves state-of-the-art (SOTA) performance across all metrics. Notably, it improves the challenging Far([email protected]) scenario by a significant +13.54%.

3D定位视觉接地文本编码多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。