arXiv:2503.04154cs.CVcs.AI2025-03中稿 · IROS 2025被引 2

通过上下文感知知识提升弱监督单目3D检测精度

CA-W3D: Leveraging Context-Aware Knowledge for Weakly Supervised Monocular 3D Detection

  • 用区域对比学习对齐单目与2D视觉模型的特征
  • 伪标签训练中引入双到一蒸馏,保留空间信息
  • 适合追求高效弱监督3D检测的开发者

弱监督单目3D检测虽减少标注成本,但难以捕捉复杂场景所需的全局上下文。现有方法多聚焦目标中心特征,忽视关键语义关系。本文提出CA-W3D,采用两阶段训练:第一阶段通过区域级对象对比匹配(ROCM),对齐可训练单目3D编码器与冻结开放词汇2D视觉定位模型的区域特征,使编码器具备场景特定属性判别能力;第二阶段引入双到一蒸馏(D2OD)的伪标签训练机制,将上下文先验注入单目编码器,同时保持推理计算效率和空间保真度。在KITTI公开数据集上的实验表明,该方法在所有指标上均超越当前最先进方法,验证了上下文感知知识的重要性。

原文摘要 · Abstract (English)

Weakly supervised monocular 3D detection, while less annotation-intensive, often struggles to capture the global context required for reliable 3D reasoning. Conventional label-efficient methods focus on object-centric features, neglecting contextual semantic relationships that are critical in complex scenes. In this work, we propose a Context-Aware Weak Supervision for Monocular 3D object detection, namely CA-W3D, to address this limitation in a two-stage training paradigm. Specifically, we first introduce a pre-training stage employing Region-wise Object Contrastive Matching (ROCM), which aligns regional object embeddings derived from a trainable monocular 3D encoder and a frozen open-vocabulary 2D visual grounding model. This alignment encourages the monocular encoder to discriminate scene-specific attributes and acquire richer contextual knowledge. In the second stage, we incorporate a pseudo-label training process with a Dual-to-One Distillation (D2OD) mechanism, which effectively transfers contextual priors into the monocular encoder while preserving spatial fidelity and maintaining computational efficiency during inference. Extensive experiments conducted on the public KITTI benchmark demonstrate the effectiveness of our approach, surpassing the SoTA method over all metrics, highlighting the importance of contextual-aware knowledge in weakly-supervised monocular 3D detection.

3D检测弱监督单目视觉上下文建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。