arXiv:2412.11412cs.CV2024-12被引 5

用2D图像生成3D数据,提升单目室内检测泛化能力

V-MIND: Building Versatile Monocular Indoor 3D Detector with Diverse 2D Annotations

  • 将大规模2D图像转为3D点云,生成伪3D框用于训练
  • 引入自校准损失降低距离误差,提升检测精度
  • 适合缺乏3D标注的场景,尤其对新类别检测有帮助

单目室内3D目标检测因VR/AR和机器人应用需求而备受关注,但受限于3D数据稀缺且标注成本高。本文提出V-MIND,通过利用公开的大规模2D数据集,结合单目深度估计和相机内参预测技术,将2D图像转化为3D点云并生成伪3D边界框。为缓解转换带来的距离误差,提出3D自校准损失以优化伪边界框;同时设计模糊性损失应对新类别引入时的歧义问题。通过与现有3D数据集联合训练,V-MIND在Omni3D数据集上实现跨多类别的最先进性能。

原文摘要 · Abstract (English)

The field of indoor monocular 3D object detection is gaining significant attention, fueled by the increasing demand in VR/AR and robotic applications. However, its advancement is impeded by the limited availability and diversity of 3D training data, owing to the labor-intensive nature of 3D data collection and annotation processes. In this paper, we present V-MIND (Versatile Monocular INdoor Detector), which enhances the performance of indoor 3D detectors across a diverse set of object classes by harnessing publicly available large-scale 2D datasets. By leveraging well-established monocular depth estimation techniques and camera intrinsic predictors, we can generate 3D training data by converting large-scale 2D images into 3D point clouds and subsequently deriving pseudo 3D bounding boxes. To mitigate distance errors inherent in the converted point clouds, we introduce a novel 3D self-calibration loss for refining the pseudo 3D bounding boxes during training. Additionally, we propose a novel ambiguity loss to address the ambiguity that arises when introducing new classes from 2D datasets. Finally, through joint training with existing 3D datasets and pseudo 3D bounding boxes derived from 2D datasets, V-MIND achieves state-of-the-art object detection performance across a wide range of classes on the Omni3D indoor dataset.

单目3D检测2D转3D数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。