用伪标签实现单目3D检测的跨域泛化,提升弱监督学习效率。
GATE3D: Generalized Attention-based Task-synergized Estimation in 3D*
- 基于2D与3D预测一致性损失,构建弱监督训练框架。
- 在KITTI和自建室内数据集上均达竞争力表现。
- 适合机器人、AR/VR等需跨场景感知的应用场景。
计算机视觉正趋向于构建可同时处理多种任务的通用模型,通常需在多领域数据集上联合训练以实现良好泛化。然而,单目3D目标检测在多领域训练中面临独特挑战:准确的3D标注数据稀缺,尤其在典型道路自动驾驶场景之外。为此,本文提出GATE3D——一种针对单目3D目标检测的弱监督新框架,利用伪标签缓解标注不足问题。现有预训练模型在非道路环境(如室内)对行人检测效果不佳,源于数据分布偏差。不同于通用2D检测模型,单目3D检测的泛化能力研究仍较匮乏。GATE3D通过2D与3D预测间的约束一致性损失有效弥合领域差异。实验表明,该模型在KITTI基准上表现优异,并在自建室内办公室数据集上验证了良好的泛化能力。结果证明,GATE3D可通过高效的预训练策略,显著加速从有限标注数据中学习,展现出在机器人、增强现实与虚拟现实中的广泛应用潜力。
原文摘要 · Abstract (English)
The emerging trend in computer vision emphasizes developing universal models capable of simultaneously addressing multiple diverse tasks. Such universality typically requires joint training across multi-domain datasets to ensure effective generalization. However, monocular 3D object detection presents unique challenges in multi-domain training due to the scarcity of datasets annotated with accurate 3D ground-truth labels, especially beyond typical road-based autonomous driving contexts. To address this challenge, we introduce a novel weakly supervised framework leveraging pseudo-labels. Current pretrained models often struggle to accurately detect pedestrians in non-road environments due to inherent dataset biases. Unlike generalized image-based 2D object detection models, achieving similar generalization in monocular 3D detection remains largely unexplored. In this paper, we propose GATE3D, a novel framework designed specifically for generalized monocular 3D object detection via weak supervision. GATE3D effectively bridges domain gaps by employing consistency losses between 2D and 3D predictions. Remarkably, our model achieves competitive performance on the KITTI benchmark as well as on an indoor-office dataset collected by us to evaluate the generalization capabilities of our framework. Our results demonstrate that GATE3D significantly accelerates learning from limited annotated data through effective pre-training strategies, highlighting substantial potential for broader impacts in robotics, augmented reality, and virtual reality applications. Project page: https://ies0411.github.io/GATE3D/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。