根据检测难度动态调整标签扰动,提升单目3D目标检测精度
Difficulty-Aware Label-Guided Denoising for Monocular 3D Object Detection
- 按检测难易程度自适应调节标签扰动强度
- 在KITTI上全难度等级均达当前最优性能
- 适合需要高鲁棒性的自动驾驶场景
单目3D目标检测在自动驾驶与机器人等领域具有成本优势,但因深度信息固有模糊性而本质困难。现有基于DETR的方法虽通过全局注意力和辅助深度预测缓解问题,仍面临深度估计不准的挑战。此外,这些方法常忽略实例级检测难度(如遮挡、远距离、截断),导致性能受限。本文提出MonoDLGD框架,依据检测不确定性自适应地扰动并重构真值标签:对较易实例施加更强扰动,对较难实例施加更弱扰动,再进行重建以提供显式几何监督。通过联合优化标签重建与3D检测任务,促进几何感知表征学习,增强对不同复杂度物体的鲁棒性。在KITTI基准上的大量实验表明,MonoDLGD在所有难度等级下均达到当前最优表现。
原文摘要 · Abstract (English)
Monocular 3D object detection is a cost-effective solution for applications like autonomous driving and robotics, but remains fundamentally ill-posed due to inherently ambiguous depth cues. Recent DETR-based methods attempt to mitigate this through global attention and auxiliary depth prediction, yet they still struggle with inaccurate depth estimates. Moreover, these methods often overlook instance-level detection difficulty, such as occlusion, distance, and truncation, leading to suboptimal detection performance. We propose MonoDLGD, a novel Difficulty-Aware Label-Guided Denoising framework that adaptively perturbs and reconstructs ground-truth labels based on detection uncertainty. Specifically, MonoDLGD applies stronger perturbations to easier instances and weaker ones into harder cases, and then reconstructs them to effectively provide explicit geometric supervision. By jointly optimizing label reconstruction and 3D object detection, MonoDLGD encourages geometry-aware representation learning and improves robustness to varying levels of object complexity. Extensive experiments on the KITTI benchmark demonstrate that MonoDLGD achieves state-of-the-art performance across all difficulty levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。