用合成分析法自动标注野外3D物体,提升真实场景检测效果
LabelAny3D: Label Any Object 3D in the Wild
- 通过2D图像重建3D场景,自动生成高质量3D框标注
- 在COCO3D数据集上,标注提升多基准检测性能
- 适合做开放词汇3D检测的算法研究者
从单目图像中检测3D空间中的物体对机器人和场景理解至关重要。尽管在室内和自动驾驶领域表现先进,现有单目3D检测模型在野外图像上仍面临挑战,主要因缺乏3D野外数据集及3D标注困难。本文提出LabelAny3D,一种基于“分析-合成”的框架,通过从2D图像重建完整3D场景,高效生成高质量3D边界框标注。基于此流程,构建了新基准COCO3D,源自MS-COCO数据集,涵盖大量现有3D数据集缺失的物体类别。实验表明,LabelAny3D生成的标注显著提升多个基准上的单目3D检测性能,优于以往自动标注方法。结果证明,基于基础模型的标注方法在真实开放世界3D识别中具有广阔前景。
原文摘要 · Abstract (English)
Detecting objects in 3D space from monocular input is crucial for applications ranging from robotics to scene understanding. Despite advanced performance in the indoor and autonomous driving domains, existing monocular 3D detection models struggle with in-the-wild images due to the lack of 3D in-the-wild datasets and the challenges of 3D annotation. We introduce LabelAny3D, an \emph{analysis-by-synthesis} framework that reconstructs holistic 3D scenes from 2D images to efficiently produce high-quality 3D bounding box annotations. Built on this pipeline, we present COCO3D, a new benchmark for open-vocabulary monocular 3D detection, derived from the MS-COCO dataset and covering a wide range of object categories absent from existing 3D datasets. Experiments show that annotations generated by LabelAny3D improve monocular 3D detection performance across multiple benchmarks, outperforming prior auto-labeling approaches in quality. These results demonstrate the promise of foundation-model-driven annotation for scaling up 3D recognition in realistic, open-world settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。