arXiv:2410.13924cs.CVcs.AI2024-10CVPR被引 10

构建超大规模3D场景数据集,推动室内场景理解性能突破

ARKit LabelMaker: A New Scale for Indoor 3D Scene Understanding

  • 基于ARKitScenes扩展,自动生成密集3D语义标签
  • 数据量超此前最大数据集3倍以上,显著提升分割准确率
  • 特别改善长尾类别表现,适合预训练与3D视觉研究者

神经网络性能随模型规模和数据量增长,这要求具备可扩展性的架构和大规模数据集。尽管变换器已用于3D视觉,但受限于训练数据不足,尚未出现类似GPT的突破。我们提出ARKit LabelMaker,一个大规模真实世界3D数据集,包含密集语义标注,其规模超过此前最大数据集三倍以上。具体而言,我们通过扩展LabelMaker流程,对ARKitScenes进行自动密集3D标注,专为大规模预训练设计。在该数据集上训练提升了多种架构的准确性,在ScanNet和ScanNet200上达到3D语义分割新纪录,尤其在长尾类别上表现显著提升。代码开源于https://labelmaker.org,数据集可在https://huggingface.co/datasets/labelmaker/arkit_labelmaker获取。

原文摘要 · Abstract (English)

Neural network performance scales with both model size and data volume, as shown in both language and image processing. This requires scaling-friendly architectures and large datasets. While transformers have been adapted for 3D vision, a `GPT-moment' remains elusive due to limited training data. We introduce ARKit LabelMaker, a large-scale real-world 3D dataset with dense semantic annotation that is more than three times larger than prior largest dataset. Specifically, we extend ARKitScenes with automatically generated dense 3D labels using an extended LabelMaker pipeline, tailored for large-scale pre-training. Training on our dataset improves accuracy across architectures, achieving state-of-the-art 3D semantic segmentation scores on ScanNet and ScanNet200, with notable gains on tail classes. Our code is available at https://labelmaker.org and our dataset at https://huggingface.co/datasets/labelmaker/arkit_labelmaker.

3D视觉数据集语义分割预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。