用互联网视频自动生成3D场景训练数据,提升模型泛化能力
Lifting Unlabeled Internet-level Data for 3D Scene Understanding

- 设计数据引擎,从网络视频中自动生成3D训练数据
- 在物体检测、分割和视觉问答任务上实现强零样本性能
- 适合需要大规模无标注数据的3D感知研究者
标注的3D场景数据稀缺且成本高,而互联网上存在大量未标注视频。本文表明,通过精心设计的数据引擎,可利用网络获取的未标注视频自动生成训练数据,辅助端到端3D场景理解模型的学习,同时结合人工标注数据集。我们识别并分析了自动化数据生成中的瓶颈,揭示了影响学习效率与效果的关键因素。为验证方法在不同感知粒度下的有效性,我们在三个任务上进行评估:低层感知如3D物体检测与实例分割,高层推理如3D空间视觉问答(VQA)和视觉语言导航(VLN)。在生成数据上训练的模型展现出优异的零样本性能,并在微调后进一步提升。结果表明,利用可获取的网络数据是构建更强大场景理解系统的一条可行路径。
原文摘要 · Abstract (English)
Annotated 3D scene data is scarce and expensive to acquire, while abundant unlabeled videos are readily available on the internet. In this paper, we demonstrate that carefully designed data engines can leverage web-curated, unlabeled videos to automatically generate training data, to facilitate end-to-end models in 3D scene understanding alongside human-annotated datasets. We identify and analyze bottlenecks in automated data generation, revealing critical factors that determine the efficiency and effectiveness of learning from unlabeled data. To validate our approach across different perception granularities, we evaluate on three tasks spanning low-level perception, i.e., 3D object detection and instance segmentation, to high-evel reasoning, i.e., 3D spatial Visual Question Answering (VQA) and Vision-Lanugage Navigation (VLN). Models trained on our generated data demonstrate strong zero-shot performance and show further improvement after finetuning. This demonstrates the viability of leveraging readily available web data as a path toward more capable scene understanding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。