arXiv:2412.14166cs.CV2024-12CVPR被引 21

用合成数据训练3D重建,效果媲美真实数据。

MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

  • 用程序生成70万场景数据,规模超真实数据50倍
  • 仅用基础几何结构,提升生成效率并保持性能
  • 纯合成数据训练模型也能达到真实数据水平

我们提出通过合成数据来扩展3D场景重建的训练规模。核心是MegaSynth,一个由70万场景组成的程序化生成3D数据集,比先前的真实数据集DL3DV大50倍以上。为实现可扩展的数据生成,关键思路是去除语义信息,无需建模复杂语义先验如物体功能和场景构成,而是采用基本空间结构与几何原型,保障可扩展性。同时控制数据复杂度以促进训练,松散对齐真实数据分布以提升真实世界泛化能力。我们探索了在MegaSynth与现有真实数据上联合训练或预训练低分辨率模型(LRMs)的效果。实验表明,联合训练或预训练可使重建质量在多个图像域中提升1.2至1.8 dB PSNR。此外,仅使用MegaSynth训练的模型表现可媲美真实数据训练模型,凸显3D重建任务的底层特性。我们还深入分析了MegaSynth对模型能力、训练稳定性和泛化性的提升作用,以及其在其他任务中的应用潜力。

原文摘要 · Abstract (English)

We propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes - over 50 times larger than the prior real dataset DL3DV - dramatically scaling the training data. To enable scalable data generation, our key idea is eliminating semantic information, removing the need to model complex semantic priors such as object affordances and scene composition. Instead, we model scenes with basic spatial structures and geometry primitives, offering scalability. Besides, we control data complexity to facilitate training while loosely aligning it with real-world data distribution to benefit real-world generalization. We explore training LRMs with both MegaSynth and available real data. Experiment results show that joint training or pre-training with MegaSynth improves reconstruction quality by 1.2 to 1.8 dB PSNR across diverse image domains. Moreover, models trained solely on MegaSynth perform comparably to those trained on real data, underscoring the low-level nature of 3D reconstruction. Additionally, we provide an in-depth analysis of MegaSynth's properties for enhancing model capability, training stability, and generalization, as well as application to other tasks.

3D重建合成数据可扩展几何建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。