仅用一张图生成高质量3D场景,突破多视角依赖
GEN3D: Generating Domain-Free 3D Scenes from a Single Image
- 从单张图像生成点云,再通过高斯点优化生成3D场景
- 在多个数据集上实现高保真、视角一致的3D重建
- 适合需要多样化3D世界模型的智能体与仿真系统
尽管神经3D重建技术取得进展,但其对密集多视角数据的依赖限制了广泛应用。3D场景生成对推动具身智能与世界模型发展至关重要,需多样且高质量场景用于学习与评估。本文提出Gen3d,一种从单张图像生成高质量、大范围、通用3D场景的新方法。首先通过提升RGBD图像生成初始点云,随后维持并扩展其世界模型,最终通过优化高斯点表示完成3D场景构建。在多个数据集上的大量实验表明,该方法具有强泛化能力,能生成高质量世界模型,并合成高保真、一致的全新视角。
原文摘要 · Abstract (English)
Despite recent advancements in neural 3D reconstruction, the dependence on dense multi-view captures restricts their broader applicability. Additionally, 3D scene generation is vital for advancing embodied AI and world models, which depend on diverse, high-quality scenes for learning and evaluation. In this work, we propose Gen3d, a novel method for generation of high-quality, wide-scope, and generic 3D scenes from a single image. After the initial point cloud is created by lifting the RGBD image, Gen3d maintains and expands its world model. The 3D scene is finalized through optimizing a Gaussian splatting representation. Extensive experiments on diverse datasets demonstrate the strong generalization capability and superior performance of our method in generating a world model and Synthesizing high-fidelity and consistent novel views.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。