用2D扩散模型生成3D物体,通过高斯点云保证多视角一致性。
GSV3D: Gaussian Splatting-based Geometric Distillation with Stable Video Diffusion for Single-Image 3D Object Generation
- 用高斯点云将2D扩散输出转为显式3D表示,约束几何一致性。
- 在多个数据集上实现顶尖的多视角一致性和3D结构准确性。
- 适合需要高质量3D模型和稳定视图的机器人与游戏应用。
基于图像的3D生成在机器人和游戏领域有广泛应用,高质量、多样化的输出及一致的3D表征至关重要。现有方法存在局限:3D扩散模型受限于数据稀缺和缺乏强预训练先验,而基于2D扩散的方法则难以保证几何一致性。本文提出一种新方法,利用2D扩散模型的隐式3D推理能力,通过基于高斯点云的几何蒸馏确保3D一致性。具体而言,所提出的高斯点云解码器将SV3D潜在输出转换为显式3D表示,显式编码空间与外观属性,通过几何约束实现多视角一致性,纠正视图不一致问题,确保鲁棒的几何一致性。结果表明,该方法可同时生成高质量、多视角一致的图像和精确3D模型,为单图像3D生成提供可扩展解决方案,弥合了2D扩散多样性与3D结构连贯性之间的差距。实验验证其在多个数据集上达到最先进的多视角一致性表现并具备强泛化能力。代码将在接受后公开。
原文摘要 · Abstract (English)
Image-based 3D generation has vast applications in robotics and gaming, where high-quality, diverse outputs and consistent 3D representations are crucial. However, existing methods have limitations: 3D diffusion models are limited by dataset scarcity and the absence of strong pre-trained priors, while 2D diffusion-based approaches struggle with geometric consistency. We propose a method that leverages 2D diffusion models' implicit 3D reasoning ability while ensuring 3D consistency via Gaussian-splatting-based geometric distillation. Specifically, the proposed Gaussian Splatting Decoder enforces 3D consistency by transforming SV3D latent outputs into an explicit 3D representation. Unlike SV3D, which only relies on implicit 2D representations for video generation, Gaussian Splatting explicitly encodes spatial and appearance attributes, enabling multi-view consistency through geometric constraints. These constraints correct view inconsistencies, ensuring robust geometric consistency. As a result, our approach simultaneously generates high-quality, multi-view-consistent images and accurate 3D models, providing a scalable solution for single-image-based 3D generation and bridging the gap between 2D Diffusion diversity and 3D structural coherence. Experimental results demonstrate state-of-the-art multi-view consistency and strong generalization across diverse datasets. The code will be made publicly available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。