用扩散模型去噪多视角图像,生成高质量3D资产。
DSplats: 3D Generation by Denoising Splats-Based Multiview Diffusion Models
- 用高斯点云重构器结合预训练扩散模型去噪多视角图。
- 单图重建在谷歌扫描物体数据集上达到PSNR 20.38、SSIM 0.842。
- 生成结果几何一致,适合需要真实感3D建模的场景。
生成高质量3D内容需要能学习复杂场景和真实物体分布的模型。近期基于高斯的3D重建方法通过前馈方式预测3D高斯,在稀疏输入图像下实现了高保真3D资产恢复。然而,这些方法常缺乏扩散模型提供的丰富先验与表达力。另一方面,已成功应用于去噪多视角图像的2D扩散模型展现出生成广泛逼真3D输出的潜力,但仍缺乏显式的3D先验与一致性。本文提出DSplats,一种新方法,直接使用基于高斯点云的重构器对多视角图像进行去噪,生成多样且真实的3D资产。为利用2D扩散模型的丰富先验,我们在重构器主干中引入预训练的潜在扩散模型以预测一组3D高斯。此外,去噪网络中嵌入的显式3D表示提供了强归纳偏置,确保了新视角生成的几何一致性。定性与定量实验表明,DSplats不仅生成高质量且空间一致的结果,还在单图到3D重建任务中树立了新标准。在Google Scanned Objects数据集上,其PSNR达20.38,SSIM为0.842,LPIPS为0.109。
原文摘要 · Abstract (English)
Generating high-quality 3D content requires models capable of learning robust distributions of complex scenes and the real-world objects within them. Recent Gaussian-based 3D reconstruction techniques have achieved impressive results in recovering high-fidelity 3D assets from sparse input images by predicting 3D Gaussians in a feed-forward manner. However, these techniques often lack the extensive priors and expressiveness offered by Diffusion Models. On the other hand, 2D Diffusion Models, which have been successfully applied to denoise multiview images, show potential for generating a wide range of photorealistic 3D outputs but still fall short on explicit 3D priors and consistency. In this work, we aim to bridge these two approaches by introducing DSplats, a novel method that directly denoises multiview images using Gaussian Splat-based Reconstructors to produce a diverse array of realistic 3D assets. To harness the extensive priors of 2D Diffusion Models, we incorporate a pretrained Latent Diffusion Model into the reconstructor backbone to predict a set of 3D Gaussians. Additionally, the explicit 3D representation embedded in the denoising network provides a strong inductive bias, ensuring geometrically consistent novel view generation. Our qualitative and quantitative experiments demonstrate that DSplats not only produces high-quality, spatially consistent outputs, but also sets a new standard in single-image to 3D reconstruction. When evaluated on the Google Scanned Objects dataset, DSplats achieves a PSNR of 20.38, an SSIM of 0.842, and an LPIPS of 0.109.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。