arXiv:2508.21542cs.CVcs.AI2025-08被引 2

仅用一张图就能生成完整3D场景的高保真重建结果。

Complete Gaussian Splats from a Single Image with Denoising Diffusion Models

  • 用扩散模型生成3D高斯点云分布,而非单一预测值。
  • 在单张图像下完成遮挡区域重建,支持360度高质量渲染。
  • 无需真实3D标注,自监督训练提升泛化能力。

高斯点阵通常需要密集视角观测,难以重建遮挡和未观测区域。本文提出一种潜在扩散模型,仅凭单张图像即可生成完整的3D场景(含遮挡部分)的高斯点云表示。由于合理表面形态存在歧义,传统方法采用回归方式预测单一“模式”,导致模糊、不真实或无法捕捉多解。为此,我们采用生成式建模,学习基于单张图像的3D高斯点云分布。为解决缺乏真实标签的问题,设计了变分自重构器,仅通过2D图像实现自监督学习,构建潜在空间并训练扩散模型。该方法能生成高保真且多样化的重建结果,有效完成遮挡表面补全,支持高质量全景渲染。

原文摘要 · Abstract (English)

Gaussian splatting typically requires dense observations of the scene and can fail to reconstruct occluded and unobserved areas. We propose a latent diffusion model to reconstruct a complete 3D scene with Gaussian splats, including the occluded parts, from only a single image during inference. Completing the unobserved surfaces of a scene is challenging due to the ambiguity of the plausible surfaces. Conventional methods use a regression-based formulation to predict a single "mode" for occluded and out-of-frustum surfaces, leading to blurriness, implausibility, and failure to capture multiple possible explanations. Thus, they often address this problem partially, focusing either on objects isolated from the background, reconstructing only visible surfaces, or failing to extrapolate far from the input views. In contrast, we propose a generative formulation to learn a distribution of 3D representations of Gaussian splats conditioned on a single input image. To address the lack of ground-truth training data, we propose a Variational AutoReconstructor to learn a latent space only from 2D images in a self-supervised manner, over which a diffusion model is trained. Our method generates faithful reconstructions and diverse samples with the ability to complete the occluded surfaces for high-quality 360-degree renderings.

3D重建扩散模型单图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。