arXiv:2503.15742cs.CV2025-03ICCV被引 2

用扩散模型提升单图3D重建的清晰度与真实感

Uncertainty-Aware Diffusion Guided Refinement of 3D Scenes

  • 用预训练扩散模型迭代优化高斯表示的粗略3D场景
  • 在KITTI-v2和RealEstate-10K上实现更清晰的新视角生成
  • 通过不确定性图指导精修,优先处理可信区域

从单张图像重建3D场景是严重欠约束的问题,导致现有方法在新视角渲染时出现不连贯和模糊现象,尤其在远离输入相机的未见区域更为严重。本文针对这一问题,提出一种基于不确定性感知的扩散引导精修方法:利用预训练的潜在视频扩散模型,对由可优化高斯参数表示的粗略场景进行迭代优化;为保证生成图像与输入图像在风格和纹理上一致,引入实时傅里叶风格迁移;设计语义不确定性量化模块,计算像素级熵并生成不确定性图,指导精修过程优先聚焦高置信度区域,舍弃高不确定部分。在真实场景数据集RealEstate-10K(域内)和KITTI-v2(域外)上进行大量实验,结果表明本方法相比现有最先进方法能生成更真实、高保真的新视角图像。

原文摘要 · Abstract (English)

Reconstructing 3D scenes from a single image is a fundamentally ill-posed task due to the severely under-constrained nature of the problem. Consequently, when the scene is rendered from novel camera views, existing single image to 3D reconstruction methods render incoherent and blurry views. This problem is exacerbated when the unseen regions are far away from the input camera. In this work, we address these inherent limitations in existing single image-to-3D scene feedforward networks. To alleviate the poor performance due to insufficient information beyond the input image's view, we leverage a strong generative prior in the form of a pre-trained latent video diffusion model, for iterative refinement of a coarse scene represented by optimizable Gaussian parameters. To ensure that the style and texture of the generated images align with that of the input image, we incorporate on-the-fly Fourier-style transfer between the generated images and the input image. Additionally, we design a semantic uncertainty quantification module that calculates the per-pixel entropy and yields uncertainty maps used to guide the refinement process from the most confident pixels while discarding the remaining highly uncertain ones. We conduct extensive experiments on real-world scene datasets, including in-domain RealEstate-10K and out-of-domain KITTI-v2, showing that our approach can provide more realistic and high-fidelity novel view synthesis results compared to existing state-of-the-art methods.

3D重建扩散模型不确定性建模新视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。