用单步扩散模型提升3D重建的逼真度和视角泛化能力。
Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models

- 通过单步扩散模型Difix,修复3D重建中的伪影与欠约束区域。
- 在两种主流3D表示上实现平均FID分数提升2倍。
- 无需修改原有流程,适配NeRF与3DGS,通用性强。
神经辐射场(NeRF)和3D高斯泼溅(3DGS)已革新3D重建与新视角合成任务。然而,从极端新视角生成逼真渲染仍具挑战,因不同表示中持续存在伪影。本文提出Difix3D+,一种基于单步扩散模型的新管道,以提升3D重建与新视角合成质量。核心是训练一个名为Difix的单步图像扩散模型,用于增强并去除由3D表示中欠约束区域引发的渲染伪影。Difix在重建阶段用于清理从3D重建生成的伪训练视图,并将其回蒸馏至3D空间,显著改善欠约束区域,提升整体3D表示质量。更重要的是,它在推理阶段作为神经增强器,有效消除由不完美3D监督与现有重建模型容量限制带来的残留伪影。Difix3D+是通用方案,兼容NeRF与3DGS,实现平均2×的FID分数提升,同时保持3D一致性。
原文摘要 · Abstract (English)
Neural Radiance Fields and 3D Gaussian Splatting have revolutionized 3D reconstruction and novel-view synthesis task. However, achieving photorealistic rendering from extreme novel viewpoints remains challenging, as artifacts persist across representations. In this work, we introduce Difix3D+, a novel pipeline designed to enhance 3D reconstruction and novel-view synthesis through single-step diffusion models. At the core of our approach is Difix, a single-step image diffusion model trained to enhance and remove artifacts in rendered novel views caused by underconstrained regions of the 3D representation. Difix serves two critical roles in our pipeline. First, it is used during the reconstruction phase to clean up pseudo-training views that are rendered from the reconstruction and then distilled back into 3D. This greatly enhances underconstrained regions and improves the overall 3D representation quality. More importantly, Difix also acts as a neural enhancer during inference, effectively removing residual artifacts arising from imperfect 3D supervision and the limited capacity of current reconstruction models. Difix3D+ is a general solution, a single model compatible with both NeRF and 3DGS representations, and it achieves an average 2$\times$ improvement in FID score over baselines while maintaining 3D consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。