用扩散模型补全稀疏视角,提升3D重建的细节与精度
Pragmatist: Multiview Conditional Diffusion Models for High-Fidelity 3D Reconstruction from Unposed Sparse Views
- 通过多视角条件扩散模型生成完整观测图像
- 在多个基准上实现优于现有方法的重建质量
- 适合需要高保真3D重建的研究者和工业应用
从稀疏、无姿态的观测中恢复3D结构极具挑战性。现有方法虽能直接从无姿态输入预测隐式表示,但缺乏几何先验,无法推测未见区域的外观,难以还原精细几何与纹理。为此,我们提出将该病态问题重述为条件新视角合成任务:从有限输入视图生成完整观测以辅助重建。首先,利用多视角条件扩散模型生成物体的完整观测;其次,通过前馈大型重建模型获取三维网格;最后,通过反演3D表示恢复输入视角姿态,并结合详细输入视图优化纹理。相比以往方法,本框架高效利用无姿态输入与生成先验,避免直接求解高度病态问题。大量实验表明,该方法在多个基准上表现优异。
原文摘要 · Abstract (English)
Inferring 3D structures from sparse, unposed observations is challenging due to its unconstrained nature. Recent methods propose to predict implicit representations directly from unposed inputs in a data-driven manner, achieving promising results. However, these methods do not utilize geometric priors and cannot hallucinate the appearance of unseen regions, thus making it challenging to reconstruct fine geometric and textural details. To tackle this challenge, our key idea is to reformulate this ill-posed problem as conditional novel view synthesis, aiming to generate complete observations from limited input views to facilitate reconstruction. With complete observations, the poses of the input views can be easily recovered and further used to optimize the reconstructed object. To this end, we propose a novel pipeline Pragmatist. First, we generate a complete observation of the object via a multiview conditional diffusion model. Then, we use a feed-forward large reconstruction model to obtain the reconstructed mesh. To further improve the reconstruction quality, we recover the poses of input views by inverting the obtained 3D representations and further optimize the texture using detailed input views. Unlike previous approaches, our pipeline improves reconstruction by efficiently leveraging unposed inputs and generative priors, circumventing the direct resolution of highly ill-posed problems. Extensive experiments show that our approach achieves promising performance in several benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。