arXiv:2412.14464cs.CVcs.GR2024-12

通过分阶段3D重建与扩散模型,从单图或少图生成高质量新视角图像。

LiftRefine: Progressively Refined View Synthesis from 3D Lifting with Volume-Triplane Representations

  • 先用体素和三平面结构将输入图升维为粗细两级3D表示
  • 在遮挡区域用扩散模型逐步补全缺失细节,提升渲染质量
  • 适合需要高保真3D重建的场景,尤其对少样本输入表现优异

我们提出一种新方法,通过单张或多张输入图像合成3D神经场以实现视图合成。针对图像到3D生成的病态性问题,设计了两阶段流程:先用重建模型将输入图像从体素升维至粗尺度3D表示,再转换为三平面精细表示;随后利用扩散模型从三平面中推测渲染图像中被遮挡区域的缺失细节。进一步引入渐进式精炼技术,迭代应用重建与扩散模型,逐步优化3D表示及其渲染效果。实证表明,该方法在合成数据集SRN-Car、真实世界数据集CO3D及大规模数据集Objaverse上均优于现有最先进方法,兼具采样效率与多视角一致性。

原文摘要 · Abstract (English)

We propose a new view synthesis method via synthesizing a 3D neural field from both single or few-view input images. To address the ill-posed nature of the image-to-3D generation problem, we devise a two-stage method that involves a reconstruction model and a diffusion model for view synthesis. Our reconstruction model first lifts one or more input images to the 3D space from a volume as the coarse-scale 3D representation followed by a tri-plane as the fine-scale 3D representation. To mitigate the ambiguity in occluded regions, our diffusion model then hallucinates missing details in the rendered images from tri-planes. We then introduce a new progressive refinement technique that iteratively applies the reconstruction and diffusion model to gradually synthesize novel views, boosting the overall quality of the 3D representations and their rendering. Empirical evaluation demonstrates the superiority of our method over state-of-the-art methods on the synthetic SRN-Car dataset, the in-the-wild CO3D dataset, and large-scale Objaverse dataset while achieving both sampling efficacy and multi-view consistency.

3D重建视图合成扩散模型三平面

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。