用像素空间扩散模型实现单图新视角生成,效果超越现有最佳方法。
Novel View Synthesis with Pixel-Space Diffusion Models
- 直接在像素空间使用现代扩散模型,端到端完成新视角生成。
- 相比之前最先进方法,重建质量显著提升,尤其在复杂场景下表现更优。
- 提出新训练方案,仅需单视图数据即可训练,泛化能力更强。
从单张输入图像生成新视角是一项挑战性任务。传统方法通过估计场景深度、图像扭曲和修补来实现,机器学习模型用于其中部分环节。近年来,生成模型越来越多地被用于新视角合成(NVS),甚至构成端到端系统。本文将现代扩散模型架构适配到像素空间的端到端NVS中,显著优于先前的最先进方法。我们探索了多种将几何信息编码进网络的方式,实验表明这些方法虽能提升性能,但影响有限,远不及改进生成模型本身的效果。此外,我们提出一种新的NVS训练方案,利用相对丰富的单视图数据集,提升了模型对域外内容场景的泛化能力。
原文摘要 · Abstract (English)
Synthesizing a novel view from a single input image is a challenging task. Traditionally, this task was approached by estimating scene depth, warping, and inpainting, with machine learning models enabling parts of the pipeline. More recently, generative models are being increasingly employed in novel view synthesis (NVS), often encompassing the entire end-to-end system. In this work, we adapt a modern diffusion model architecture for end-to-end NVS in the pixel space, substantially outperforming previous state-of-the-art (SOTA) techniques. We explore different ways to encode geometric information into the network. Our experiments show that while these methods may enhance performance, their impact is minor compared to utilizing improved generative models. Moreover, we introduce a novel NVS training scheme that utilizes single-view datasets, capitalizing on their relative abundance compared to their multi-view counterparts. This leads to improved generalization capabilities to scenes with out-of-domain content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。