用双图对训练扩散模型,让修复图像保持3D一致性。
3D-Consistent Image Inpainting with Diffusion Models
- 用同一场景的两张图作为输入,引导修复过程
- 在合成与真实数据集上均优于现有方法
- 无需3D标注,适合图像修复与视觉生成任务
针对基于扩散模型的图像修复中存在的3D不一致性问题,我们提出一种使用同一场景图像对的生成模型。通过在去噪过程中引入场景的另一视角,构建归纳偏置,使模型在2D去噪训练中恢复3D先验,无需显式3D监督。利用额外图像作为上下文引导训练无条件扩散模型,实现掩码与非掩码区域的语义一致修复,并保证3D一致性。我们在一个合成数据集和三个真实世界数据集上评估该方法,结果表明其生成的修复图像在语义和3D结构上均具一致性,性能超越当前最优方法。
原文摘要 · Abstract (English)
We address the problem of 3D inconsistency of image inpainting based on diffusion models. We propose a generative model using image pairs that belong to the same scene. To achieve the 3D-consistent and semantically coherent inpainting, we modify the generative diffusion model by incorporating an alternative point of view of the scene into the denoising process. This creates an inductive bias that allows to recover 3D priors while training to denoise in 2D, without explicit 3D supervision. Training unconditional diffusion models with additional images as in-context guidance allows to harmonize the masked and non-masked regions while repainting and ensures the 3D consistency. We evaluate our method on one synthetic and three real-world datasets and show that it generates semantically coherent and 3D-consistent inpaintings and outperforms the state-of-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。