用扩散模型实现布局调整与物体外观一致的图像编辑
Consistent Image Layout Editing with Diffusion Models
- 通过多概念学习保留物体视觉特征,实现布局重排
- 利用扩散模型中间特征保持物体外观一致性,效果优于现有方法
- 适合需要精准布局修改且保持真实感的图像编辑场景
尽管大规模文本到图像扩散模型在图像生成和编辑中取得显著进展,现有方法在真实图像的布局编辑方面仍存在困难。部分工作虽尝试解决此问题,但或无法有效调整布局,或难以保持物体编辑后的视觉一致性。为此,本文提出一种新型图像布局编辑方法,不仅能将真实图像重新排列至指定布局,还能确保物体外观与编辑前一致。具体而言,该方法包含两个关键组件:首先,采用多概念学习方案从单张图像中学习不同物体的概念,对保持视觉一致性至关重要;其次,利用扩散模型中间特征中的语义一致性,直接将物体外观信息投影至目标区域。此外,引入一种新颖的初始化噪声设计,以促进布局重排过程。大量实验表明,该方法在布局对齐和视觉一致性方面均优于先前工作。
原文摘要 · Abstract (English)
Despite the great success of large-scale text-to-image diffusion models in image generation and image editing, existing methods still struggle to edit the layout of real images. Although a few works have been proposed to tackle this problem, they either fail to adjust the layout of images, or have difficulty in preserving visual appearance of objects after the layout adjustment. To bridge this gap, this paper proposes a novel image layout editing method that can not only re-arrange a real image to a specified layout, but also can ensure the visual appearance of the objects consistent with their appearance before editing. Concretely, the proposed method consists of two key components. Firstly, a multi-concept learning scheme is used to learn the concepts of different objects from a single image, which is crucial for keeping visual consistency in the layout editing. Secondly, it leverages the semantic consistency within intermediate features of diffusion models to project the appearance information of objects to the desired regions directly. Besides, a novel initialization noise design is adopted to facilitate the process of re-arranging the layout. Extensive experiments demonstrate that the proposed method outperforms previous works in both layout alignment and visual consistency for the task of image layout editing
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。