用双对比去噪得分实现单参考文生图编辑,保形又改内容。
Single-Reference Text-to-Image Manipulation with Dual Contrastive Denoising Score
- 引入双对比损失,利用扩散模型中间特征进行无辅助网络的编辑
- 在真实图像上实现灵活修改与结构保持,无需微调预训练模型
- 支持零样本图像到图像转换,适合直接编辑已有图像
大规模文生图生成模型展现出强大的多样性和高质量图像合成能力。然而,直接用于编辑真实图像仍面临两大挑战:用户难以写出精确描述输入图像所有视觉细节的理想文本提示;现有方法虽能在特定区域引入期望变化,却常大幅改变输入内容,并在非目标区域引入意外改动。为此,我们提出双对比去噪得分(Dual Contrastive Denoising Score)框架,充分利用文生图扩散模型丰富的生成先验。受无配对图像到图像翻译中对比学习的启发,我们在框架中引入简洁的双对比损失。该方法利用潜空间扩散模型自注意力层的丰富空间信息,无需依赖额外网络。我们的方法实现了输入输出间的内容灵活修改与结构保留,同时支持零样本图像到图像转换。大量实验表明,本方法在真实图像编辑任务上优于现有方法,且可直接使用预训练的文生图扩散模型,无需额外训练。
原文摘要 · Abstract (English)
Large-scale text-to-image generative models have shown remarkable ability to synthesize diverse and high-quality images. However, it is still challenging to directly apply these models for editing real images for two reasons. First, it is difficult for users to come up with a perfect text prompt that accurately describes every visual detail in the input image. Second, while existing models can introduce desirable changes in certain regions, they often dramatically alter the input content and introduce unexpected changes in unwanted regions. To address these challenges, we present Dual Contrastive Denoising Score, a simple yet powerful framework that leverages the rich generative prior of text-to-image diffusion models. Inspired by contrastive learning approaches for unpaired image-to-image translation, we introduce a straightforward dual contrastive loss within the proposed framework. Our approach utilizes the extensive spatial information from the intermediate representations of the self-attention layers in latent diffusion models without depending on auxiliary networks. Our method achieves both flexible content modification and structure preservation between input and output images, as well as zero-shot image-to-image translation. Through extensive experiments, we show that our approach outperforms existing methods in real image editing while maintaining the capability to directly utilize pretrained text-to-image diffusion models without further training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。