用视觉信息增强文本提示,实现更精准的图像编辑。
Vision-guided and Mask-enhanced Adaptive Denoising for Prompt-based Image Editing
- 引入图像嵌入辅助文本提示,提升生成引导能力。
- 通过自注意力迭代优化编辑区域定位,精度更高。
- 按区域自适应调整去噪强度,关键部分修改更充分。
基于文本的图像编辑在文生图扩散模型基础上快速发展,但现有方法仍存在三大问题:1)文本提示引导能力有限;2)未能充分挖掘词与图像块、块与块之间的关联关系以精确定位编辑区域;3)每个去噪步骤对所有区域采用统一编辑强度。为此,本文提出视觉引导与掩码增强的自适应编辑方法(ViMAEdit),包含三项创新设计:首先,利用基于CLIP的图像嵌入估计策略,将图像特征作为显式引导信号融入去噪过程;其次,设计自注意力引导的迭代区域定位策略,通过自注意力图迭代挖掘块间关系,优化交叉注意力中的词-块对应关系;最后,提出空间自适应方差引导采样机制,增强关键区域的采样方差以强化编辑效果。实验表明,ViMAEdit在多项指标上均优于现有方法。
原文摘要 · Abstract (English)
Text-to-image diffusion models have demonstrated remarkable progress in synthesizing high-quality images from text prompts, which boosts researches on prompt-based image editing that edits a source image according to a target prompt. Despite their advances, existing methods still encounter three key issues: 1) limited capacity of the text prompt in guiding target image generation, 2) insufficient mining of word-to-patch and patch-to-patch relationships for grounding editing areas, and 3) unified editing strength for all regions during each denoising step. To address these issues, we present a Vision-guided and Mask-enhanced Adaptive Editing (ViMAEdit) method with three key novel designs. First, we propose to leverage image embeddings as explicit guidance to enhance the conventional textual prompt-based denoising process, where a CLIP-based target image embedding estimation strategy is introduced. Second, we devise a self-attention-guided iterative editing area grounding strategy, which iteratively exploits patch-to-patch relationships conveyed by self-attention maps to refine those word-to-patch relationships contained in cross-attention maps. Last, we present a spatially adaptive variance-guided sampling, which highlights sampling variances for critical image regions to promote the editing capability. Experimental results demonstrate the superior editing capacity of ViMAEdit over all existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。