通过增强提示词实现更精准的文本引导图像编辑
Prompt Augmentation for Self-supervised Text-guided Image Manipulation
- 将单个提示词扩展为多个目标提示,定位编辑区域
- 设计对比损失使修改区移开、保留区靠近,提升一致性
- 适合需要精细控制图像局部变化的研究者和创作者
文本引导图像编辑在创意与实用领域有广泛应用。尽管图像生成研究取得进展,但仍面临图像变换连贯性与上下文保持的双重挑战。为此,本文提出提示词增强方法,将单一输入提示扩展为多个目标提示,强化文本语境并实现局部编辑。具体而言,利用增强后的提示界定预期操作区域,并提出一种针对性的对比损失,通过移动编辑区域、拉近保留区域来驱动有效编辑。考虑到图像编辑的连续特性,进一步引入相似性概念,构建软对比损失。新损失被融入扩散模型,在公开数据集和生成图像上均取得优于或相当当前最优方法的编辑效果。
原文摘要 · Abstract (English)
Text-guided image editing finds applications in various creative and practical fields. While recent studies in image generation have advanced the field, they often struggle with the dual challenges of coherent image transformation and context preservation. In response, our work introduces prompt augmentation, a method amplifying a single input prompt into several target prompts, strengthening textual context and enabling localised image editing. Specifically, we use the augmented prompts to delineate the intended manipulation area. We propose a Contrastive Loss tailored to driving effective image editing by displacing edited areas and drawing preserved regions closer. Acknowledging the continuous nature of image manipulations, we further refine our approach by incorporating the similarity concept, creating a Soft Contrastive Loss. The new losses are incorporated to the diffusion model, demonstrating improved or competitive image editing results on public datasets and generated images over state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。