用单对图像生成文本提示,实现高保真图像编辑。
Textualize Visual Prompt for Image Editing via Diffusion Bridge
- 通过扩散桥在文本引导下转换前后图像分布。
- 仅需一对图像即可生成可迁移的文本嵌入,效果媲美多模型方法。
- 适合需要快速、通用图像编辑的开发者与设计师。
视觉提示(一对编辑前后的图像)能表达难以描述的图像变化,在图像编辑中表现优异。但现有方法依赖预训练的文本引导图像到图像生成模型,需提供文本、编辑前和编辑后三者组成的三元组进行重训练,限制了可扩展性和泛化能力。本文提出一种基于任意单个文生图模型的框架,无需显式图像到图像模型,显著提升泛化与可扩展性。具体地,利用概率流常微分方程构建扩散桥,在文本引导下实现前后图像分布的转移;通过优化文本嵌入,自适应将视觉提示中的编辑变换转化为文本表示,无需额外模型。同时引入差异注意力控制,在优化过程中解耦文本嵌入与前后图像的不变特征,使其仅捕捉细微变换,并可推广至不同图像。在真实图像上的实验验证了该方法在泛化性、上下文一致性及高保真度方面具有竞争力,仅需一对图像作为视觉提示即可实现精细编辑。
原文摘要 · Abstract (English)
Visual prompt, a pair of before-and-after edited images, can convey indescribable imagery transformations and prosper in image editing. However, current visual prompt methods rely on a pretrained text-guided image-to-image generative model that requires a triplet of text, before, and after images for retraining over a text-to-image model. Such crafting triplets and retraining processes limit the scalability and generalization of editing. In this paper, we present a framework based on any single text-to-image model without reliance on the explicit image-to-image model thus enhancing the generalizability and scalability. Specifically, by leveraging the probability-flow ordinary equation, we construct a diffusion bridge to transfer the distribution between before-and-after images under the text guidance. By optimizing the text via the bridge, the framework adaptively textualizes the editing transformation conveyed by visual prompts into text embeddings without other models. Meanwhile, we introduce differential attention control during text optimization, which disentangles the text embedding from the invariance of the before-and-after images and makes it solely capture the delicate transformation and generalize to edit various images. Experiments on real images validate competitive results on the generalization, contextual coherence, and high fidelity for delicate editing with just one image pair as the visual prompt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。