通过改进注意力与潜在表示,实现更精准的图像视频编辑。
ProEdit: Inversion-based Editing From Prompts Done Right
- 引入KV-mix混合源目标区域的键值特征,减少源图干扰。
- 设计潜空间偏移机制,消除反演潜变量对生成的影响。
- 可无缝集成到现有编辑框架,支持图像视频多场景应用。
基于反演的视觉编辑提供了一种无需训练的有效图像或视频编辑方法。现有方法通常在采样过程中注入源图信息以保持编辑一致性,但这种策略过度依赖源图信息,导致目标图像中属性修改失败(如姿势、数量或颜色变化)。本文提出ProEdit,从注意力和潜在表示两方面解决该问题。在注意力层面,提出KV-mix,混合编辑区域的源图与目标图键值特征,降低源图影响并保持背景一致;在潜在层面,提出Latents-Shift,扰动源图潜在表示中的编辑区域,消除反演潜变量对采样过程的影响。在多个图像与视频编辑基准测试中,本方法达到当前最优性能。此外,其设计为即插即用,可无缝集成至现有反演与编辑方法,如RF-Solver、FireFlow和UniEdit。
原文摘要 · Abstract (English)
Inversion-based visual editing provides an effective and training-free way to edit an image or a video based on user instructions. Existing methods typically inject source image information during the sampling process to maintain editing consistency. However, this sampling strategy overly relies on source information, which negatively affects the edits in the target image (e.g., failing to change the subject's atributes like pose, number, or color as instructed). In this work, we propose ProEdit to address this issue both in the attention and the latent aspects. In the attention aspect, we introduce KV-mix, which mixes KV features of the source and the target in the edited region, mitigating the influence of the source image on the editing region while maintaining background consistency. In the latent aspect, we propose Latents-Shift, which perturbs the edited region of the source latent, eliminating the influence of the inverted latent on the sampling. Extensive experiments on several image and video editing benchmarks demonstrate that our method achieves SOTA performance. In addition, our design is plug-and-play, which can be seamlessly integrated into existing inversion and editing methods, such as RF-Solver, FireFlow and UniEdit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。