提出新方法让流模型无需调参就能灵活编辑图像
Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing
- 分两阶段优化反演,提升图像投影精度
- 通过文本特征调控实现刚性与非刚性编辑统一控制
- 支持文字、数量、表情等多种编辑类型,适合通用图像编辑场景
利用流变换器的大型生成先验实现无调参图像编辑,需真实反演将图像映射到模型空间,并具备灵活的不变性控制机制以保留非目标内容。然而,现有扩散反演在基于流的模型中表现不佳,且不变性控制难以兼顾刚性与非刚性编辑任务。为此,我们系统分析了流变换器中的反演与不变性控制。具体地,发现欧拉反演结构类似DDIM,但更易受近似误差影响。因此提出两阶段反演:先优化速度估计,再补偿残余误差,紧密依赖模型先验,显著提升编辑效果。同时提出基于自适应层归一化的不变性控制机制,将文本提示变化关联至图像语义。该机制可同时保留非目标内容,并支持刚性与非刚性操作,实现包括视觉文字、数量、面部表情等在内的多样化编辑。在多种场景下的实验验证表明,本框架实现灵活精准的编辑,充分释放流变换器在通用图像编辑中的潜力。
原文摘要 · Abstract (English)
Leveraging the large generative prior of the flow transformer for tuning-free image editing requires authentic inversion to project the image into the model's domain and a flexible invariance control mechanism to preserve non-target contents. However, the prevailing diffusion inversion performs deficiently in flow-based models, and the invariance control cannot reconcile diverse rigid and non-rigid editing tasks. To address these, we systematically analyze the \textbf{inversion and invariance} control based on the flow transformer. Specifically, we unveil that the Euler inversion shares a similar structure to DDIM yet is more susceptible to the approximation error. Thus, we propose a two-stage inversion to first refine the velocity estimation and then compensate for the leftover error, which pivots closely to the model prior and benefits editing. Meanwhile, we propose the invariance control that manipulates the text features within the adaptive layer normalization, connecting the changes in the text prompt to image semantics. This mechanism can simultaneously preserve the non-target contents while allowing rigid and non-rigid manipulation, enabling a wide range of editing types such as visual text, quantity, facial expression, etc. Experiments on versatile scenarios validate that our framework achieves flexible and accurate editing, unlocking the potential of the flow transformer for versatile image editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。