提出并行解耦编辑与速度衰减,提升复杂图像编辑的准确性与一致性。
FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing
- 将复杂编辑拆分为多个子任务并行处理,避免逐轮编辑误差累积。
- 通过衰减垂直于编辑方向的速度分量,更好保留源图像结构。
- 构建复杂编辑基准Complex-PIE-Bench,验证方法在多目标场景下的优势。
随着预训练文本到图像流匹配模型的兴起,基于文本的图像编辑性能显著提升,尤其在仅含单一编辑目标的简单编辑任务上表现优异。然而,面对日益增长的复杂编辑需求——即包含多个编辑目标的任务——现有方法面临挑战。单轮编辑存在长文本依赖问题,多轮编辑则受累积不一致性的制约,难以兼顾语义对齐与源图像一致性。本文提出FlowDC,将复杂编辑分解为多个子编辑效果,并在编辑过程中并行叠加。同时观察到,与编辑位移方向正交的速度分量会破坏源结构完整性,因此我们对速度进行分解,并衰减该正交部分以增强源一致性。为评估复杂编辑能力,我们构建了复杂编辑基准Complex-PIE-Bench。在两个基准上的实验表明,FlowDC优于现有方法。我们还对模块设计进行了详尽消融分析。
原文摘要 · Abstract (English)
With the surge of pre-trained text-to-image flow matching models, text-based image editing performance has gained remarkable improvement, especially for \underline{simple editing} that only contains a single editing target. To satisfy the exploding editing requirements, the \underline{complex editing} which contains multiple editing targets has posed as a more challenging task. However, current complex editing solutions: single-round and multi-round editing are limited by long text following and cumulative inconsistency, respectively. Thus, they struggle to strike a balance between semantic alignment and source consistency. In this paper, we propose \textbf{FlowDC}, which decouples the complex editing into multiple sub-editing effects and superposes them in parallel during the editing process. Meanwhile, we observed that the velocity quantity that is orthogonal to the editing displacement harms the source structure preserving. Thus, we decompose the velocity and decay the orthogonal part for better source consistency. To evaluate the effectiveness of complex editing settings, we construct a complex editing benchmark: Complex-PIE-Bench. On two benchmarks, FlowDC shows superior results compared with existing methods. We also detail the ablations of our module designs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。