支持多轮迭代的图像编辑框架,解决连续修改一致性难题。
Multi-turn Consistent Image Editing
- 用流匹配实现精准图像逆向重建,减少误差积累。
- 采用双目标LQR稳定采样,提升多轮编辑视觉质量。
- 通过自适应注意力增强可编辑性,适合需要反复调整的创作场景。
许多实际应用如交互式修图、艺术内容生成和产品设计,需要灵活且迭代的图像编辑能力。然而,现有方法多聚焦单步修改,常因用户意图模糊、变换复杂或需逐步优化而表现不佳,导致结果不一致或不符合预期。为此,我们提出一种多轮图像编辑框架,支持用户持续迭代修正,逐步获得更满意的成果。该方法利用流匹配实现精准图像逆向重建,并采用双目标线性二次调节器(LQR)进行稳定采样,有效缓解误差累积问题。此外,通过分析Transformer各层作用,引入自适应注意力突出方法,在保持多轮一致性的同时增强可编辑性。大量实验表明,相比现有方法,本框架显著提升了编辑成功率与视觉保真度。
原文摘要 · Abstract (English)
Many real-world applications, such as interactive photo retouching, artistic content creation, and product design, require flexible and iterative image editing. However, existing image editing methods primarily focus on achieving the desired modifications in a single step, which often struggles with ambiguous user intent, complex transformations, or the need for progressive refinements. As a result, these methods frequently produce inconsistent outcomes or fail to meet user expectations. To address these challenges, we propose a multi-turn image editing framework that enables users to iteratively refine their edits, progressively achieving more satisfactory results. Our approach leverages flow matching for accurate image inversion and a dual-objective Linear Quadratic Regulators (LQR) for stable sampling, effectively mitigating error accumulation. Additionally, by analyzing the layer-wise roles of transformers, we introduce a adaptive attention highlighting method that enhances editability while preserving multi-turn coherence. Extensive experiments demonstrate that our framework significantly improves edit success rates and visual fidelity compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。