arXiv:2510.17803cs.CV2025-10SIGGRAPH被引 13

无需训练即可实现高一致性、精准的图像视频编辑

ConsistEdit: Highly Consistent and Precise Training-free Visual Editing

  • 基于MM-DiT架构设计注意力控制机制,分步调节查询/键/值向量
  • 在多轮编辑中保持视觉一致性,支持结构与纹理的精细修改
  • 适用于视频编辑与多区域修改,可渐进调节一致性强度

近期训练自由的注意力控制方法实现了对已有生成模型的灵活高效文本引导编辑。然而,现有方法难以同时保证强编辑力度与源图一致性,这一缺陷在多轮和视频编辑中尤为突出,易导致视觉误差累积。多数方法仅强制全局一致性,限制了对特定属性(如纹理)的独立修改能力,阻碍细粒度编辑。随着从U-Net到MM-DiT的架构转变,生成性能显著提升,并引入了图文模态融合新机制。通过对MM-DiT的深入分析,我们发现其注意力机制的三个关键特性。基于此,提出ConsistEdit——一种专为MM-DiT设计的新注意力控制方法,结合纯视觉注意力控制、掩码引导的预注意力融合及对查询、键、值令牌的差异化操作,实现一致且与提示对齐的编辑。大量实验证明,ConsistEdit在多种图像与视频编辑任务中达到顶尖表现,涵盖结构一致与不一致场景。不同于以往方法,它是首个在所有推理步骤与注意力层上无需手工设计即可完成编辑的方法,显著提升可靠性与一致性,支持鲁棒的多轮与多区域编辑,并可渐进调节结构一致性,实现更精细控制。

原文摘要 · Abstract (English)

Recent advances in training-free attention control methods have enabled flexible and efficient text-guided editing capabilities for existing generation models. However, current approaches struggle to simultaneously deliver strong editing strength while preserving consistency with the source. This limitation becomes particularly critical in multi-round and video editing, where visual errors can accumulate over time. Moreover, most existing methods enforce global consistency, which limits their ability to modify individual attributes such as texture while preserving others, thereby hindering fine-grained editing. Recently, the architectural shift from U-Net to MM-DiT has brought significant improvements in generative performance and introduced a novel mechanism for integrating text and vision modalities. These advancements pave the way for overcoming challenges that previous methods failed to resolve. Through an in-depth analysis of MM-DiT, we identify three key insights into its attention mechanisms. Building on these, we propose ConsistEdit, a novel attention control method specifically tailored for MM-DiT. ConsistEdit incorporates vision-only attention control, mask-guided pre-attention fusion, and differentiated manipulation of the query, key, and value tokens to produce consistent, prompt-aligned edits. Extensive experiments demonstrate that ConsistEdit achieves state-of-the-art performance across a wide range of image and video editing tasks, including both structure-consistent and structure-inconsistent scenarios. Unlike prior methods, it is the first approach to perform editing across all inference steps and attention layers without handcraft, significantly enhancing reliability and consistency, which enables robust multi-round and multi-region editing. Furthermore, it supports progressive adjustment of structural consistency, enabling finer control.

视觉编辑MM-DiT一致性无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。