arXiv:2606.14125cs.CVcs.AI2026-06中稿 · KDD

改进文本条件控制,让扩散模型图像编辑更稳定精准

Conditioning Matters: Stabilizing Inversion and Attention in Diffusion Image Editing

论文配图:Conditioning Matters: Stabilizing Inversion and Attention in Diffusion Image Editing
图 1 · 摘自论文原文
  • 通过优化文本条件信号提升反演稳定性
  • 在PIE-Bench上实现更高重建质量与编辑保真度
  • 适合需要精确控制编辑效果的研究者

基于反演的图像编辑无需训练即可灵活控制,但仍面临反演精度低和编辑保真度与背景保留之间的权衡问题。尽管近期方法改进了反演形式或注意力交互,但文本条件对扩散动态和编辑行为的影响仍被忽视。我们通过实证与理论证明,文本条件的精确性通过调节扩散速度场的几何结构,影响反演稳定性及跨分支注意力的一致性,进而决定背景保留与语义保真度。基于此分析,我们提出SimEdit,一个条件感知框架,包含:(a) 条件细化,构建语义更精确、结构对齐更好的条件信号以促进稳定反演与一致注意力操作;(b) 词级跨分支注意力控制,分离编辑相关与结构保留成分,并在注意力操作中异步调节。在PIE-Bench上的大量实验表明,SimEdit在反演重建质量和编辑性能上均优于现有注意力调控方法。代码已公开于https://github.com/zju-pi/SimEdit。

原文摘要 · Abstract (English)

Inversion-based image editing offers flexible and training-free control but still struggles with inversion accuracy and the trade-off between editing fidelity and background preservation. While recent methods improve inversion formulations or attention interactions, the role of textual conditioning in shaping diffusion dynamics and editing behavior remains underexplored. We show both empirically and theoretically that the precision of textual conditioning influences inversion stability by modulating the geometry of the diffusion velocity field, while also affecting the consistency of cross-branch attention during editing. These effects directly impact background preservation and semantic fidelity. Building on this analysis, we propose SimEdit, a conditioning-aware framework with two complementary components: (a) conditioning refinement, which constructs conditioning signals with improved semantic precision and structural alignment to facilitate stable inversion and consistent attention manipulation, and (b) token-wise cross-branch attention control, which separates edit-relevant and structure-preserving components and modulates them asymmetrically during attention manipulation. Extensive experiments on PIE-Bench demonstrate that SimEdit consistently improves both inversion reconstruction quality and editing performance over previous attention-manipulation approaches. Our code is available at https://github.com/zju-pi/SimEdit.

图像编辑扩散模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。