通过协同位置与语义信息,提升复杂形变图像编辑的准确性
The Devil is in Attention Sharing: Improving Complex Non-rigid Image Editing Faithfulness via Attention Synergy
- 动态调节位置与语义特征权重,避免编辑失真
- 在去噪每一步量化所需编辑程度,实现精准控制
- 适合需要高保真形变编辑的视觉生成任务
无需训练的大型扩散模型已实现图像编辑的实用化,但准确执行复杂非刚性编辑(如姿态或形状变化)仍极具挑战。我们发现根本原因在于现有注意力共享机制中的注意力坍塌:位置嵌入或语义特征单一主导视觉内容检索,导致过编辑或欠编辑。为此,我们提出 SynPS,一种协同利用位置嵌入与语义信息的方法,实现高保真非刚性图像编辑。首先提出一种编辑度量,量化每个去噪步骤所需的编辑强度。基于该度量,设计注意力协同管道,动态调节位置嵌入的影响,使 SynPS 在语义修改与保真度之间取得平衡。通过自适应融合位置与语义线索,有效避免过编辑与欠编辑。在公开及新构建基准上的大量实验表明,本方法性能卓越且忠实度更高。
原文摘要 · Abstract (English)
Training-free image editing with large diffusion models has become practical, yet faithfully performing complex non-rigid edits (e.g., pose or shape changes) remains highly challenging. We identify a key underlying cause: attention collapse in existing attention sharing mechanisms, where either positional embeddings or semantic features dominate visual content retrieval, leading to over-editing or under-editing. To address this issue, we introduce SynPS, a method that Synergistically leverages Positional embeddings and Semantic information for faithful non-rigid image editing. We first propose an editing measurement that quantifies the required editing magnitude at each denoising step. Based on this measurement, we design an attention synergy pipeline that dynamically modulates the influence of positional embeddings, enabling SynPS to balance semantic modifications and fidelity preservation. By adaptively integrating positional and semantic cues, SynPS effectively avoids both over- and under-editing. Extensive experiments on public and newly curated benchmarks demonstrate the superior performance and faithfulness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。