arXiv:2505.09571cs.CV2025-05被引 3

用自注意力机制提升图像编辑效率与质量

Don't Forget your Inverse DDIM for Image Editing

  • 基于逆向DDIM的自注意力引导,高效重构未编辑区域
  • 47名用户测试中全部偏好SAGE,7项量化指标排名第一
  • 适合需要快速高质量图像编辑的研究者和设计师

文本到图像生成领域因扩散模型取得显著进展,但真实图像编辑仍面临计算成本高或重建效果差的问题。本文提出SAGE(自注意力引导图像编辑)技术,基于预训练扩散模型,改进DDIM算法,利用扩散U-Net的自注意力层生成注意力图,构建逆向DDIM过程中的重建目标。该方法无需精确重建整张输入图像,即可高效恢复未编辑区域,直接解决图像编辑核心难题。定量与定性评估表明,SAGE优于其他方法;一项经统计验证的用户研究显示,47名参与者全部更倾向SAGE。在10项量化分析中,SAGE位列第一7次,第二、第三各2次。

原文摘要 · Abstract (English)

The field of text-to-image generation has undergone significant advancements with the introduction of diffusion models. Nevertheless, the challenge of editing real images persists, as most methods are either computationally intensive or produce poor reconstructions. This paper introduces SAGE (Self-Attention Guidance for image Editing) - a novel technique leveraging pre-trained diffusion models for image editing. SAGE builds upon the DDIM algorithm and incorporates a novel guidance mechanism utilizing the self-attention layers of the diffusion U-Net. This mechanism computes a reconstruction objective based on attention maps generated during the inverse DDIM process, enabling efficient reconstruction of unedited regions without the need to precisely reconstruct the entire input image. Thus, SAGE directly addresses the key challenges in image editing. The superiority of SAGE over other methods is demonstrated through quantitative and qualitative evaluations and confirmed by a statistically validated comprehensive user study, in which all 47 surveyed users preferred SAGE over competing methods. Additionally, SAGE ranks as the top-performing method in seven out of 10 quantitative analyses and secures second and third places in the remaining three.

图像编辑扩散模型自注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。