通过校准扩散路径,实现更精准的图像去物操作。
Erase Diffusion: Empowering Object Removal Through Calibrating Diffusion Pathways
- 设计链式修正优化机制,模拟物体逐步消失过程
- 在OpenImages V5上达到当前最佳性能,真实场景效果显著
- 适合需要高精度去物的图像编辑与修复任务
去物修复旨在精确移除掩码区域内的目标物体,同时保持周围内容的整体一致性。尽管基于扩散的方法在图像修复领域取得显著进展,但仍存在意外物体或伪影出现的问题。我们认为,现有标准优化范式建立的扩散路径不准确,限制了去物效果。为此,提出新型去物扩散方法EraDiff,通过改进优化范式与网络结构,提升结果的连贯性与消除效果。首先引入链式修正优化(CRO)机制,设计专用于去物目标的扩散过渡路径,模拟物体在优化中逐步消除的过程,使模型更准确捕捉去物意图。此外,为缓解采样路径中伪影带来的偏差,提出简单有效的自修正注意力(SRA)机制,通过调整自注意力激活来校准采样路径,有效避开伪影并增强生成内容连贯性。该方法在OpenImages V5数据集上达到当前最优性能,并在真实场景中表现突出。
原文摘要 · Abstract (English)
Erase inpainting, or object removal, aims to precisely remove target objects within masked regions while preserving the overall consistency of the surrounding content. Despite diffusion-based methods have made significant strides in the field of image inpainting, challenges remain regarding the emergence of unexpected objects or artifacts. We assert that the inexact diffusion pathways established by existing standard optimization paradigms constrain the efficacy of object removal. To tackle these challenges, we propose a novel Erase Diffusion, termed EraDiff, aimed at unleashing the potential power of standard diffusion in the context of object removal. In contrast to standard diffusion, the EraDiff adapts both the optimization paradigm and the network to improve the coherence and elimination of the erasure results. We first introduce a Chain-Rectifying Optimization (CRO) paradigm, a sophisticated diffusion process specifically designed to align with the objectives of erasure. This paradigm establishes innovative diffusion transition pathways that simulate the gradual elimination of objects during optimization, allowing the model to accurately capture the intent of object removal. Furthermore, to mitigate deviations caused by artifacts during the sampling pathways, we develop a simple yet effective Self-Rectifying Attention (SRA) mechanism. The SRA calibrates the sampling pathways by altering self-attention activation, allowing the model to effectively bypass artifacts while further enhancing the coherence of the generated content. With this design, our proposed EraDiff achieves state-of-the-art performance on the OpenImages V5 dataset and demonstrates significant superiority in real-world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。