arXiv:2506.18438cs.CV2025-06中稿 · IEEE Transactions …被引 2

无需微调即可精准编辑真实图像,保持物体细节与背景完整

CPAM: Context-Preserving Adaptive Manipulation for Zero-Shot Real Image Editing

  • 通过自适应注意力机制分离控制对象与背景
  • 在多个扩散模型上实现零样本编辑,提升纹理与身份一致性
  • 适合需要快速、高保真图像修改的设计师与研究人员

利用文本描述对自然图像进行编辑仍是文本到图像扩散模型中的重大挑战,尤其在保持生成一致性及处理复杂非刚性物体方面。现有方法常难以保留纹理与身份特征,需大量微调,并在特定区域或对象编辑时影响背景细节。本文提出一种零样本框架——上下文保持自适应操控(CPAM),针对复杂非刚性真实图像编辑问题。我们设计了保持适应模块,调整自注意力机制以独立控制对象与背景,确保形状、纹理与身份不变,同时背景不受干扰。此外,引入局部提取模块,减少跨注意力机制中非目标区域的干扰。结合多种掩码引导策略,实现多样化的图像操作。该方法可无缝集成于SD1.5、SD2.1和SDXL等多种扩散模型,展现良好泛化能力。在新构建的真实图像编辑基准数据集IMBA上的大量实验表明,人类评估者更偏好本方法,显著优于现有先进技术。代码与数据将公开发布。

原文摘要 · Abstract (English)

Editing natural images using textual descriptions in text-to-image diffusion models remains a significant challenge, particularly in achieving consistent generation and handling complex, non-rigid objects. Existing methods often struggle to preserve textures and identity, require extensive fine-tuning, and exhibit limitations in editing specific spatial regions or objects while retaining background details. This paper proposes Context-Preserving Adaptive Manipulation (CPAM), a novel zero-shot framework for complicated, non-rigid real image editing. Specifically, we propose a preservation adaptation module that adjusts self-attention mechanisms to preserve and independently control the object and background effectively. This ensures that the objects' shapes, textures, and identities are maintained while keeping the background undistorted during the editing process using the mask guidance technique. Additionally, we develop a localized extraction module to mitigate the interference with the non-desired modified regions during conditioning in cross-attention mechanisms. We also introduce various mask-guidance strategies to facilitate diverse image manipulation tasks in a simple manner. CPAM can be seamlessly integrated with multiple diffusion backbones, including SD1.5, SD2.1, and SDXL, demonstrating strong generalization across different model architectures. Extensive experiments on our newly constructed Image Manipulation BenchmArk (IMBA), a robust benchmark dataset specifically designed for real image editing, demonstrate that our proposed method is the preferred choice among human raters, outperforming existing state-of-the-art editing techniques. The source code and data will be publicly released at the project page: https://vdkhoi20.github.io/CPAM

图像编辑扩散模型零样本掩码引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。