arXiv:2503.16795cs.CV2025-03被引 13

通过精准定位语义实现更精细的图像编辑,背景保留更好。

DCEdit: Dual-Level Controlled Image Editing via Precisely Localized Semantics

  • 用视觉与文本自注意力增强跨注意力图,精确定位需编辑区域。
  • 在特征和潜在空间双层级控制,实现更细粒度编辑效果。
  • 适用于需要高精度修改且保持背景一致的图像编辑场景。

本文提出一种基于扩散模型的新型文本引导图像编辑方法。文本引导图像编辑的核心挑战在于精确识别并修改目标语义区域,而现有方法在此方面表现不足。我们的方法引入精确语义定位策略,利用视觉与文本自注意力机制增强跨注意力图,作为区域提示以提升编辑性能。进一步提出双层级控制机制,在特征与潜在表示层面同时融入区域提示,实现更精细的编辑控制。为全面评估本方法与其它基于DiT的方案,我们构建了RW-800基准数据集,包含高分辨率图像、长描述性文本、真实世界图像及新的文本编辑任务。在主流PIE-Bench和新提出的RW-800基准上的实验结果表明,该方法在保留背景和实现精准编辑方面均表现出优越性能。

原文摘要 · Abstract (English)

This paper presents a novel approach to improving text-guided image editing using diffusion-based models. Text-guided image editing task poses key challenge of precisly locate and edit the target semantic, and previous methods fall shorts in this aspect. Our method introduces a Precise Semantic Localization strategy that leverages visual and textual self-attention to enhance the cross-attention map, which can serve as a regional cues to improve editing performance. Then we propose a Dual-Level Control mechanism for incorporating regional cues at both feature and latent levels, offering fine-grained control for more precise edits. To fully compare our methods with other DiT-based approaches, we construct the RW-800 benchmark, featuring high resolution images, long descriptive texts, real-world images, and a new text editing task. Experimental results on the popular PIE-Bench and RW-800 benchmarks demonstrate the superior performance of our approach in preserving background and providing accurate edits.

图像编辑扩散模型语义定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。