arXiv:2505.05101cs.CV2025-05被引 5

无需训练,精准编辑复杂图像中的多个对象。

MDE-Edit: Masked Dual-Editing for Multi-Object Image Editing via Diffusion Models

  • 通过双损失优化扩散模型噪声特征,实现精确定位。
  • 在复杂场景下编辑准确率显著优于现有方法。
  • 适合需要高精度多对象编辑的设计师与研究者。

多对象编辑旨在修改复杂场景中多个对象或区域的同时保持结构一致性。该任务在重叠或交互对象场景中面临两大挑战:(1)注意力错位导致目标对象定位不准,引发编辑不完整或偏移;(2)属性-对象错配,颜色或纹理变化因跨注意力泄漏未能对齐目标区域,造成语义冲突(如颜色渗入非目标区域)。现有方法受限于全局交叉注意力机制带来的注意力稀释与空间干扰,或掩码方法在多对象场景中因特征纠缠无法精确绑定属性与几何区域。为此,我们提出一种无需训练、仅在推理阶段优化的MDE-Edit方法,通过两个关键损失函数——对象对齐损失(OAL)使多层交叉注意力与分割掩码对齐以精确定位对象,颜色一致性损失(CCL)在掩码内增强目标属性注意力并抑制向邻近区域的泄漏。该双损失设计确保了局部化且连贯的多对象编辑。大量实验表明,MDE-Edit在编辑准确率和视觉质量上均优于现有最优方法,为复杂多对象图像操作提供了稳健解决方案。

原文摘要 · Abstract (English)

Multi-object editing aims to modify multiple objects or regions in complex scenes while preserving structural coherence. This task faces significant challenges in scenarios involving overlapping or interacting objects: (1) Inaccurate localization of target objects due to attention misalignment, leading to incomplete or misplaced edits; (2) Attribute-object mismatch, where color or texture changes fail to align with intended regions due to cross-attention leakage, creating semantic conflicts (\textit{e.g.}, color bleeding into non-target areas). Existing methods struggle with these challenges: approaches relying on global cross-attention mechanisms suffer from attention dilution and spatial interference between objects, while mask-based methods fail to bind attributes to geometrically accurate regions due to feature entanglement in multi-object scenarios. To address these limitations, we propose a training-free, inference-stage optimization approach that enables precise localized image manipulation in complex multi-object scenes, named MDE-Edit. MDE-Edit optimizes the noise latent feature in diffusion models via two key losses: Object Alignment Loss (OAL) aligns multi-layer cross-attention with segmentation masks for precise object positioning, and Color Consistency Loss (CCL) amplifies target attribute attention within masks while suppressing leakage to adjacent regions. This dual-loss design ensures localized and coherent multi-object edits. Extensive experiments demonstrate that MDE-Edit outperforms state-of-the-art methods in editing accuracy and visual quality, offering a robust solution for complex multi-object image manipulation tasks.

图像编辑扩散模型多对象掩码优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。