arXiv:2607.20883cs.CV2026-07

让一键图像编辑精准定位修改区域,效果更稳更准。

WhereEdit: Mask-aware Local Latent Editing for One-Step Image Editing

论文配图:WhereEdit: Mask-aware Local Latent Editing for One-Step Image Editing
图 1 · 摘自论文原文
  • 基于模型内部特征自动找可编辑区域,实现局部自适应调节。
  • 在PIE-Bench上优于现有方法,保持高效生成的同时提升编辑质量。
  • 适合需要精确控制修改位置的图像编辑场景。

近期的一键文本到图像(T2I)模型实现了高效的图像生成,为实时图像编辑带来新机遇。然而,现有方法主要依赖文本条件进行语义变换,缺乏对修改位置的显式空间控制。更重要的是,即使引入空间约束,这些方法在目标区域仍难以实现强而稳定的语义修改。本文从空间可控视角重新审视一键图像编辑,识别出两个关键挑战:可编辑区域发现与有效局部语义变换。我们发现现有方法存在全局语义迁移,限制了在单步设置下高密度局部编辑。为此,提出 extbf{WhereEdit}框架,将一键编辑重构为局部自适应编辑。WhereEdit自动从模型内部特征中识别语义相关区域,并应用自适应局部调制,在增强目标区域编辑的同时,保持非目标区域和结构一致性。在PIE-Bench基准上的实验表明,WhereEdit持续优于现有方法,在保持单步生成效率的同时实现更优的编辑质量。额外的区域级监督实验进一步凸显了显式空间推理对高质量一键图像编辑的重要性。

原文摘要 · Abstract (English)

Recent one-step text-to-image (T2I) models enable efficient image synthesis and provide new opportunities for real-time image editing. However, existing one-step editing methods primarily rely on text conditioning for semantic transformation, lacking explicit spatial control over \textit{where} to edit. More importantly, even when spatial constraints are introduced, these methods often struggle to achieve strong and stable semantic modifications within the target regions. In this work, we revisit one-step image editing from a spatially controlled perspective and identify two key challenges: discovering editable regions and achieving effective localized semantic transformation. We reveal that existing methods perform global semantic transport, which limits high-intensity local editing under the one-step setting. To address this issue, we propose \textbf{WhereEdit}, a framework that reformulates one-step editing as localized adaptive editing. WhereEdit automatically identifies semantically relevant regions from internal model features and applies adaptive local modulation to enhance target-region editing while preserving non-target areas and structural consistency. Experiments on the PIE-Bench benchmark demonstrate that WhereEdit consistently outperforms existing one-step image editing methods, achieving superior editing quality while maintaining the efficiency of one-step generation. Additional experiments with region-level supervision further highlight the importance of explicit spatial reasoning for high-quality one-step image editing.

图像编辑局部控制生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。