arXiv:2409.08272cs.CVcs.GR2024-09AAAI被引 3

只用一个点击点就能动态生成掩码,实现精准局部图像编辑。

Click2Mask: Local Editing with Dynamic Mask Generation

  • 仅需一个点击点,动态生成目标区域掩码。
  • 在多种评测中表现优于或媲美顶尖方法。
  • 适合希望快速编辑图像的普通用户使用。

生成模型的进步已让图像生成与编辑变得对非专业人士也易于操作。本文聚焦于局部图像编辑,特别是向粗略指定区域添加新内容的任务。现有方法通常需要精确掩码或详细位置描述,过程繁琐且易出错。我们提出 Click2Mask,通过仅需一个参考点(外加内容描述)即可简化流程。在融合潜空间扩散(BLD)过程中,掩码会动态围绕该点生长,并由基于掩码的 CLIP 语义损失引导。Click2Mask 克服了依赖分割和微调方法的局限,提供更友好且上下文准确的解决方案。实验表明,Click2Mask 不仅大幅降低用户操作成本,还在人工评价与自动指标上达到或超越当前最优水平。关键贡献包括简化用户输入、自由添加不受现有分割约束的对象,以及动态掩码方法可集成至其他编辑框架。

原文摘要 · Abstract (English)

Recent advancements in generative models have revolutionized image generation and editing, making these tasks accessible to non-experts. This paper focuses on local image editing, particularly the task of adding new content to a loosely specified area. Existing methods often require a precise mask or a detailed description of the location, which can be cumbersome and prone to errors. We propose Click2Mask, a novel approach that simplifies the local editing process by requiring only a single point of reference (in addition to the content description). A mask is dynamically grown around this point during a Blended Latent Diffusion (BLD) process, guided by a masked CLIP-based semantic loss. Click2Mask surpasses the limitations of segmentation-based and fine-tuning dependent methods, offering a more user-friendly and contextually accurate solution. Our experiments demonstrate that Click2Mask not only minimizes user effort but also enables competitive or superior local image manipulations compared to SoTA methods, according to both human judgement and automatic metrics. Key contributions include the simplification of user input, the ability to freely add objects unconstrained by existing segments, and the integration potential of our dynamic mask approach within other editing methods.

图像编辑扩散模型动态掩码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。