arXiv:2412.01223cs.CVcs.AI2024-12被引 3

PainterNet让图像修复更贴合用户提示,支持灵活编辑习惯。

PainterNet: Adaptive Image Inpainting with Actual-Token Attention and Diverse Mask Control

  • 引入局部提示与注意力控制点,增强模型对局部区域的关注。
  • 在多个指标上优于现有方法,尤其在文本一致性与图像质量上提升显著。
  • 适配主流扩散模型,适合需要精准可控修复的用户使用。

近期扩散模型在图像修复任务中表现优异,能生成高质量、逼真的修复内容。然而,现有方法在图像与文本语义一致性及用户编辑习惯匹配方面仍存在不足。为此,我们提出PainterNet,一个可灵活嵌入各类扩散模型的插件。为使修复内容更贴合用户输入提示,我们设计了局部提示输入、注意力控制点(ACP)以及实际标记注意力损失(ATAL),以增强模型对局部区域的关注。此外,我们重新设计了训练与测试数据集中的掩码生成算法,模拟用户真实掩码操作习惯,并构建了定制化训练数据集PainterData和基准测试集PainterBench。大量实验表明,PainterNet在图像质量、全局与局部文本一致性等关键指标上均超越现有最先进模型。

原文摘要 · Abstract (English)

Recently, diffusion models have exhibited superior performance in the area of image inpainting. Inpainting methods based on diffusion models can usually generate realistic, high-quality image content for masked areas. However, due to the limitations of diffusion models, existing methods typically encounter problems in terms of semantic consistency between images and text, and the editing habits of users. To address these issues, we present PainterNet, a plugin that can be flexibly embedded into various diffusion models. To generate image content in the masked areas that highly aligns with the user input prompt, we proposed local prompt input, Attention Control Points (ACP), and Actual-Token Attention Loss (ATAL) to enhance the model's focus on local areas. Additionally, we redesigned the MASK generation algorithm in training and testing dataset to simulate the user's habit of applying MASK, and introduced a customized new training dataset, PainterData, and a benchmark dataset, PainterBench. Our extensive experimental analysis exhibits that PainterNet surpasses existing state-of-the-art models in key metrics including image quality and global/local text consistency.

图像修复扩散模型注意力机制用户习惯

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。