arXiv:2412.20062cs.CV2024-12被引 2

MADiff通过预测掩码和增强注意力,精准实现时尚图像文本编辑。

MADiff: Text-Guided Fashion Image Editing with Mask Prediction and Attention-Enhanced Diffusion

  • 用轻量UNet融合前景、姿态与大模型提示,精准预测编辑区域掩码。
  • 通过注意力处理器重构噪声图,显著提升时尚图像编辑强度。
  • 专为时尚编辑设计,适合需要细粒度控制的设计师或电商应用。

文本引导图像编辑模型在通用领域取得显著进展,但在时尚领域面临两个挑战:(1)编辑区域定位不准;(2)编辑力度弱。为此,本文提出MADiff模型。首先,设计MaskNet,将前景区域、densepose及大语言模型生成的掩码提示输入轻量级UNet,以精准预测编辑区域掩码。其次,提出注意力增强扩散模型,将噪声图、注意力图与MaskNet输出的掩码送入注意力处理器,生成优化后的噪声图,并融入扩散模型,使生成图像更贴合目标文本提示。为填补时尚图像编辑无基准数据集的空白,构建Fashion-E数据集,包含训练集28390对图像-文本,评估集2639对用于四类时尚任务。大量实验表明,所提方法在掩码预测精度和编辑强度上均显著优于现有最先进方法。

原文摘要 · Abstract (English)

Text-guided image editing model has achieved great success in general domain. However, directly applying these models to the fashion domain may encounter two issues: (1) Inaccurate localization of editing region; (2) Weak editing magnitude. To address these issues, the MADiff model is proposed. Specifically, to more accurately identify editing region, the MaskNet is proposed, in which the foreground region, densepose and mask prompts from large language model are fed into a lightweight UNet to predict the mask for editing region. To strengthen the editing magnitude, the Attention-Enhanced Diffusion Model is proposed, where the noise map, attention map, and the mask from MaskNet are fed into the proposed Attention Processor to produce a refined noise map. By integrating the refined noise map into the diffusion model, the edited image can better align with the target prompt. Given the absence of benchmarks in fashion image editing, we constructed a dataset named Fashion-E, comprising 28390 image-text pairs in the training set, and 2639 image-text pairs for four types of fashion tasks in the evaluation set. Extensive experiments on Fashion-E demonstrate that our proposed method can accurately predict the mask of editing region and significantly enhance editing magnitude in fashion image editing compared to the state-of-the-art methods.

图像编辑时尚生成扩散模型文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。