arXiv:2409.04750cs.CV2024-09被引 3

无需训练,通过条件与掩码引导实现电商图像风格一致生成

Training-Free Style Consistent Image Synthesis with Condition and Mask Guidance in E-Commerce

  • 在QKV层面修改注意力图,不破坏商品主体结构
  • 使用共享键值增强跨注意力相似性,提升风格一致性
  • 适合需要快速部署的电商图像生成场景

在电商领域,生成风格一致的图像是一项常见任务。当前方法多基于扩散模型,已取得优异效果。本文提出在UNet与图像条件融合时,引入QKV(查询/键/值)层级的概念,对自注意力和跨注意力的注意力图进行修改。目标是在不改变电商品牌图像主体结构的前提下,采用无训练方法,通过预设条件引导生成。该方法利用共享键值(KV)增强跨注意力中的相似性,并从注意力图中生成掩码引导,巧妙控制风格一致性的生成过程。实验表明,该方法在实际应用中表现良好。

原文摘要 · Abstract (English)

Generating style-consistent images is a common task in the e-commerce field, and current methods are largely based on diffusion models, which have achieved excellent results. This paper introduces the concept of the QKV (query/key/value) level, referring to modifications in the attention maps (self-attention and cross-attention) when integrating UNet with image conditions. Without disrupting the product's main composition in e-commerce images, we aim to use a train-free method guided by pre-set conditions. This involves using shared KV to enhance similarity in cross-attention and generating mask guidance from the attention map to cleverly direct the generation of style-consistent images. Our method has shown promising results in practical applications.

图像生成扩散模型电商图像无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。