让图像生成模型按提示词个性化学习新概念,效果更准更像。
Per-Query Visual Concept Learning
- 针对每个提示词和随机种子定制个性化过程
- 结合自注意力与交叉注意力损失,提升概念相似性
- 适配多种生成模型,适合需要精准图像定制的场景
视觉概念学习(又称文本到图像个性化)是向预训练模型注入新概念的过程,广泛应用于商品摆放、娱乐和个性化设计。本文表明,通过引入一种针对提示词和噪声种子特异的个性化步骤,并采用基于自注意力与交叉注意力的双损失项,可显著提升现有方法性能。我们利用先前设计用于捕捉身份特征的PDM特征,进一步优化个性化语义相似性。在六种不同个性化方法及多种基础文生图模型(包括UNet和DiT架构)上进行评估,结果表明该方法在多个基准上均实现显著改进,甚至优于已有按查询个性化的先进方法。
原文摘要 · Abstract (English)
Visual concept learning, also known as Text-to-image personalization, is the process of teaching new concepts to a pretrained model. This has numerous applications from product placement to entertainment and personalized design. Here we show that many existing methods can be substantially augmented by adding a personalization step that is (1) specific to the prompt and noise seed, and (2) using two loss terms based on the self- and cross- attention, capturing the identity of the personalized concept. Specifically, we leverage PDM features -- previously designed to capture identity -- and show how they can be used to improve personalized semantic similarity. We evaluate the benefit that our method gains on top of six different personalization methods, and several base text-to-image models (both UNet- and DiT-based). We find significant improvements even over previous per-query personalization methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。