用文本控制图像中多个连续属性,无需配对数据即可精准调节。
Att-Adapter: A Robust and Precise Domain-Specific Multi-Attributes T2I Diffusion Adapter via Conditional Variational Autoencoder
- 通过条件变分自编码器学习单个适配器,实现多属性联合控制。
- 在两个公开数据集上超越所有LoRA基线,控制范围更广且属性解耦更好。
- 无需合成配对数据,可灵活扩展至多个属性,适合快速定制化生成。
文本到图像扩散模型在生成高质量图像方面取得了显著进展。然而,在仅使用文本引导的情况下,对新领域(如眼睛张开度或汽车宽度等数值属性)的连续属性进行精确控制,尤其是同时控制多个属性,仍是一大挑战。为此,我们提出属性适配器(Att-Adapter),一个可即插即用的模块,用于在预训练扩散模型中实现细粒度、多属性控制。该方法从一组无配对的样本图像中学习单一控制适配器,这些图像包含多种视觉属性。Att-Adapter利用解耦的交叉注意力模块,自然融合多属性与文本条件。我们进一步引入条件变分自编码器(CVAE)以缓解过拟合,适应视觉世界的多样性。在两个公开数据集上的评估表明,Att-Adapter在控制连续属性方面优于所有基于LoRA的基线方法。此外,该方法实现了更广的控制范围,并提升了多属性间的解耦性,超越了基于StyleGAN的技术。值得注意的是,Att-Adapter具有灵活性,训练无需配对的合成数据,且可轻松扩展至单模型中的多个属性。
原文摘要 · Abstract (English)
Text-to-Image (T2I) Diffusion Models have achieved remarkable performance in generating high quality images. However, enabling precise control of continuous attributes, especially multiple attributes simultaneously, in a new domain (e.g., numeric values like eye openness or car width) with text-only guidance remains a significant challenge. To address this, we introduce the Attribute (Att) Adapter, a novel plug-and-play module designed to enable fine-grained, multi-attributes control in pretrained diffusion models. Our approach learns a single control adapter from a set of sample images that can be unpaired and contain multiple visual attributes. The Att-Adapter leverages the decoupled cross attention module to naturally harmonize the multiple domain attributes with text conditioning. We further introduce Conditional Variational Autoencoder (CVAE) to the Att-Adapter to mitigate overfitting, matching the diverse nature of the visual world. Evaluations on two public datasets show that Att-Adapter outperforms all LoRA-based baselines in controlling continuous attributes. Additionally, our method enables a broader control range and also improves disentanglement across multiple attributes, surpassing StyleGAN-based techniques. Notably, Att-Adapter is flexible, requiring no paired synthetic data for training, and is easily scalable to multiple attributes within a single model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。