通过嵌入层增强评估扩散模型鲁棒性,发现文本编码器会干扰评估结果。
Evaluating Robustness in Latent Diffusion Models via Embedding Level Augmentation
- 在嵌入层而非完整架构上测试模型鲁棒性,避免文本编码器干扰
- 提出新数据增强方法,暴露模型对多样提示的脆弱性
- 为Dreambooth微调后的模型设计专用鲁棒性评估流程
潜在扩散模型(LDMs)在图像生成和视频合成等任务中表现优异,但普遍缺乏鲁棒性,这一问题尚未得到充分研究。本文提出多种方法弥补该空白:首先,假设应剥离文本编码器来评估LDMs的鲁棒性,因完整架构会混淆生成器与编码器的问题;其次,设计新型数据增强技术,揭示模型在处理多样化文本提示时的脆弱性;随后,利用Dreambooth对Stable Diffusion 3和Stable Diffusion XL进行多任务微调,融合上述增强方法;最后,构建专门针对Dreambooth微调后模型的鲁棒性评估流水线。
原文摘要 · Abstract (English)
Latent diffusion models (LDMs) achieve state-of-the-art performance across various tasks, including image generation and video synthesis. However, they generally lack robustness, a limitation that remains not fully explored in current research. In this paper, we propose several methods to address this gap. First, we hypothesize that the robustness of LDMs primarily should be measured without their text encoder, because if we take and explore the whole architecture, the problems of image generator and text encoders wll be fused. Second, we introduce novel data augmentation techniques designed to reveal robustness shortcomings in LDMs when processing diverse textual prompts. We then fine-tune Stable Diffusion 3 and Stable Diffusion XL models using Dreambooth, incorporating these proposed augmentation methods across multiple tasks. Finally, we propose a novel evaluation pipeline specifically tailored to assess the robustness of LDMs fine-tuned via Dreambooth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。