通过动态调整生成策略,提升文本到图像模型的奖励微调效率
DiffExp: Efficient Exploration in Reward Fine-tuning for Text-to-Image Diffusion Models
- 动态调节无分类器引导强度以增加图像多样性
- 随机加权文本提示词,挖掘更强奖励信号
- 显著提升在线生成时的探索效率,适合图像生成优化场景
将文本到图像扩散模型进行奖励微调已被证明能有效提升性能,但现有方法因在线样本生成导致收敛缓慢。因此,获取多样且高奖励信号的样本对提升样本效率和整体表现至关重要。本文提出DiffExp,一种简单而高效的奖励微调探索策略。该方法采用两项关键机制:(a) 动态调整无分类器引导的尺度以增强样本多样性;(b) 随机加权文本提示中的短语,以挖掘高质量奖励信号。实验证明,这些策略显著提升了在线生成过程中的探索能力,改进了近期如DDPO和AlignProp等方法的样本效率。
原文摘要 · Abstract (English)
Fine-tuning text-to-image diffusion models to maximize rewards has proven effective for enhancing model performance. However, reward fine-tuning methods often suffer from slow convergence due to online sample generation. Therefore, obtaining diverse samples with strong reward signals is crucial for improving sample efficiency and overall performance. In this work, we introduce DiffExp, a simple yet effective exploration strategy for reward fine-tuning of text-to-image models. Our approach employs two key strategies: (a) dynamically adjusting the scale of classifier-free guidance to enhance sample diversity, and (b) randomly weighting phrases of the text prompt to exploit high-quality reward signals. We demonstrate that these strategies significantly enhance exploration during online sample generation, improving the sample efficiency of recent reward fine-tuning methods, such as DDPO and AlignProp.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。