通过自熵正则化提升扩散模型的生成稳定性和多样性
SEE-DPO: Self Entropy Enhanced Direct Preference Optimization
- 引入自熵正则化机制,增强生成过程中的探索能力
- 有效缓解奖励劫持问题,提升图像质量与多样性
- 适合关注扩散模型训练稳定性与生成效果的研究者
直接偏好优化(DPO)已成功用于对齐大语言模型与人类偏好,并被应用于提升文本到图像扩散模型的质量。然而,基于DPO的方法如SPO、Diffusion-DPO和D3PO在长期训练中极易过拟合与奖励劫持,尤其是在生成模型拟合分布外样本时。为克服这些挑战并稳定扩散模型训练,本文在从人类反馈中强化学习中引入自熵正则化机制。该机制通过鼓励更广泛的探索和更强的鲁棒性,改善DPO训练过程。实验表明,该方法有效缓解奖励劫持,显著提升整个潜在空间中的图像质量与多样性,在关键生成指标上达到领先水平。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has been successfully used to align large language models (LLMs) according to human preferences, and more recently it has also been applied to improving the quality of text-to-image diffusion models. However, DPO-based methods such as SPO, Diffusion-DPO, and D3PO are highly susceptible to overfitting and reward hacking, especially when the generative model is optimized to fit out-of-distribution during prolonged training. To overcome these challenges and stabilize the training of diffusion models, we introduce a self-entropy regularization mechanism in reinforcement learning from human feedback. This enhancement improves DPO training by encouraging broader exploration and greater robustness. Our regularization technique effectively mitigates reward hacking, leading to improved stability and enhanced image quality across the latent space. Extensive experiments demonstrate that integrating human feedback with self-entropy regularization can significantly boost image diversity and specificity, achieving state-of-the-art results on key image generation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。