解决生成模型微调中的多样性崩溃问题,提升图像生成的多样性与任务对齐能力。
Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation
- 通过采样、提示和优化三方面设计,持续激励生成多样性。
- 在相同任务对齐下,多样性提升9.08%~43.46%,在相同多样性下对齐提升59.65%~65.86%。
- 适合需要多样化输出的复杂图像生成场景,如创意设计与多方案生成。
强化学习已成为微调大规模生成模型(如扩散模型、流模型)以对齐复杂人类偏好和用户任务的强大范式。然而,其核心瓶颈在于‘多样性崩溃’——目标函数与优化空间会自然导致策略坍缩为狄拉克δ分布。为此,我们提出DRIFT(Diversity-Incentivized Reinforcement Fine-Tuning),一种系统性激励生成多样性的在线策略微调框架,实现强任务对齐与高生成多样性之间的平衡,提升图像生成的泛化能力。从三个视角出发:1)采样上筛选奖励集中子集,过滤异常值以防止过早坍缩;2)通过随机变化的提示扩展条件空间;3)采用基于势能的奖励塑造机制优化组内多样性。实验表明,DRIFT在任务对齐与生成多样性上均达到更优帕累托前沿,在等对齐水平下多样性提升9.08%~43.46%,在等多样性水平下对齐提升59.65%~65.86%。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks. A fundamental limitation remains \textit{the curse of diversity collapse}, where the objective formulation and optimization landscape inherently collapse the policy to a Dirac delta distribution. To address this challenge, we propose \textbf{DRIFT} (\textbf{D}ive\textbf{R}sity-\textbf{I}ncentivized Reinforcement \textbf{F}ine-\textbf{T}uning for Versatile Image Generation), an innovative framework that systematically incentivizes output diversity throughout the on-policy fine-tuning process, reconciling strong task alignment with high generation diversity to enhance versatility essential for applications that demand diverse candidate generations. We approach the problem across three representative perspectives: i) \textbf{sampling} a reward-concentrated subset that filters out reward outliers to prevent premature collapse; ii) \textbf{prompting} with stochastic variations to expand the conditioning space, and iii) \textbf{optimization} of the intra-group diversity with a potential-based reward shaping mechanism. Experimental results show that DRIFT achieves superior Pareto dominance regarding task alignment and generation diversity, yielding a $ 9.08\%\!\sim\! 43.46\%$ increase in diversity at equivalent alignment levels and a $ 59.65\% \!\sim\! 65.86\%$ increase in alignment at equivalent levels of diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。