解决扩散模型生成多样性下降问题,提升人类偏好对齐效果。
Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement Learning
- 通过方向性修正奖励信号,防止模型陷入单一高分输出模式。
- 在新基准DivGenBench上验证,生成多样性提升37.6%,人类偏好得分更高。
- 适合关注生成质量与多样性平衡的研究者和应用开发者。
近期研究通过人类反馈强化学习显著提升了文本到图像扩散模型与人类偏好的对齐。然而,现有方法虽在自动化奖励指标上表现优异,却常引发偏好模式坍塌(PMC)——模型过度收敛于少数高分输出(如单调风格或过度曝光图像),严重损害生成多样性。本文首次量化该现象,提出DivGenBench基准以评估PMC程度。我们发现,坍塌源于对奖励模型内在偏见的过度优化。为此,提出方向解耦对齐(D²-Align)框架:在冻结奖励模型的前提下,学习其嵌入空间中的方向性修正,并将其应用于优化过程,有效防止模式坍塌,维持生成多样性。综合定性与定量评估显示,该方法在人类偏好对齐方面表现更优。
原文摘要 · Abstract (English)
Recent studies have demonstrated significant progress in aligning text-to-image diffusion models with human preference via Reinforcement Learning from Human Feedback. However, while existing methods achieve high scores on automated reward metrics, they often lead to Preference Mode Collapse (PMC)-a specific form of reward hacking where models converge on narrow, high-scoring outputs (e.g., images with monolithic styles or pervasive overexposure), severely degrading generative diversity. In this work, we introduce and quantify this phenomenon, proposing DivGenBench, a novel benchmark designed to measure the extent of PMC. We posit that this collapse is driven by over-optimization along the reward model's inherent biases. Building on this analysis, we propose Directional Decoupling Alignment (D$^2$-Align), a novel framework that mitigates PMC by directionally correcting the reward signal. Specifically, our method first learns a directional correction within the reward model's embedding space while keeping the model frozen. This correction is then applied to the reward signal during the optimization process, preventing the model from collapsing into specific modes and thereby maintaining diversity. Our comprehensive evaluation, combining qualitative analysis with quantitative metrics for both quality and diversity, reveals that D$^2$-Align achieves superior alignment with human preference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。