让一致性模型高效实现人类反馈强化学习,提升生成质量与训练稳定性。
ROCM: RLHF on consistency models
- 直接优化奖励,利用一阶梯度替代策略梯度,训练更高效。
- 在多个自动指标和人工评估中表现优于或媲美传统方法。
- 引入分布正则化防止奖励欺骗,增强模型泛化能力,适合高要求生成任务。
扩散模型在图像、音频、视频生成等连续域生成任务中取得突破,但其迭代采样过程导致生成速度慢、训练效率低,结合人类反馈强化学习(RLHF)时因奖励稀疏和时间跨度长而加剧挑战。一致性模型通过单步或多步高效生成缓解此问题,显著降低计算成本。本文提出一种直接奖励优化框架,用于对一致性模型实施RLHF,引入分布正则化以增强训练稳定性和防止奖励欺骗。我们研究了多种f-散度作为正则化策略,在奖励最大化与模型一致性间取得平衡。相比策略梯度方法,本方法采用一阶梯度,更具效率且对超参数不敏感。实验表明,该方法在多个自动评估指标和人工评价中达到竞争性或更优性能。此外,分析显示不同正则化技术能有效提升模型泛化能力,防止过拟合。
原文摘要 · Abstract (English)
Diffusion models have revolutionized generative modeling in continuous domains like image, audio, and video synthesis. However, their iterative sampling process leads to slow generation and inefficient training, challenges that are further exacerbated when incorporating Reinforcement Learning from Human Feedback (RLHF) due to sparse rewards and long time horizons. Consistency models address these issues by enabling single-step or efficient multi-step generation, significantly reducing computational costs. In this work, we propose a direct reward optimization framework for applying RLHF to consistency models, incorporating distributional regularization to enhance training stability and prevent reward hacking. We investigate various $f$-divergences as regularization strategies, striking a balance between reward maximization and model consistency. Unlike policy gradient methods, our approach leverages first-order gradients, making it more efficient and less sensitive to hyperparameter tuning. Empirical results show that our method achieves competitive or superior performance compared to policy gradient based RLHF methods, across various automatic metrics and human evaluation. Additionally, our analysis demonstrates the impact of different regularization techniques in improving model generalization and preventing overfitting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。