发现熵正则化在在线强化学习中易导致过优化,不如约束KL的方法稳定。
Failure Modes of Maximum Entropy RLHF
- 将SimPO理论化为最大熵强化学习,解释其参考无监督机制
- 在线场景下熵正则化常引发过优化与不稳定的KL动态,即使学习率保守也难避免
- 适合关注在线偏好学习中稳定性问题的研究者和实践者
本文表明,简单偏好优化(SimPO)可被推导为最大熵强化学习,为其无参考方法提供了理论基础。受SimPO在离线偏好优化中优异表现的启发,我们探究最大熵强化学习在在线RLHF设置下的效果。实验发现,最大熵强化学习在不同模型规模下频繁出现过优化及不稳定的KL动态,部分配置即使使用保守学习率仍持续存在过优化。与保持训练稳定的KL约束方法相比,熵正则化无法可靠防止奖励劫持,在实验中反而与过优化的出现相关。即便在训练稳定的情况下,熵正则化也不是稳定因素。最后,我们讨论了为何SimPO在离线场景成功而最大熵强化学习在在线场景失败的可能原因。研究结果表明,无参考方法在在线与离线偏好学习中面临截然不同的挑战。
原文摘要 · Abstract (English)
In this paper, we show that Simple Preference Optimization (SimPO) can be derived as Maximum Entropy Reinforcement Learning, providing a theoretical foundation for this reference-free method. Motivated by SimPO's strong performance in offline preference optimization, we investigate whether Maximum Entropy RL can achieve similar results in online RLHF settings. Our experiments find that Maximum Entropy RL frequently exhibits overoptimization and unstable KL dynamics across model scales, with overoptimization persisting even at conservative learning rates for some configurations. Unlike KL-constrained methods that maintain stable training, entropy regularization fails to reliably prevent reward hacking and, in our experiments, correlates with the onset of overoptimization rather than guarding against it. Even in configurations where training remains stable, entropy regularization is not the stabilizing factor. Lastly, we discuss possible explanations for why SimPO succeeds in offline settings while Maximum Entropy RL struggles in online scenarios. Our findings suggest that reference-free approaches may face distinct challenges when applied to online versus offline preference learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。