arXiv:2506.07492cs.LGstat.ML2025-06ICML被引 6

提出EXPO框架,无需隐式奖励就能更好优化语言模型偏好。

Explicit Preference Optimization: No Need for an Implicit Reward Model

  • 用显式正则化替代隐式奖励重构偏好优化
  • 实验证明可避免DPO的次优正则化问题
  • 适合追求可解释性与稳定性的模型调优者

大型语言模型的生成结果通常通过强化学习从人类反馈(RLHF)进行偏好微调。由于RLHF需独立训练奖励模型再用于策略更新,过程复杂,研究转向更直接的替代方法。直接偏好优化(DPO)及其变体通过重参数化技巧引入隐式奖励,将学习简化为单一损失函数最小化。然而我们证明,这类方法仍存在次优正则化和反直觉插值行为,源于其依赖的重参数化机制。为此,我们提出显式偏好优化框架EXPO,无需重参数化即可实现显式奖励。通过从头设计直观正则化项,透明规避了关键DPO变体的潜在缺陷,且在理论上满足期望的正则化要求。实验结果验证分析,并展示EXPO的有效性。

原文摘要 · Abstract (English)

The generated responses of large language models (LLMs) are often fine-tuned to human preferences through a process called reinforcement learning from human feedback (RLHF). As RLHF relies on a challenging training sequence, whereby a separate reward model is independently learned and then later applied to LLM policy updates, ongoing research effort has targeted more straightforward alternatives. In this regard, direct preference optimization (DPO) and its many offshoots circumvent the need for a separate reward training step. Instead, through the judicious use of a reparameterization trick that induces an \textit{implicit} reward, DPO and related methods consolidate learning to the minimization of a single loss function. And yet despite demonstrable success in some real-world settings, we prove that DPO-based objectives are nonetheless subject to sub-optimal regularization and counter-intuitive interpolation behaviors, underappreciated artifacts of the reparameterizations upon which they are based. To this end, we introduce an \textit{explicit} preference optimization framework termed EXPO that requires no analogous reparameterization to achieve an implicit reward. Quite differently, we merely posit intuitively-appealing regularization factors from scratch that transparently avoid the potential pitfalls of key DPO variants, provably satisfying regularization desiderata that prior methods do not. Empirical results serve to corroborate our analyses and showcase the efficacy of EXPO.

偏好优化大模型微调可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。