新方法同时学偏好与奖励,让大模型更稳更准。
Simultaneous Reward Distillation and Preference Learning: Get You a Language Model Who Can Do Both
- 联合建模奖励与偏好,避免噪声标签导致的策略崩溃
- 在Ultrafeedback和TL;DR数据集上超越DPO等方法,平均奖励更高
- 适合需要高鲁棒性、抗干扰能力强的对话系统应用
传统基于强化学习的对齐方法显式最大化来自独立奖励模型的期望奖励。近期的监督对齐方法如直接偏好优化(DPO)绕过该步骤以避免模型漂移和奖励过拟合问题。尽管因简单而流行,但依赖布拉德利-特瑞比对偏好形式的DPO等方法,在面对非确定性或噪声偏好标签(如人类对两个输出评分信心不足)时仍可能导致退化策略。本文提出DRDO(直接奖励蒸馏与策略优化),同时建模奖励与偏好,以避免此类退化。DRDO直接模仿来自理想源的奖励,同时通过新颖的偏好似然形式学习人类偏好。在Ultrafeedback和TL;DR数据集上的实验表明,使用DRDO训练的策略在预期奖励上优于DPO和e-DPO,并且对噪声偏好信号及分布外(OOD)设置更具鲁棒性。
原文摘要 · Abstract (English)
Traditional RLHF-based LLM alignment methods explicitly maximize the expected rewards from a separate reward model. More recent supervised alignment methods like Direct Preference Optimization (DPO) circumvent this phase to avoid problems including model drift and reward overfitting. Although popular due to its simplicity, DPO and similar direct alignment methods which rely heavily on the Bradley-Terry-based pairwise preference formulation can still lead to degenerate policies when challenged by non-deterministic or noisy preference labels, for example human scoring of two candidate outputs with low confidence. This paper introduces DRDO (Direct Reward Distillation and policy-Optimization), which simultaneously models rewards and preferences to avoid such degeneracy. DRDO directly mimics rewards assigned by an oracle while learning human preferences with a novel preference likelihood formulation. Results on the Ultrafeedback and TL;DR datasets demonstrate that DRDO-trained policies surpass methods such as DPO and e-DPO in terms of expected rewards and are more robust, on average, to noisy preference signals as well as out-of-distribution (OOD) settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。