arXiv:2504.06020cs.AIcs.CL2025-04NeurIPS被引 7

分离提示与响应的奖励,提升模型对新输入的泛化能力

Information-Theoretic Reward Decomposition for Generalizable RLHF

  • 从信息论视角分解奖励为无提示和相关提示两部分
  • 在标准测试中提升对未见数据的对齐与泛化表现
  • 无需额外模型,适合需高鲁棒性的强化学习场景

通用奖励模型在人类反馈强化学习(RLHF)中至关重要,因其能正确评估未见过的提示-响应对。然而,现有奖励模型通常仅通过增大优选与劣选响应之间的奖励差距进行训练,忽视了响应所依赖的提示。这导致当模型评估分布外的提示-响应对时,因忽略提示影响而泛化能力差。为此,本文将奖励值分解为两个独立分量:仅由响应决定的无提示奖励,以及同时依赖提示与响应的提示相关奖励。该分解基于信息论视角,无需额外模型。随后提出一种新奖励学习算法,优先选择无提示奖励值高的样本。通过简单示例验证,提取的两部分奖励能有效刻画奖励模型的不同成分。标准评估显示,该方法显著提升奖励模型的对齐性能与泛化能力。

原文摘要 · Abstract (English)

A generalizable reward model is crucial in Reinforcement Learning from Human Feedback (RLHF) as it enables correctly evaluating unseen prompt-response pairs. However, existing reward models lack this ability, as they are typically trained by increasing the reward gap between chosen and rejected responses, while overlooking the prompts that the responses are conditioned on. Consequently, when the trained reward model is evaluated on prompt-response pairs that lie outside the data distribution, neglecting the effect of prompts may result in poor generalization of the reward model. To address this issue, we decompose the reward value into two independent components: prompt-free reward and prompt-related reward. Prompt-free reward represents the evaluation that is determined only by responses, while the prompt-related reward reflects the reward that derives from both the prompt and the response. We extract these two components from an information-theoretic perspective, which requires no extra models. Subsequently, we propose a new reward learning algorithm by prioritizing data samples based on their prompt-free reward values. Through toy examples, we demonstrate that the extracted prompt-free and prompt-related rewards effectively characterize two parts of the reward model. Further, standard evaluations show that our method improves both the alignment performance and the generalization capability of the reward model.

强化学习奖励建模泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。