arXiv:2604.25872cs.LGcs.AI2026-04被引 1

错误奖励未必有害,有些反而能提升语言模型训练效果。

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient

论文配图:When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
图 1 · 摘自论文原文
  • 按误差对奖励进行分类,分析其对策略优化的影响
  • 部分错误奖励可避免模型陷入平庸输出,提升最终表现
  • 为人类反馈强化学习提供更有效的奖励评估方法

通过强化学习训练语言模型时常依赖不完美的代理奖励,因为精确定义理想行为的真值奖励很少存在。传统评估代理奖励质量的方法(如排序准确率)将错误奖励视为纯粹有害。本文指出,并非所有与真值的偏差都同等有害。通过理论分析策略梯度优化中概率集中于哪些输出,我们根据误差对真值奖励提升的影响,对奖励误差进行分类。分析表明,尽管通常被视为有害,某些奖励误差实际上可能无害甚至有益,能防止策略在真值奖励中等的输出上停滞。基于此理论,提出两个实际应用:一是在人类反馈强化学习(RLHF)中,设计考虑误差危害性的奖励模型评估指标,相比标准排序准确率,这些指标与微调后模型性能的相关性更高,但对奖励模型的稳健评估仍有挑战;二是在可验证奖励场景下提供奖励设计洞见。核心发现是,代理奖励的有效性高度依赖于其与初始策略和学习算法的相互作用。

原文摘要 · Abstract (English)

Training language models via reinforcement learning often relies on imperfect proxy rewards, since ground truth rewards that precisely define the intended behavior are rarely available. Standard metrics for assessing the quality of proxy rewards, such as ranking accuracy, treat incorrect rewards as strictly harmful. In this work, however, we highlight that not all deviations from the ground truth are equal. By theoretically analyzing which outputs attract probability during policy gradient optimization, we categorize reward errors according to their effect on the increase in ground truth reward. The analysis establishes that reward errors, though conventionally viewed as harmful, can also be benign or even beneficial by preventing the policy from stalling around outputs with mediocre ground truth reward. We then present two practical implications of our theory. First, for reinforcement learning from human feedback (RLHF), we develop reward model evaluation metrics that account for the harmfulness of reward errors. Compared to standard ranking accuracy, these metrics typically correlate better with the performance of a language model after RLHF, yet gaps remain in robustly evaluating reward models. Second, we provide insights for reward design in settings with verifiable rewards. A key theme underlying our results is that the effectiveness of a proxy reward function depends heavily on its interaction with the initial policy and learning algorithm.

强化学习语言模型奖励设计策略梯度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。