arXiv:2602.21765cs.LGcs.AI2026-02

揭示强化学习对齐中奖励偏移与KL裁剪的泛化误差来源

Generalisation of RLHF under Reward Shift and Clipped KL Regularisation

  • 构建考虑奖励偏移与KL裁剪的理论框架
  • 发现泛化误差来自采样、奖励偏移和裁剪三类误差
  • 给出最优裁剪阈值与数据预算分配建议

大型语言模型的对齐与适应高度依赖人类反馈强化学习(RLHF),但其泛化能力的理论理解仍不充分,尤其是在奖励模型基于早期或混合行为策略的偏好数据训练,而RLHF在当前策略的轨迹上优化时。本文发展了考虑两个关键因素的泛化理论:(1)奖励偏移——奖励模型训练数据来自早期或混合行为策略;(2)裁剪的KL正则化——通过采样估计对数概率比并裁剪以稳定训练,引入误差。我们给出了RLHF的泛化界,表明泛化误差由提示词与轨迹的采样误差、奖励偏移误差以及KL裁剪误差构成。还讨论了两种特殊情况:(1)使用有限空间上的均匀先验初始化参数;(2)将梯度下降视为奥恩斯坦-乌伦贝克过程。理论推导带来两项实际启示:(1)确定最优的KL裁剪阈值;(2)优化提示词、轨迹与偏好数据的资源分配。

原文摘要 · Abstract (English)

Alignment and adaptation in large language models heavily rely on reinforcement learning from human feedback (RLHF); yet, theoretical understanding of its generalisability remains premature, especially when the learned reward could shift, and the KL control is estimated and clipped. To address this issue, we develop generalisation theory for RLHF that explicitly accounts for (1) \emph{reward shift}: reward models are trained on preference data from earlier or mixed behaviour policies while RLHF optimises the current policy on its own rollouts; and (2) \emph{clipped KL regularisation}: the KL regulariser is estimated from sampled log-probability ratios and then clipped for stabilisation, resulting in an error to RLHF. We present generalisation bounds for RLHF, suggesting that the generalisation error stems from a sampling error from prompts and rollouts, a reward shift error, and a KL clipping error. We also discuss special cases of (1) initialising RLHF parameters with a uniform prior over a finite space, and (2) training RLHF by stochastic gradient descent, as an Ornstein-Uhlenbeck process. The theory yields practical implications in (1) optimal KL clipping threshold, and (2) budget allocation in prompts, rollouts, and preference data.

RLHF泛化理论奖励偏移KL裁剪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。