arXiv:2601.16403cs.LG2026-01被引 5

揭示RLHF在高维下泛化能力的理论机制

Towards a Theoretical Understanding to the Generalization of RLHF

  • 基于算法稳定性分析,构建线性奖励模型下的泛化理论
  • 在特征覆盖条件下,政策模型泛化误差为O(n⁻¹⁄²)
  • 适用于梯度上升等实际训练方法,支持实证效果

强化学习从人类反馈(RLHF)及其变体已成为对齐大语言模型与人类意图的主流方法。尽管在实践中表现有效,其在高维设置下的理论泛化性质仍不明确。本文在线性奖励模型框架下,通过算法稳定性理论,建立了LLM的RLHF泛化理论。与依赖奖励模型最大似然估计一致性的现有工作不同,本分析采用符合实际的端到端学习框架。具体而言,在关键的“特征覆盖”条件下,政策模型的经验最优解具有O(n⁻¹⁄²)阶的泛化界。该结果可推广至基于梯度的学习算法,如梯度上升(GA)和随机梯度上升(SGA)。因此,我们的结论为经过RLHF后大语言模型的实证泛化性能提供了新的理论依据。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) and its variants have emerged as the dominant approaches for aligning Large Language Models with human intent. While empirically effective, the theoretical generalization properties of these methods in high-dimensional settings remain to be explored. To this end, we build the generalization theory on RLHF of LLMs under the linear reward model, through the framework of algorithmic stability. In contrast to the existing works built upon the consistency of maximum likelihood estimations on reward model, our analysis is presented under an end-to-end learning framework, which is consistent with practice. Concretely, we prove that under a key \textbf{feature coverage} condition, the empirical optima of policy model have a generalization bound of order $\mathcal{O}(n^{-\frac{1}{2}})$. Moreover, the results can be extrapolated to parameters obtained by gradient-based learning algorithms, i.e., Gradient Ascent (GA) and Stochastic Gradient Ascent (SGA). Thus, we argue that our results provide new theoretical evidence for the empirically observed generalization of LLMs after RLHF.

强化学习模型对齐泛化理论大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。