提出统一框架,用多种散度正则化提升人类反馈强化学习效率
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses

- 基于采样原则设计两类新算法,统一处理多种散度正则化
- 理论证明可实现O(log T)后悔值和O(1/T)次优间隙
- 首次为通用f-散度正则化在线RLHF提供性能保证,适合算法研究者
基于人类反馈的强化学习(RLHF)已成为大语言模型后训练的核心技术。尽管现有方法多采用反向KL正则化,近期实证研究已开始探索前向KL、卡方等替代散度作为正则项。然而,对一般f-散度正则化的统一理论理解仍不充分。本文构建了在线RLHF中通用f-散度正则化目标的完整理论框架。不针对每种散度单独处理,而是从函数类整体视角出发,提出两种基于不同采样原理的算法:第一种将经典乐观性原则扩展,引入精心设计的探索奖励;第二种利用在f-散度正则化下最优策略对奖励扰动的敏感性。理论分析表明,两类算法均可实现O(log T)后悔值和O(1/T)次优间隙,证明了其可证明的有效性,据我们所知,这是首次为在线RLHF在通用f-散度正则化下的性能提供理论边界。
原文摘要 · Abstract (English)
Reinforcement Learning from Human Feedback (RLHF) has become a cornerstone technique for post-training large language models. While most existing approaches rely on the reverse KL-regularization, recent empirical studies have begun exploring alternative divergences (e.g., forward KL, chi-squared) as regularizers in RLHF. However, a unified theoretical understanding of general $f$-divergence regularization remains under-explored. To fill this gap, this work develops a comprehensive theoretical framework for online RLHF with a general $f$-divergence regularized objective. Rather than treating each possible divergence function individually, we adopt a holistic perspective across the entire function class and propose two algorithms based on distinct sampling principles. The first extends the classical optimism principle with a carefully designed exploration bonus, while the second introduces a new method that exploits the sensitivity of the optimal policy to reward perturbations under $f$-divergence regularization. Theoretical analysis shows that $O(\log T)$ regret and $O(1/T)$ sub-optimality gap are achievable, establishing provable efficiency of both algorithms and, to the best of our knowledge, the first performance bounds for online RLHF under general $f$-divergence regularization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。