提出在线人类反馈强化学习的后验采样算法,理论证明其收益损失随时间呈根号增长。
Thompson Sampling in Online RLHF with General Function Approximation
- 基于汤普森采样设计无模型后验采样算法,利用贝尔曼可消除维数衡量函数类复杂度。
- 理论证明算法在时间 T 内的累计损失为 O(√T),依赖于任务时长、函数类复杂度等因子。
- 首次建立平方贝尔曼误差的集中不等式,对分析同类算法具有独立参考价值。
强化学习从人类反馈(RLHF)在对齐大语言模型与人类偏好方面取得了显著的实证成功,从理论上研究其统计效率具有重要意义。本文考虑在线 RLHF 设置,其中偏好数据在学习过程中逐步揭示,研究动作值函数近似问题。受汤普森采样启发,我们设计了一种无需模型的后验采样算法,并提供了理论保证。具体地,采用贝尔曼可消除(BE)维度作为函数类的复杂度度量,建立了该算法的 $O(igsqrt{T})$ 误差界,其余乘性因子依赖于任务时长、BE 维度及函数类的对数包络数。进一步,在分析中首次基于最大似然估计的泛化界,建立了平方贝尔曼误差的集中型不等式,这一结果在获得类似埃尔德尔型误差界中起关键作用,可能具有独立研究价值。
原文摘要 · Abstract (English)
Reinforcement learning from human feedback (RLHF) has achieved great empirical success in aligning large language models (LLMs) with human preference, and it is of great importance to study the statistical efficiency of RLHF algorithms from a theoretical perspective. In this work, we consider the online RLHF setting where the preference data is revealed during the learning process and study action value function approximation. We design a model-free posterior sampling algorithm for online RLHF inspired by Thompson sampling and provide its theoretical guarantee. Specifically, we adopt Bellman eluder (BE) dimension as the complexity measure of the function class and establish $O(\sqrt{T})$ regret bound for the proposed algorithm with other multiplicative factor depending on the horizon, BE dimension and the $log$-bracketing number of the function class. Further, in the analysis, we first establish the concentration-type inequality of the squared Bellman error bound based on the maximum likelihood estimator (MLE) generalization bound, which plays the crucial rules in obtaining the eluder-type regret bound and may be of independent interest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。