arXiv:2411.04625cs.LGstat.ML2024-11NeurIPS被引 30

首次证明KL正则化可将样本复杂度从1/ε²降至1/ε,显著提升强化学习效率。

Sharp Analysis for KL-Regularized Contextual Bandits and RLHF

  • 提出尖锐理论分析,揭示KL正则化在上下文赌博机与RLHF中的优势
  • 发现当ε足够小时,样本复杂度降为1/ε,远优于传统方法的1/ε²
  • 证明参考策略充分覆盖时,混合采样策略仅需加性依赖覆盖率

反Kullback-Leibler(KL)正则化已成为强化学习(RL)与基于人类反馈的强化学习(RLHF)中提升策略优化的关键技术,迫使学习策略贴近参考策略。尽管其有效性已在实践中得到验证,但现有理论分析仍显示与无正则化问题相同的$/mathcal{O}(1 / ε^2)$样本复杂度。本文首次通过尖锐分析揭示了KL正则化的根本优势:在上下文赌博机和RLHF中,当ε足够小时,样本复杂度可降至$/mathcal{O}(1 / ε)$。进一步研究数据覆盖的作用——在离线RLHF中,覆盖假设常被用于连接参考策略与最优策略,但通常带来覆盖率系数的乘性依赖;而在线RLHF中其影响尚不明确。以往分析通常要求显式探索或对奖励函数类施加额外结构假设。本文表明,在参考策略具备充分覆盖的前提下,简单的两阶段混合采样策略即可实现仅含覆盖率系数加性依赖的样本复杂度。结果为理解KL正则化与数据覆盖在RLHF中的作用提供了全面视角,指导更高效算法的设计。

原文摘要 · Abstract (English)

Reverse-Kullback-Leibler (KL) regularization has emerged to be a predominant technique used to enhance policy optimization in reinforcement learning (RL) and reinforcement learning from human feedback (RLHF), which forces the learned policy to stay close to a reference policy. While the effectiveness and necessity of KL-regularization have been empirically demonstrated in various practical scenarios, current theoretical analysis of KL-regularized RLHF still obtains the same $\mathcal{O}(1 / ε^2)$ sample complexity as problems without KL-regularization. To understand the fundamental distinction between policy learning objectives with KL-regularization and ones without KL-regularization, we are the first to theoretically demonstrate the power of KL-regularization by providing a sharp analysis for KL-regularized contextual bandits and RLHF, revealing an $\mathcal{O}(1 / ε)$ sample complexity when $ε$ is sufficiently small. We further explore the role of data coverage in contextual bandits and RLHF. While the coverage assumption is commonly employed in offline RLHF to link the samples from the reference policy to the optimal policy, often at the cost of a multiplicative dependence on the coverage coefficient, its impact on the sample complexity of online RLHF remains unclear. Previous theoretical analyses of online RLHF typically require explicit exploration and additional structural assumptions on the reward function class. In contrast, we show that with sufficient coverage from the reference policy, a simple two-stage mixed sampling strategy can achieve a sample complexity with only an additive dependence on the coverage coefficient. Our results provide a comprehensive understanding of the roles of KL-regularization and data coverage in RLHF, shedding light on the design of more efficient RLHF algorithms.

强化学习KL正则样本复杂度数据覆盖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。