提出GEB框架,让强化学习中的探索更主动发现未知区域。
General Exploratory Bonus for Optimistic Exploration in RLHF
- 用参考模型调节奖励,纠正现有方法的保守倾向
- 在多种发散度设置下,显著提升对齐任务性能
- 适用于大语言模型,兼顾理论严谨与实际效果
乐观探索是提升人类反馈强化学习(RLHF)样本效率的核心,但现有探索奖励方法常因KL或α-分歧正则化而无意中偏向参考模型的高概率区域,导致行为保守而非发现不确定区域。本文通过理论分析揭示此偏差根源,并提出通用探索奖励(GEB),其通过参考依赖的奖励调节可证明满足乐观原则。GEB统一了已有启发式奖励作为特例,并自然扩展至整个α-分歧族。实验表明,GEB在多个分歧设置和大语言模型骨干上均持续优于基线,在对齐任务中表现更优。结果证明GEB为RLHF中的乐观探索提供了兼具理论基础与实用价值的解决方案。
原文摘要 · Abstract (English)
Optimistic exploration is central to improving sample efficiency in reinforcement learning with human feedback, yet existing exploratory bonus methods to incentivize exploration often fail to realize optimism. We provide a theoretical analysis showing that current formulations, under KL or $α$-divergence regularization, unintentionally bias exploration toward high-probability regions of the reference model, thereby reinforcing conservative behavior instead of promoting discovery of uncertain regions. To address this pitfall, we introduce the General Exploratory Bonus (GEB), a novel theoretical framework that provably satisfies the optimism principle. GEB counteracts divergence-induced bias via reference-dependent reward regulation and unifies prior heuristic bonuses as special cases, while extending naturally across the full $α$-divergence family. Empirically, GEB consistently outperforms baselines on alignment tasks across multiple divergence settings and large language model backbones. These results demonstrate that GEB offers both a principled and practical solution for optimistic exploration in RLHF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。