用随机探索与实验设计,高效获取人类偏好以学习奖励。
Efficient Preference-Based Reinforcement Learning: Randomized Exploration Meets Experimental Design
- 基于随机探索设计元算法,避免乐观方法的计算瓶颈。
- 在温和假设下实现后悔率与最终迭代的理论保证。
- 批量查询+最优实验设计,减少偏好询问次数,支持并行部署。
我们研究在一般马尔可夫决策过程中的从人类反馈中进行强化学习,其中智能体通过轨迹级偏好比较学习。该场景的核心挑战是设计能选择高信息量偏好查询的算法,以识别潜在奖励函数,同时保证理论性能。我们提出一种基于随机探索的元算法,避免了乐观方法带来的计算复杂性,保持可计算性。在温和的强化学习预言机假设下,建立了后悔率和最后迭代的理论保证。为降低查询复杂度,我们引入并分析了一种改进算法,该算法批量收集轨迹对,并利用最优实验设计选择高信息量的比较查询。批量结构也支持偏好查询的并行化,符合实际部署需求。实验表明,所提方法在较少偏好查询下,性能可媲美基于奖励的强化学习。
原文摘要 · Abstract (English)
We study reinforcement learning from human feedback in general Markov decision processes, where agents learn from trajectory-level preference comparisons. A central challenge in this setting is to design algorithms that select informative preference queries to identify the underlying reward while ensuring theoretical guarantees. We propose a meta-algorithm based on randomized exploration, which avoids the computational challenges associated with optimistic approaches and remains tractable. We establish both regret and last-iterate guarantees under mild reinforcement learning oracle assumptions. To improve query complexity, we introduce and analyze an improved algorithm that collects batches of trajectory pairs and applies optimal experimental design to select informative comparison queries. The batch structure also enables parallelization of preference queries, which is relevant in practical deployment as feedback can be gathered concurrently. Empirical evaluation confirms that the proposed method is competitive with reward-based reinforcement learning while requiring a small number of preference queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。