arXiv:2606.19328cs.LGcs.AI2026-06

用不确定性平衡规划提升偏好强化学习采样效率

UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning

论文配图:UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning
图 1 · 摘自论文原文
  • 通过联合评估奖励、动态和价值函数的不确定性,主动规划探索路径
  • 在Meta-World上样本效率显著优于无模型和非乐观基线方法
  • 理论保证亚线性遗憾,无需人工设计探索策略

基于偏好的强化学习通过行为的成对比较学习奖励模型,避免了显式奖励设计。然而现有方法通常依赖被动数据收集,在学习初期样本效率低。我们提出一种基于模型的方法,通过联合推理奖励、动态和价值函数中的不确定性,主动引导探索。所提方法UBP2利用奖励、动态和价值函数的集成模型,根据统一得分(包含期望奖励、终端价值和认知不确定性)评估候选轨迹。该目标下的规划自然实现利用与信息获取之间的权衡,无需人为探索启发式。在标准正则性假设下,我们建立了有限和无限时域设置下的亚线性遗憾保证。实验在Meta-World基准上显示,UBP2在样本效率上显著优于无模型偏好方法和非乐观模型基线。

原文摘要 · Abstract (English)

Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing the need for explicit reward design. However, existing methods typically rely on passive data collection and suffer from poor sample efficiency, especially during the early stages of learning. We introduce a model-based approach that actively directs exploration by jointly reasoning over uncertainties in the reward, dynamics, and value functions. Our method, Uncertainty-Balanced Preference Planning (UBP2), uses ensembles of reward, dynamics, and value function models to evaluate candidate trajectories according to a unified score that combines expected reward, terminal value, and epistemic uncertainty. Planning under this objective yields an explicit tradeoff between exploitation and information acquisition without requiring ad hoc exploration heuristics. Under standard regularity assumptions, we establish sublinear regret guarantees for both finite-horizon and infinite-horizon settings. Empirically, experiments on the Meta-World benchmark show UBP2 achieves substantially higher sample efficiency than model-free preference-based methods and non-optimistic model-based baselines.

强化学习偏好学习高效采样不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。