arXiv:2604.02349cs.LGcs.AI2026-04

用数据集内探索提升离线偏好强化学习的效率

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration

论文配图:OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration
图 1 · 摘自论文原文
  • 基于数据集内探索设计高效查询策略
  • 减少查询次数,性能显著优于现有方法
  • 适合需要少样本人类反馈的机器人任务

基于偏好的强化学习(PbRL)可避免复杂的奖励设计,更契合人类意图,在实际应用中前景广阔。然而,获取人类偏好反馈成本高、耗时长,成为制约PbRL发展的主要障碍。本文针对离线PbRL中查询效率低的问题,指出两大原因:探索效率低与奖励函数过优化。为此,提出新算法OPRIDE(Offline PbRL via In-Dataset Exploration),包含两个核心机制:一种基于信息量最大化的合理探索策略,以及用于缓解奖励过优化的折扣调度机制。实验表明,OPRIDE显著优于现有方法,仅用更少查询即达成优异性能,并提供理论效率保障。在多种运动、操作和导航任务上的结果验证了该方法的有效性与通用性。

原文摘要 · Abstract (English)

Preference-based reinforcement learning (PbRL) can help avoid sophisticated reward designs and align better with human intentions, showing great promise in various real-world applications. However, obtaining human feedback for preferences can be expensive and time-consuming, which forms a strong barrier for PbRL. In this work, we address the problem of low query efficiency in offline PbRL, pinpointing two primary reasons: inefficient exploration and overoptimization of learned reward functions. In response to these challenges, we propose a novel algorithm, \textbf{O}ffline \textbf{P}b\textbf{R}L via \textbf{I}n-\textbf{D}ataset \textbf{E}xploration (OPRIDE), designed to enhance the query efficiency of offline PbRL. OPRIDE consists of two key features: a principled exploration strategy that maximizes the informativeness of the queries and a discount scheduling mechanism aimed at mitigating overoptimization of the learned reward functions. Through empirical evaluations, we demonstrate that OPRIDE significantly outperforms prior methods, achieving strong performance with notably fewer queries. Moreover, we provide theoretical guarantees of the algorithm's efficiency. Experimental results across various locomotion, manipulation, and navigation tasks underscore the efficacy and versatility of our approach.

强化学习偏好学习离线学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。