arXiv:2509.26605cs.AIcs.LG2025-09被引 2

用人类偏好反馈优化专家演示策略,提升机器人训练效率与安全性

Fine-tuning Behavioral Cloning Policies with Preference-Based Reinforcement Learning

  • 先从无奖励演示数据学初始策略,再通过人类偏好在线微调
  • 在MuJoCo环境中,相比纯行为克隆和纯在线强化学习,遗憾值更低
  • 首次给出离线到在线方法的理论分析,适合需高效安全训练的场景

将强化学习应用于机器人、工业和医疗领域面临两大挑战:难以设计精确奖励函数,以及探索过程危险且数据消耗大。本文提出两阶段框架:首先从无奖励的专家演示数据集中学习安全初始策略,然后在线使用人类偏好反馈进行微调。我们首次对这种离线到在线的方法进行了严谨分析,提出了BRIDGE算法,通过不确定性加权目标统一融合两种信号。推导出后悔界随离线演示数量增加而减小,明确建立了离线数据量与在线样本效率的关系。在离散与连续控制的MuJoCo环境中验证了BRIDGE,结果表明其后悔值低于独立的行为克隆和在线偏好强化学习。本工作为设计更高效的交互式智能体奠定了理论基础。

原文摘要 · Abstract (English)

Deploying reinforcement learning (RL) in robotics, industry, and health care is blocked by two obstacles: the difficulty of specifying accurate rewards and the risk of unsafe, data-hungry exploration. We address this by proposing a two-stage framework that first learns a safe initial policy from a reward-free dataset of expert demonstrations, then fine-tunes it online using preference-based human feedback. We provide the first principled analysis of this offline-to-online approach and introduce BRIDGE, a unified algorithm that integrates both signals via an uncertainty-weighted objective. We derive regret bounds that shrink with the number of offline demonstrations, explicitly connecting the quantity of offline data to online sample efficiency. We validate BRIDGE in discrete and continuous control MuJoCo environments, showing it achieves lower regret than both standalone behavioral cloning and online preference-based RL. Our work establishes a theoretical foundation for designing more sample-efficient interactive agents.

强化学习行为克隆人类偏好样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。