arXiv:2605.15971cs.RO2026-05

用人类实时偏好指导机器人学习,提升安全性和效率。

OHP-RL: Online Human Preference as Guidance in Reinforcement Learning for Robot Manipulation

论文配图:OHP-RL: Online Human Preference as Guidance in Reinforcement Learning for Robot Manipulation
图 1 · 摘自论文原文
  • 通过状态依赖的偏好门控制人类干预时机与强度
  • 在真实机械臂上实现更高成功率、更快收敛和更少人工干预
  • 适合需要人机协作的复杂操作场景

强化学习虽能使机器人自主习得技能,但其在现实部署中受限于低效且不安全的探索过程。人机协同干预提供了一种实用方案,但现有方法通常仅将干预视为辅助训练信号,未能充分挖掘其关于何时及如何引导自主性的丰富信息。人类干预常体现行为选择的相对偏好,而非精确动作模仿。为此,我们提出在线人类偏好引导强化学习(OHP-RL)框架,利用人类干预作为偏好信息指导策略学习。OHP-RL引入状态依赖的偏好门,自适应调节人类干预对策略学习的影响程度与时机。该设计使智能体可在间歇性、不完美的人类反馈下受益,同时保持自主探索与稳定策略优化。我们在Franka机械臂上评估了三个高接触力复杂操作任务。所有任务中,OHP-RL均实现更高成功率、更快收敛速度,并显著降低人类干预频率。此外,学习到的策略在整个训练过程中表现出更稳定且符合人类偏好的行为。

原文摘要 · Abstract (English)

While reinforcement learning (RL) enables robots to acquire skills autonomously, its real-world deployment is severely limited by inefficient and unsafe exploration. Human-in-the-loop interventions offer a practical solution, yet existing methods typically exploit these interventions as auxiliary training signals, without fully capturing the richer information they provide about when and how autonomy should be guided. Human interventions often encode relative preferences over behavior under safety and task constraints, rather than prescribing exact actions to imitate. Motivated by this perspective, we propose Online Human Preference as Guidance in Reinforcement Learning (OHP-RL), a framework that leverages human interventions as preference information to guide policy learning. OHP-RL introduces a state-dependent preference gate that adaptively regulates when and to what extent human interventions should shape policy learning. This design enables the agent to benefit from intermittent and imperfect human feedback while preserving autonomous exploration and stable policy optimization. We evaluate OHP-RL on three challenging real-world contact-rich manipulation tasks on a Franka robot. Across all tasks, OHP-RL consistently achieves strong success rates, faster convergence, and substantially lower human intervention effort than prior approaches. Moreover, the learned policies exhibit more stable and human-aligned behavior throughout training.

强化学习人机协作机器人操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。