arXiv:2506.14648cs.ROcs.AI2025-06被引 2

通过智能选样本和引导探索,让人类反馈更省力、机器人学习更快。

SENIOR: Efficient Query Selection and Preference-Guided Exploration in Preference-based Reinforcement Learning

  • 用运动差异筛选易比较的轨迹片段,减少人类标注负担。
  • 在6个仿真和4个真实机器人任务中,反馈效率和收敛速度均领先。
  • 适合需要高效人机交互的复杂机器人控制场景。

基于人类偏好的强化学习(PbRL)通过人类偏好学习奖励模型,避免了复杂的奖励设计,但其反馈效率和样本效率仍不足。本文提出一种名为SENIOR的新方法,通过高效查询选择与偏好引导探索机制,显著提升人类反馈效率并加速策略学习。核心思想包含两方面:(1) 运动差异选择(MDS):基于状态分布的核密度估计,筛选出运动明显且方向不同的轨迹片段对,使人类更易判断偏好;(2) 偏好引导探索(PGE):鼓励代理向高偏好但低访问状态探索,持续获取有价值样本。两者协同显著加快奖励与策略学习进程。实验表明,在来自仿真和真实世界的六个复杂机器人操作任务及四个现实任务中,SENIOR在人类反馈效率和策略收敛速度上优于其他五种现有方法。视频演示见项目主页:https://2025senior.github.io/

原文摘要 · Abstract (English)

Preference-based Reinforcement Learning (PbRL) methods provide a solution to avoid reward engineering by learning reward models based on human preferences. However, poor feedback- and sample- efficiency still remain the problems that hinder the application of PbRL. In this paper, we present a novel efficient query selection and preference-guided exploration method, called SENIOR, which could select the meaningful and easy-to-comparison behavior segment pairs to improve human feedback-efficiency and accelerate policy learning with the designed preference-guided intrinsic rewards. Our key idea is twofold: (1) We designed a Motion-Distinction-based Selection scheme (MDS). It selects segment pairs with apparent motion and different directions through kernel density estimation of states, which is more task-related and easy for human preference labeling; (2) We proposed a novel preference-guided exploration method (PGE). It encourages the exploration towards the states with high preference and low visits and continuously guides the agent achieving the valuable samples. The synergy between the two mechanisms could significantly accelerate the progress of reward and policy learning. Our experiments show that SENIOR outperforms other five existing methods in both human feedback-efficiency and policy convergence speed on six complex robot manipulation tasks from simulation and four real-worlds. Videos can be found on our project website: https://2025senior.github.io/

强化学习人机交互机器人控制偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。