arXiv:2509.20541cs.RO2025-09被引 1

只在学习停滞时才请求人类反馈,显著减少人力成本。

Selective Progress-Aware Querying for Human-in-the-Loop Reinforcement Learning

  • 根据学习进展决定是否请求反馈,避免无效提问。
  • 仅用一半反馈量达成与持续反馈相当的机器人任务成功率。
  • 适合需要节省人力、追求高效训练的真实机器人场景。

人类反馈可显著加速机器人学习,但在现实场景中反馈成本高且有限。现有基于人类反馈的强化学习方法多假设反馈充足,限制了其在物理机器人上的实用性。本文提出SPARQ,一种基于学习进展的智能查询策略,仅在学习停滞或退化时才请求反馈,从而减少不必要的反馈调用。我们在PyBullet仿真环境中的UR5机械臂抓取积木任务上评估SPARQ,对比无反馈、随机查询和始终查询三种基线。实验表明,SPARQ在仅消耗约一半反馈预算的情况下,达到接近完全反馈的近乎完美的任务成功率;相比随机查询,学习更稳定高效;远优于无反馈训练。结果表明,基于进展的主动选择性反馈策略可使人类在环强化学习更高效、更适用于真实人类投入受限的机器人部署场景。

原文摘要 · Abstract (English)

Human feedback can greatly accelerate robot learning, but in real-world settings, such feedback is costly and limited. Existing human-in-the-loop reinforcement learning (HiL-RL) methods often assume abundant feedback, limiting their practicality for physical robot deployment. In this work, we introduce SPARQ, a progress-aware query policy that requests feedback only when learning stagnates or worsens, thereby reducing unnecessary oracle calls. We evaluate SPARQ on a simulated UR5 cube-picking task in PyBullet, comparing against three baselines: no feedback, random querying, and always querying. Our experiments show that SPARQ achieves near-perfect task success, matching the performance of always querying while consuming about half the feedback budget. It also provides more stable and efficient learning than random querying, and significantly improves over training without feedback. These findings suggest that selective, progress-based query strategies can make HiL-RL more efficient and scalable for robots operating under realistic human effort constraints.

强化学习人机交互机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。