arXiv:2512.24288cs.RO2025-12被引 4

用不完美的人工干预加速机器人学习,提升真实场景下的操作效率。

Real-world Reinforcement Learning from Suboptimal Interventions

  • 基于状态感知的拉格朗日约束强化学习,动态评估人类干预可靠性
  • 实验表明,达成90%成功率时间减少50%以上,长任务成功率100%
  • 适合需要人机协作、容忍错误干预的真实机器人训练场景

现实世界强化学习为在线训练精确灵巧的机器人操作策略提供了前景,使机器人能从自身经验中学习并逐步减少人工干预。然而,以往方法通常假设人工干预在全状态空间内均为最优,忽略了即使专家也难以在所有状态下持续提供最优动作或完全避免错误。盲目混合干预数据与机器人收集的数据会继承强化学习的样本低效性,而单纯模仿干预数据则可能最终降低强化学习可达到的性能上限。如何利用可能不理想且带噪声的人工干预来加速学习而不受其限制,仍是开放问题。为此,我们提出SiLRI,一种用于真实世界机器人操作任务的状态式拉格朗日强化学习算法。具体而言,我们将在线操作问题建模为约束强化学习优化问题,其中每个状态的约束边界由人类干预的不确定性决定。随后引入状态式拉格朗日乘子,并通过极小极大优化求解,联合优化策略与拉格朗日乘子以达到鞍点。基于人机共驾遥操作系统,我们在多样化的实际操作任务上进行了实验验证。结果表明,SiLRI有效利用了人类的非最优干预,在至少50%的时间内缩短了达到90%成功率所需时间,相较于最先进的方法HIL-SERL;并在其他强化学习方法难以成功地长时序操作任务中实现了100%成功率。

原文摘要 · Abstract (English)

Real-world reinforcement learning (RL) offers a promising approach to training precise and dexterous robotic manipulation policies in an online manner, enabling robots to learn from their own experience while gradually reducing human labor. However, prior real-world RL methods often assume that human interventions are optimal across the entire state space, overlooking the fact that even expert operators cannot consistently provide optimal actions in all states or completely avoid mistakes. Indiscriminately mixing intervention data with robot-collected data inherits the sample inefficiency of RL, while purely imitating intervention data can ultimately degrade the final performance achievable by RL. The question of how to leverage potentially suboptimal and noisy human interventions to accelerate learning without being constrained by them thus remains open. To address this challenge, we propose SiLRI, a state-wise Lagrangian reinforcement learning algorithm for real-world robot manipulation tasks. Specifically, we formulate the online manipulation problem as a constrained RL optimization, where the constraint bound at each state is determined by the uncertainty of human interventions. We then introduce a state-wise Lagrange multiplier and solve the problem via a min-max optimization, jointly optimizing the policy and the Lagrange multiplier to reach a saddle point. Built upon a human-as-copilot teleoperation system, our algorithm is evaluated through real-world experiments on diverse manipulation tasks. Experimental results show that SiLRI effectively exploits human suboptimal interventions, reducing the time required to reach a 90% success rate by at least 50% compared with the state-of-the-art RL method HIL-SERL, and achieving a 100% success rate on long-horizon manipulation tasks where other RL methods struggle to succeed. Project website: https://silri-rl.github.io/.

强化学习机器人操作人机协作在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。