arXiv:2509.19301cs.ROcs.LG2025-09被引 42

用轻量残差修正行为克隆策略,实现高效安全的现实机器人强化学习。

Residual Off-Policy RL for Finetuning Behavior Cloning Policies

  • 以行为克隆策略为基底,用少量离线数据学习每步残差修正
  • 仅需稀疏二元奖励,在高自由度机械臂上提升抓取性能
  • 首次在带灵巧手的人形机器人上成功实现实时强化学习

近期行为克隆(BC)进展使视觉-运动控制策略取得显著成效,但受限于人类示范质量、数据收集的人工成本及离线数据的边际效益递减。相比之下,强化学习(RL)通过与环境自主交互训练智能体,在多个领域表现优异。然而,直接在真实机器人上训练RL策略仍具挑战,主要因样本效率低、安全风险及长时程任务中稀疏奖励的学习困难,尤其对高自由度(DoF)系统而言。本文提出一种结合BC与RL优势的残差学习框架。方法以BC策略为黑箱基础,通过高效离线策略学习轻量级每步残差修正。实验表明,该方法仅需稀疏二元奖励信号,即可有效提升高自由度系统在仿真与真实世界中的操作性能。特别地,据我们所知,这是首个在具备灵巧手的人形机器人上成功实现真实世界强化学习训练的案例。结果在多种视觉任务中达到当前最优水平,为实际部署强化学习提供了可行路径。

原文摘要 · Abstract (English)

Recent advances in behavior cloning (BC) have enabled impressive visuomotor control policies. However, these approaches are limited by the quality of human demonstrations, the manual effort required for data collection, and the diminishing returns from offline data. In comparison, reinforcement learning (RL) trains an agent through autonomous interaction with the environment and has shown remarkable success in various domains. Still, training RL policies directly on real-world robots remains challenging due to sample inefficiency, safety concerns, and the difficulty of learning from sparse rewards for long-horizon tasks, especially for high-degree-of-freedom (DoF) systems. We present a recipe that combines the benefits of BC and RL through a residual learning framework. Our approach leverages BC policies as black-box bases and learns lightweight per-step residual corrections via sample-efficient off-policy RL. We demonstrate that our method requires only sparse binary reward signals and can effectively improve manipulation policies on high-degree-of-freedom (DoF) systems in both simulation and the real world. In particular, we demonstrate, to the best of our knowledge, the first successful real-world RL training on a humanoid robot with dexterous hands. Our results demonstrate state-of-the-art performance in various vision-based tasks, pointing towards a practical pathway for deploying RL in the real world.

强化学习行为克隆残差学习真实机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。