arXiv:2508.12252cs.RO2025-08中稿 · CoRL被引 12

用机械臂指导人形机器人在真实世界中自主学习,减少人工干预。

Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids

  • 机械臂作为教师,实时提供保护与训练反馈
  • 实测实现人形机器人从零开始学翻滚动作,步行速度追踪误差低于5%
  • 适合希望降低真实环境训练成本的研究者与工程师

基于仿真的强化学习已显著提升人形机器人行走任务表现,但直接在真实世界中从头训练或迁移预训练策略仍极为罕见,限制了人形机器人的潜力。真实世界学习虽能克服仿真到现实的差距,却面临安全、奖励设计和学习效率等挑战。为此,我们提出机器人教机器人(RTR)框架,由机械臂教师主动支持并引导人形机器人学生。RTR系统提供保护、学习调度、奖励信号、扰动、故障检测及自动重置功能,实现低人工干预下的高效长期真实世界训练。此外,我们提出一种新型强化学习流程,通过在真实世界中优化单一动态编码的潜在变量,稳定并促进仿真到现实的迁移。我们在两项挑战性真实任务中验证该方法:精确速度追踪的步行策略微调,以及从零开始学习人形机器人摆起动作,展示了RTR式系统在真实世界人形机器人学习中的巨大潜力。

原文摘要 · Abstract (English)

Simulation-based reinforcement learning (RL) has significantly advanced humanoid locomotion tasks, yet direct real-world RL from scratch or adapting from pretrained policies remains rare, limiting the full potential of humanoid robots. Real-world learning, despite being crucial for overcoming the sim-to-real gap, faces substantial challenges related to safety, reward design, and learning efficiency. To address these limitations, we propose Robot-Trains-Robot (RTR), a novel framework where a robotic arm teacher actively supports and guides a humanoid robot student. The RTR system provides protection, learning schedule, reward, perturbation, failure detection, and automatic resets. It enables efficient long-term real-world humanoid training with minimal human intervention. Furthermore, we propose a novel RL pipeline that facilitates and stabilizes sim-to-real transfer by optimizing a single dynamics-encoded latent variable in the real world. We validate our method through two challenging real-world humanoid tasks: fine-tuning a walking policy for precise speed tracking and learning a humanoid swing-up task from scratch, illustrating the promising capabilities of real-world humanoid learning realized by RTR-style systems. See https://robot-trains-robot.github.io/ for more info.

人形机器人强化学习真实世界训练自主学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。