让机器人从模拟到现实精准完成投掷、举重等高难度动作
Bridging the Sim-to-Real Gap for Athletic Loco-Manipulation
- 用无监督驱动网络利用真实数据缩小仿真与现实差距
- 结合参考轨迹预训练,使机器人在真实环境实现精准动作
- 适合研究机器人动态操控与仿真迁移的学者
实现机器人在运动中的复杂操控需超越传统跟踪奖励,转向能驱动真正动态、目标导向行为的任务奖励。例如“尽可能远地投球”或“快速举起重物”等指令,能激发机器人的敏捷性与爆发力。然而,仅使用任务奖励会引发两大挑战:奖励被滥用(奖励黑客),且探索过程缺乏有效引导。为此,本文提出两阶段训练流程:首先引入无监督驱动网络(UAN),利用真实数据在无需扭矩传感的情况下缩小复杂执行机构的仿真-现实差距;其次采用预训练与微调策略,以参考轨迹作为初始提示,引导探索方向。通过这些创新,机器人在仿真中学习的举、抛、拖等动作可高保真迁移到真实世界。
原文摘要 · Abstract (English)
Achieving athletic loco-manipulation on robots requires moving beyond traditional tracking rewards - which simply guide the robot along a reference trajectory - to task rewards that drive truly dynamic, goal-oriented behaviors. Commands such as "throw the ball as far as you can" or "lift the weight as quickly as possible" compel the robot to exhibit the agility and power inherent in athletic performance. However, training solely with task rewards introduces two major challenges: these rewards are prone to exploitation (reward hacking), and the exploration process can lack sufficient direction. To address these issues, we propose a two-stage training pipeline. First, we introduce the Unsupervised Actuator Net (UAN), which leverages real-world data to bridge the sim-to-real gap for complex actuation mechanisms without requiring access to torque sensing. UAN mitigates reward hacking by ensuring that the learned behaviors remain robust and transferable. Second, we use a pre-training and fine-tuning strategy that leverages reference trajectories as initial hints to guide exploration. With these innovations, our robot athlete learns to lift, throw, and drag with remarkable fidelity from simulation to reality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。