看野生动物视频学走路,机器人也能学会复杂动作。
Reinforcement Learning from Wild Animal Videos
- 用自然纪录片视频训练动作识别模型,提取动物运动特征。
- 通过视频分类得分作为奖励,让仿真机器人学会行走、跳跃等技能。
- 无需参考轨迹或特定奖励,直接迁移到真实四足机器人上。
我们提出一种从互联网野生动物视频中学习腿式机器人运动技能的方法,即通过观看大量自然纪录片中的动物视频来获取多样化的运动范例。为此,我们引入了强化学习从野生动物视频(RLWAV)的方法,将这些运动模式转化为物理机器人可执行的动作。首先,在大规模动物视频数据集上训练一个视频分类器,用于识别自然栖息地中动物的运动行为。接着,在物理仿真环境中训练一个多技能策略,以第三视角摄像头拍摄的机器人运动视频的分类得分作为强化学习的奖励信号。最后,我们将训练好的策略直接迁移至真实的四足机器人Solo上。令人惊讶的是,尽管动物与机器人在环境和形态上存在巨大差异,该方法仍成功使机器人学习到行走、跳跃和静止等多种技能,且不依赖参考轨迹或针对特定技能设计的奖励函数。
原文摘要 · Abstract (English)
We propose to learn legged robot locomotion skills by watching thousands of wild animal videos from the internet, such as those featured in nature documentaries. Indeed, such videos offer a rich and diverse collection of plausible motion examples, which could inform how robots should move. To achieve this, we introduce Reinforcement Learning from Wild Animal Videos (RLWAV), a method to ground these motions into physical robots. We first train a video classifier on a large-scale animal video dataset to recognize actions from RGB clips of animals in their natural habitats. We then train a multi-skill policy to control a robot in a physics simulator, using the classification score of a third-person camera capturing videos of the robot's movements as a reward for reinforcement learning. Finally, we directly transfer the learned policy to a real quadruped Solo. Remarkably, despite the extreme gap in both domain and embodiment between animals in the wild and robots, our approach enables the policy to learn diverse skills such as walking, jumping, and keeping still, without relying on reference trajectories nor skill-specific rewards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。