无需人工奖励或操作,机器人仅看视频就能学会抓放动作。
Online World Modeling Enables Real-World Inverse Reinforcement Learning from Observation
- 通过观察视频直接逆向学习奖励函数,实现无监督视觉操控。
- 40分钟内成功率达82%,远超传统方法的0%和12%。
- 首次实现真实世界跨任务正向在线迁移,适合零样本学习场景。
当前机器人学习受限于人工设计奖励、仿真建模或动作监督(如遥操作),需大量领域知识与人力投入。本文提出基于观察的逆强化学习(IRLfO),仅依赖任务观察(如视频)进行学习。由于强化学习方法在该场景下挑战大,此前实现实体机器人学习尚不现实。我们首次实现了从零开始在真实世界中学习视觉抓放任务,并首次展示跨视觉操控任务的正向在线迁移。在40分钟内,MPAIL2模型成功率高达82%,而同等交互与示范预算下的强化学习和行为克隆方法分别仅达0%和12%。项目互动页面含训练视频:https://uwrobotlearning.github.io/mpail2/
原文摘要 · Abstract (English)
Current methods in robot learning are fundamentally bottlenecked by one or more of: hand-designed rewards, simulation modeling, or action supervision (e.g. teleoperation) each requiring significant domain expertise, engineering effort, and robot-operator labor. Towards eliminating these bottlenecks, this work pursues observational learning via Inverse Reinforcement Learning from Observation (IRLfO) in which only access to task observations (e.g. video) is assumed. Due to the challenging setting and limitations of RL methods, IRLfO has thus far remained impractical for real-world robot learning. Here, we present the first IRL method to learn visual manipulation in the real world from scratch, and the first real-world demonstration of positive online transfer across visual manipulation tasks from scratch. In under 40 minutes, MPAIL2 learns pick-and-place from scratch to 82% success, where RL and BC with equal interaction and demonstration budgets reach only 0% and 12% despite their reward and action supervision. Interactive project page with training videos: https://uwrobotlearning.github.io/mpail2/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。