让机器人超越受限专家,学出更快更优的执行路径。
When a Robot is More Capable than a Human: Learning from Constrained Demonstrators
- 用演示数据推断状态奖励,实现任务进度度量
- 通过时间插值自标注未知状态奖励,提升学习效率
- 实测在真实机械臂上比行为克隆快10倍
学习从示范中使专家能通过运动教学、操纵杆控制和仿真到现实迁移等接口教机器人完成复杂任务。然而,这些接口常因间接控制、设置限制和硬件安全而制约专家展示最优行为。例如,操纵杆仅能二维移动机械臂,尽管机器人本身处于高维空间。因此,受限专家提供的示范导致学习策略表现不佳。核心问题在于:机器人能否学会优于受限专家示范的策略?我们通过允许代理超越直接模仿专家动作,探索更短更高效的轨迹来解决此问题。利用示范数据推断仅基于状态的奖励信号以衡量任务进展,并通过时间插值对未知状态进行自标注奖励。该方法在样本效率和任务完成时间上均优于常见模仿学习。在真实WidowX机械臂上,任务仅需12秒完成,比行为克隆快10倍,视频展示见https://sites.google.com/view/constrainedexpert。
原文摘要 · Abstract (English)
Learning from demonstrations enables experts to teach robots complex tasks using interfaces such as kinesthetic teaching, joystick control, and sim-to-real transfer. However, these interfaces often constrain the expert's ability to demonstrate optimal behavior due to indirect control, setup restrictions, and hardware safety. For example, a joystick can move a robotic arm only in a 2D plane, even though the robot operates in a higher-dimensional space. As a result, the demonstrations collected by constrained experts lead to suboptimal performance of the learned policies. This raises a key question: Can a robot learn a better policy than the one demonstrated by a constrained expert? We address this by allowing the agent to go beyond direct imitation of expert actions and explore shorter and more efficient trajectories. We use the demonstrations to infer a state-only reward signal that measures task progress, and self-label reward for unknown states using temporal interpolation. Our approach outperforms common imitation learning in both sample efficiency and task completion time. On a real WidowX robotic arm, it completes the task in 12 seconds, 10x faster than behavioral cloning, as shown in real-robot videos on https://sites.google.com/view/constrainedexpert .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。