单人示范教多机器人协作,效率媲美专业同步演示。
R2BC: Multi-Agent Imitation Learning from Single-Agent Demonstrations
- 人类逐个操控机器人示范,系统自动整合为团队行为
- 四组仿真任务中表现接近甚至超过同步示范的基准
- 适合真实场景下缺乏多智能体协同示范的机器人训练
模仿学习(IL)是人类教授机器人的一种自然方式,尤其在高质量示范易获取时。尽管单机器人场景中广泛应用,但将此类方法扩展至多智能体系统仍较少研究,尤其是在单个操作者需为协作机器人团队提供示范的情况下。本文提出轮转行为克隆(R2BC),使单一人类操作者可通过顺序、单智能体示范有效训练多机器人系统。该方法允许人类依次操控每个机器人,并逐步传授团队协作行为,无需在联合动作空间中进行同步示范。我们在四个多智能体仿真任务中验证,R2BC性能可达到甚至超越基于理想同步示范的基准方法。最后,我们使用真实人类示范,在两个物理机器人任务上部署了R2BC。
原文摘要 · Abstract (English)
Imitation Learning (IL) is a natural way for humans to teach robots, particularly when high-quality demonstrations are easy to obtain. While IL has been widely applied to single-robot settings, relatively few studies have addressed the extension of these methods to multi-agent systems, especially in settings where a single human must provide demonstrations to a team of collaborating robots. In this paper, we introduce and study Round-Robin Behavior Cloning (R2BC), a method that enables a single human operator to effectively train multi-robot systems through sequential, single-agent demonstrations. Our approach allows the human to teleoperate one agent at a time and incrementally teach multi-agent behavior to the entire system, without requiring demonstrations in the joint multi-agent action space. We show that R2BC methods match, and in some cases surpass, the performance of an oracle behavior cloning approach trained on privileged synchronized demonstrations across four multi-agent simulated tasks. Finally, we deploy R2BC on two physical robot tasks trained using real human demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。