arXiv:2603.02783cs.ROcs.LG2026-03中稿 · publication at the…

用人类示范训练机器人集群,让多机协同行为自然模仿真实动作。

Generative adversarial imitation learning for robot swarms: Learning from human demonstrations and trained policies

  • 基于生成对抗模仿学习框架,从人类示范中提取集体行为模式。
  • 在6个任务中表现接近示范水平,仿真与实机测试效果一致。
  • 适用于需群体协作的机器人系统,尤其适合缺乏标注数据场景。

在模仿学习中,机器人需从期望行为的示范中学习。现有集群机器人模仿学习大多使用已有策略的轨迹作为示范。本文提出一种基于生成对抗模仿学习的框架,旨在从人类示范中学习集体行为。该框架在六个不同任务中进行评估,既接受人工示范,也接受由PPO训练策略生成的示范。结果表明,模仿学习过程能够习得定性上有意义的行为,其性能与提供示范相当。此外,我们将在仿真中学习到的策略部署于一组TurtleBot 4机器人上进行真实机器人实验,所展现的行为保持了可视可辨特征,且性能与仿真中相当。

原文摘要 · Abstract (English)

In imitation learning, robots are supposed to learn from demonstrations of the desired behavior. Most of the work in imitation learning for swarm robotics provides the demonstrations as rollouts of an existing policy. In this work, we provide a framework based on generative adversarial imitation learning that aims to learn collective behaviors from human demonstrations. Our framework is evaluated across six different missions, learning both from manual demonstrations and demonstrations derived from a PPO-trained policy. Results show that the imitation learning process is able to learn qualitatively meaningful behaviors that perform similarly well as the provided demonstrations. Additionally, we deploy the learned policies on a swarm of TurtleBot 4 robots in real-robot experiments. The exhibited behaviors preserved their visually recognizable character and their performance is comparable to the one achieved in simulation.

模仿学习集群机器人生成对抗实机部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。