通过课程学习提升多机器人长时协调能力,从全局演示中训练出鲁棒分布式策略。
Curriculum Imitation Learning of Distributed Multi-Robot Policies
- 用渐进式轨迹长度课程学习,增强长期行为的稳定性与准确性。
- 仅用第三人称状态演示,模拟各机器人本地感知,适配真实传感器噪声。
- 无需专家动作或机载测量,可从全局演示中学习鲁棒分布式控制策略。
多机器人系统(MRS)的控制策略学习面临长期协调困难与真实训练数据获取难的挑战。本文在模仿学习框架下同时解决这两点:首先,将课程学习重心从机器人数规模扩展转向提升长期协调能力,提出逐步增加专家轨迹长度的课程策略,稳定训练并提高长期行为精度;其次,提出一种仅利用第三人称全局状态演示即可近似各机器人本体感知的方法,通过过滤邻居、坐标系转换和模拟机载传感器波动,将理想轨迹转化为局部可观测信息。两项贡献结合形成物理启发的可扩展分布式策略生成方法。在两个任务上测试不同团队规模与噪声水平,结果表明课程学习显著提升长期准确率,感知估计方法使策略对现实不确定性具备鲁棒性。两者协同实现从全局演示中学习鲁棒分布式控制器,无需专家动作或机载测量。
原文摘要 · Abstract (English)
Learning control policies for multi-robot systems (MRS) remains a major challenge due to long-term coordination and the difficulty of obtaining realistic training data. In this work, we address both limitations within an imitation learning framework. First, we shift the typical role of Curriculum Learning in MRS, from scalability with the number of robots, to focus on improving long-term coordination. We propose a curriculum strategy that gradually increases the length of expert trajectories during training, stabilizing learning and enhancing the accuracy of long-term behaviors. Second, we introduce a method to approximate the egocentric perception of each robot using only third-person global state demonstrations. Our approach transforms idealized trajectories into locally available observations by filtering neighbors, converting reference frames, and simulating onboard sensor variability. Both contributions are integrated into a physics-informed technique to produce scalable, distributed policies from observations. We conduct experiments across two tasks with varying team sizes and noise levels. Results show that our curriculum improves long-term accuracy, while our perceptual estimation method yields policies that are robust to realistic uncertainty. Together, these strategies enable the learning of robust, distributed controllers from global demonstrations, even in the absence of expert actions or onboard measurements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。