arXiv:2412.19669cs.ROcs.LG2024-12被引 16

用学习方法替代传统优化,实现万级机器人快速协同控制。

Toward Scalable Multirobot Control: Fast Policy Learning in Distributed MPC

  • 通过在线分布式强化学习生成闭环控制策略,避免反复求解优化问题。
  • 在万级机器人系统上实现快速学习与部署,计算效率提升显著。
  • 适合大规模多机器人系统,尤其适用于动态变化的复杂任务场景。

分布式模型预测控制(DMPC)在多机器人系统(MRS)中具有实现最优协同控制的潜力,但其实时实现依赖于数值优化工具在线周期性计算局部控制序列,该过程计算量大,难以扩展至大规模非线性系统。本文提出一种新型分布式学习型预测控制(DLPC)框架,用于实现可扩展的多机器人控制。与传统基于开环控制序列的DMPC方法不同,本方法采用一种计算高效、分布式的策略学习算法,直接生成显式的闭环DMPC策略,无需数值求解器。策略学习在每个预测区间内以增量式、前向方式进行,通过在线分布式演员-评论家架构实现。控制策略以滚动时域方式持续更新,保证闭环稳定性。所学策略可部署于规模可变的MRS中,显著提升可扩展性和迁移能力。此外,我们还引入受力场启发的策略学习方法,应对多机器人安全学习挑战。通过大规模轮式机器人和多旋翼无人机的协同任务实验验证了方法的有效性、可扩展性与高效性。结果表明,该方法可在10,000台机器人规模下实现快速策略学习与部署。

原文摘要 · Abstract (English)

Distributed model predictive control (DMPC) is promising in achieving optimal cooperative control in multirobot systems (MRS). However, real-time DMPC implementation relies on numerical optimization tools to periodically calculate local control sequences online. This process is computationally demanding and lacks scalability for large-scale, nonlinear MRS. This article proposes a novel distributed learning-based predictive control (DLPC) framework for scalable multirobot control. Unlike conventional DMPC methods that calculate open-loop control sequences, our approach centers around a computationally fast and efficient distributed policy learning algorithm that generates explicit closed-loop DMPC policies for MRS without using numerical solvers. The policy learning is executed incrementally and forward in time in each prediction interval through an online distributed actor-critic implementation. The control policies are successively updated in a receding-horizon manner, enabling fast and efficient policy learning with the closed-loop stability guarantee. The learned control policies could be deployed online to MRS with varying robot scales, enhancing scalability and transferability for large-scale MRS. Furthermore, we extend our methodology to address the multirobot safe learning challenge through a force field-inspired policy learning approach. We validate our approach's effectiveness, scalability, and efficiency through extensive experiments on cooperative tasks of large-scale wheeled robots and multirotor drones. Our results demonstrate the rapid learning and deployment of DMPC policies for MRS with scales up to 10,000 units.

多机器人控制强化学习分布式系统高效部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。