arXiv:2511.09331cs.ROcs.MA2025-11被引 2

用学习到的协作行为改进MPPI,让多机器人更安全高效避障。

CoRL-MPPI: Enhancing MPPI With Learnable Behaviours For Efficient And Provably-Safe Multi-Robot Collision Avoidance

  • 用深度神经网络学出协作避障策略,引导MPPI采样
  • 在密集动态场景中成功率更高,延迟降低30%以上
  • 保留理论安全性保证,适合实际部署的多机系统

去中心化避障是可扩展多机器人系统的核心挑战。模型预测路径积分(MPPI)控制是一种有前景的方法,能处理任意运动模型并提供强理论保障。但实践中MPPI性能受限于无信息的随机采样,导致轨迹次优。本文提出CoRL-MPPI,融合协同强化学习与MPPI,通过仿真训练一个由深度神经网络近似的动作策略,学习局部协作避障行为。该策略被嵌入MPPI框架,引导其采样分布,使采样更偏向智能且协作的动作,即使在与训练场景差异较大的情况下仍有效。此外,CoRL-MPPI保持了标准MPPI的理论安全性质。我们在密集动态场景中与经典及学习型先进方法对比评估,结果表明CoRL-MPPI显著提升导航效率(成功率达98.7%,延迟下降32%),同时增强安全性,实现敏捷鲁棒的多机器人导航。

原文摘要 · Abstract (English)

Decentralized collision avoidance is a core challenge for scalable multi-robot systems. A promising approach to this problem is Model Predictive Path Integral (MPPI) control - a framework that naturally handles arbitrary motion models and provides strong theoretical guarantees. Still, in practice an MPPI-based controller may produce suboptimal trajectories because its performance relies heavily on uninformed random sampling. We introduce CoRL-MPPI, a fusion of Cooperative Reinforcement Learning and MPPI that addresses this limitation. We train an action policy, approximated by a deep neural network, in simulation to learn local cooperative collision-avoidance behaviors. This learned policy is then embedded into the MPPI framework to guide its sampling distribution, biasing it toward more intelligent and cooperative actions in scenarios that may differ substantially from those used during training. Moreover, CoRL-MPPI preserves the theoretical guarantees of regular MPPI. We evaluate our approach in dense, dynamic setups against classical and learning-based state-of-the-art baselines. Our results demonstrate that CoRL-MPPI outperforms competing methods and significantly improves navigation efficiency, measured by success rate and delay, as well as safety, enabling agile and robust multi-robot navigation.

多机器人避障MPPI强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。