arXiv:2509.15799eess.SYcs.AI2025-09

用分层强化学习+模型预测控制,让多智能体在复杂环境中更安全高效协作。

Hierarchical Reinforcement Learning with Low-Level MPC for Multi-Agent Control

  • 高层用RL选目标区域,低层用MPC保证动作安全可行
  • 在捕食者-猎物任务中,奖励更高、更安全、结果更稳定
  • 适合需要高安全性与可靠性的多智能体系统场景

在动态且约束密集的环境中实现安全与协调行为仍是基于学习的控制面临的重大挑战。纯端到端学习通常样本效率差且可靠性有限,而基于模型的方法依赖预设参考轨迹,泛化能力不足。本文提出一种分层框架,结合强化学习(RL)进行战术决策与模型预测控制(MPC)实现低层执行。对于多智能体系统,高层策略从结构化的兴趣区域(ROIs)中选择抽象目标,而MPC确保运动在动态上可行且安全。在捕食者-猎物基准测试中,该方法在奖励、安全性和一致性方面均优于端到端及屏蔽式强化学习基线,验证了结构化学习与模型化控制结合的优势。

原文摘要 · Abstract (English)

Achieving safe and coordinated behavior in dynamic, constraint-rich environments remains a major challenge for learning-based control. Pure end-to-end learning often suffers from poor sample efficiency and limited reliability, while model-based methods depend on predefined references and struggle to generalize. We propose a hierarchical framework that combines tactical decision-making via reinforcement learning (RL) with low-level execution through Model Predictive Control (MPC). For the case of multi-agent systems this means that high-level policies select abstract targets from structured regions of interest (ROIs), while MPC ensures dynamically feasible and safe motion. Tested on a predator-prey benchmark, our approach outperforms end-to-end and shielding-based RL baselines in terms of reward, safety, and consistency, underscoring the benefits of combining structured learning with model-based control.

多智能体强化学习模型预测控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。