通过安全约束提升多智能体强化学习的稳定性和收敛性。
Tackling Uncertainties in Multi-Agent Reinforcement Learning through Integration of Agent Termination Dynamics
- 引入基于屏障函数的安全损失,融合系统故障特征。
- 在星际争霸微操任务中收敛更快,安全性与完成率双优。
- 适合需要高可靠性与安全探索的复杂多智能体场景。
多智能体强化学习(MARL)在解决复杂现实任务中受到广泛关注,但环境中的固有随机性和不确定性给高效、鲁棒的策略学习带来挑战。尽管分布式强化学习在单智能体设置中成功应对风险与不确定性,但在MARL中的应用仍受限。本文提出一种新方法,将分布学习与基于安全性的损失函数结合,以改善协作式MARL任务的收敛性。具体而言,我们引入一种基于屏障函数的损失,利用系统内在故障识别出的安全度量,融入策略学习过程。该额外损失项有助于降低风险,并在训练初期促进更安全的探索。我们在星际争霸II微操基准上评估了该方法,结果表明其在收敛速度、安全性和任务完成率方面均优于现有最先进基线。研究结果表明,融入安全考量可显著提升复杂多智能体环境中的学习性能。
原文摘要 · Abstract (English)
Multi-Agent Reinforcement Learning (MARL) has gained significant traction for solving complex real-world tasks, but the inherent stochasticity and uncertainty in these environments pose substantial challenges to efficient and robust policy learning. While Distributional Reinforcement Learning has been successfully applied in single-agent settings to address risk and uncertainty, its application in MARL is substantially limited. In this work, we propose a novel approach that integrates distributional learning with a safety-focused loss function to improve convergence in cooperative MARL tasks. Specifically, we introduce a Barrier Function based loss that leverages safety metrics, identified from inherent faults in the system, into the policy learning process. This additional loss term helps mitigate risks and encourages safer exploration during the early stages of training. We evaluate our method in the StarCraft II micromanagement benchmark, where our approach demonstrates improved convergence and outperforms state-of-the-art baselines in terms of both safety and task completion. Our results suggest that incorporating safety considerations can significantly enhance learning performance in complex, multi-agent environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。