为多智能体强化学习设计安全约束框架,减少违规并提升协作。
Think Smart, Act SMARL! Analyzing Probabilistic Logic Shields for Multi-Agent Reinforcement Learning
- 将概率逻辑约束融入价值更新与策略梯度,实现安全学习。
- 在多玩家博弈中显著降低违规率,合作效果更优。
- 适合需要安全合规的多智能体系统研究者使用。
安全强化学习对实际应用至关重要,而多智能体交互带来了额外的安全挑战。尽管概率逻辑屏蔽(PLS)在单智能体强化学习中已证明有效,但其在多智能体环境中的泛化能力尚不明确。本文通过在去中心化多智能体环境中进行广泛分析,提出了一种通用框架——屏蔽式多智能体强化学习(SMARL),以引导多智能体强化学习向符合规范的成果发展。主要贡献包括:(1) 提出一种新的概率逻辑时序差分(PLTD)更新方法,将概率约束直接嵌入独立Q学习的价值更新;(2) 设计一种带概率逻辑的策略梯度方法,用于屏蔽式PPO,实现多智能体强化学习的形式化安全保证;(3) 在对称与非对称多智能体博弈基准上进行全面评估,结果表明在规范约束下违规更少、协作显著提升。这些成果使SMARL成为有效的均衡选择机制,推动更安全、社会契合的多智能体系统发展。
原文摘要 · Abstract (English)
Safe reinforcement learning (RL) is crucial for real-world applications, and multi-agent interactions introduce additional safety challenges. While Probabilistic Logic Shields (PLS) has been a powerful proposal to enforce safety in single-agent RL, their generalizability to multi-agent settings remains unexplored. In this paper, we address this gap by conducting extensive analyses of PLS within decentralized, multi-agent environments, and in doing so, propose $\textbf{Shielded Multi-Agent Reinforcement Learning (SMARL)}$ as a general framework for steering MARL towards norm-compliant outcomes. Our key contributions are: (1) a novel Probabilistic Logic Temporal Difference (PLTD) update for shielded, independent Q-learning, which incorporates probabilistic constraints directly into the value update process; (2) a probabilistic logic policy gradient method for shielded PPO with formal safety guarantees for MARL; and (3) comprehensive evaluation across symmetric and asymmetrically shielded $n$-player game-theoretic benchmarks, demonstrating fewer constraint violations and significantly better cooperation under normative constraints. These results position SMARL as an effective mechanism for equilibrium selection, paving the way toward safer, socially aligned multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。