提出协作式安全通信框架,让多智能体提前预判风险并协同避险。
Co2PO: Coordinated Constrained Policy Optimization for Multi-Agent RL
- 通过共享黑板广播意图和产量信号,实现风险感知的主动通信。
- 在多个复杂安全基准上超越现有方法,收益更高且满足约束条件。
- 适合需要高安全性的多智能体系统,如自动驾驶编队、机器人协作。
受限多智能体强化学习面临探索与安全约束之间的根本矛盾。现有主流方法如拉格朗日法依赖全局惩罚或中心化评判器,在违规发生后才响应,常抑制探索并导致过度保守。本文提出Co2PO,一种基于通信增强的新型多智能体框架,通过选择性、风险感知的通信实现协调驱动的安全性。Co2PO引入共享黑板架构,广播位置意图和产量信号,由学习得到的危险预测器在更长时序范围内前瞻性预测潜在违规。将这些预测整合进约束优化目标,使智能体能在不牺牲性能的前提下提前预判并规避集体风险。我们在一系列复杂多智能体安全基准上评估Co2PO,结果表明其相比领先受限基线获得更高回报,并在部署时收敛至符合成本约束的策略。消融实验进一步验证了风险触发通信、自适应门控和共享记忆组件的必要性。
原文摘要 · Abstract (English)
Constrained multi-agent reinforcement learning (MARL) faces a fundamental tension between exploration and safety-constrained optimization. Existing leading approaches, such as Lagrangian methods, typically rely on global penalties or centralized critics that react to violations after they occur, often suppressing exploration and leading to over-conservatism. We propose Co2PO, a novel MARL communication-augmented framework that enables coordination-driven safety through selective, risk-aware communication. Co2PO introduces a shared blackboard architecture for broadcasting positional intent and yield signals, governed by a learned hazard predictor that proactively forecasts potential violations over an extended temporal horizon. By integrating these forecasts into a constrained optimization objective, Co2PO allows agents to anticipate and navigate collective hazards without the performance trade-offs inherent in traditional reactive constraints. We evaluate Co2PO across a suite of complex multi-agent safety benchmarks, where it achieves higher returns compared to leading constrained baselines while converging to cost-compliant policies at deployment. Ablation studies further validate the necessity of risk-triggered communication, adaptive gating, and shared memory components.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。