检测多智能体强化学习中故障或敌对智能体,提升交通信号协同控制的可靠性。
BARD-MARL: Byzantine-Agent Detection for Learned Communication in Multi-Agent Reinforcement Learning

- 融合策略图特征与贝叶斯信任统计,双路诊断识别异常智能体。
- 100智能体场景下,对固定动作和协同攻击均达0.982 AUC-ROC。
- 揭示了不同攻击类型下检测方法的互补性,适合可信系统设计者。
学习通信虽能提升多智能体强化学习中的协作效率,但也引入信任问题:训练后的策略可能经由故障或敌对智能体传递信息。本文研究自适应交通信号控制中基于学习通信的拜占庭智能体检测问题。提出BARD-MARL,作为在BayesG之上的后验诊断层,其本身不构成论文贡献。BARD-MARL结合两类智能体级证据:从状态-动作轨迹中提取的策略图特征,以及基于BayesG潜在掩码概率计算的贝叶斯信任统计。在SUMO交通网格中测试固定动作、观测翻转、随机噪声及协同攻击场景,结果表明这两类信号互补而非始终主导。在25智能体网格中,观测翻转攻击下BARD-MARL达到0.843 AUC-ROC,仅用策略图检测则在协同攻击下达0.917。在100智能体网格中,统一版BARD-MARL对10%固定动作和10%协同攻击均达0.982 AUC-ROC。研究证明学习通信策略蕴含可利用的诊断线索,但可信鲁棒性声明需依赖攻击特定消融实验,并明确区分协调、检测与缓解机制。
原文摘要 · Abstract (English)
Learned communication improves coordination in cooperative multi-agent reinforcement learning, but it also creates a trust problem: a trained policy may route information through agents that have become faulty or adversarial. This paper studies Byzantine-agent detection for learned-communication MARL in adaptive traffic signal control. We propose BARD-MARL, a post-hoc diagnostic layer on top of BayesG, which is used as an attributed communication substrate rather than as a contribution of this paper. BARD-MARL combines two agent-level evidence streams: policy-graph features extracted from state-action trajectories and Bayesian trust statistics computed from BayesG latent mask probabilities. Across fixed-action, observation-flip, random-noise, and coordinated attacks in SUMO traffic grids, the results show that these signals are complementary rather than universally dominant. On a 25-agent grid, BARD-MARL reaches 0.843 AUC-ROC under a 10% observation-flip attack, while policy-graph-only detection reaches 0.917 AUC-ROC under a 10% coordinated attack. On a 100-agent grid, the unified BARD-MARL variant reaches 0.982 AUC-ROC for both 10% fixed-action and 10% coordinated attacks. The study shows that learned communication policies expose useful diagnostic evidence, but credible resilience claims require attack-specific ablations and explicit separation between coordination, detection, and mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。