arXiv:2409.07775cs.AIcs.CR2024-09被引 6

仅在单个智能体中植入隐蔽后门,即可攻陷整个多智能体系统。

A Spatiotemporal Stealthy Backdoor Attack against Cooperative Multi-Agent Deep Reinforcement Learning

  • 用时空行为模式作触发器,隐蔽性强且无需固定视觉图案。
  • 通过奖励反转与单向引导,使单个智能体破坏整体团队性能。
  • 实验成功率91.6%,正常表现波动仅3.7%,适合研究安全防护者参考。

近期研究显示,协作式多智能体深度强化学习(c-MADRL)面临后门攻击威胁。一旦检测到后门触发信号,系统将执行异常行为导致失败或达成恶意目标。然而,现有后门攻击存在诸多问题:如固定视觉触发模式缺乏隐蔽性、需额外网络训练或激活、所有智能体均被植入后门等。为此,本文提出一种新型后门攻击方法,仅在单个智能体中嵌入后门即可攻击整个多智能体团队。首先,引入对抗性时空行为模式作为触发器,而非手动注入的固定视觉图案或瞬时状态,有效保障后门的隐蔽性与实用性。其次,在训练阶段通过奖励反转与单向引导篡改被攻陷智能体的原始奖励函数,确保其对团队产生负面影响。我们在两个经典c-MADRL算法VDN和QMIX上,于流行环境SMAC中验证该攻击。实验结果表明,该攻击在保持低清洁性能方差率(3.7%)的同时,达到91.6%的高攻击成功率。

原文摘要 · Abstract (English)

Recent studies have shown that cooperative multi-agent deep reinforcement learning (c-MADRL) is under the threat of backdoor attacks. Once a backdoor trigger is observed, it will perform abnormal actions leading to failures or malicious goals. However, existing proposed backdoors suffer from several issues, e.g., fixed visual trigger patterns lack stealthiness, the backdoor is trained or activated by an additional network, or all agents are backdoored. To this end, in this paper, we propose a novel backdoor attack against c-MADRL, which attacks the entire multi-agent team by embedding the backdoor only in a single agent. Firstly, we introduce adversary spatiotemporal behavior patterns as the backdoor trigger rather than manual-injected fixed visual patterns or instant status and control the attack duration. This method can guarantee the stealthiness and practicality of injected backdoors. Secondly, we hack the original reward function of the backdoored agent via reward reverse and unilateral guidance during training to ensure its adverse influence on the entire team. We evaluate our backdoor attacks on two classic c-MADRL algorithms VDN and QMIX, in a popular c-MADRL environment SMAC. The experimental results demonstrate that our backdoor attacks are able to reach a high attack success rate (91.6\%) while maintaining a low clean performance variance rate (3.7\%).

后门攻击多智能体强化学习安全威胁

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。