仅通过一个代理植入后门,即可操控整个多智能体系统。
BLAST: A Stealthy Backdoor Leverage Attack against Cooperative Multi-Agent Deep Reinforcement Learning based Systems
- 用动态行为模式做触发器,隐蔽性强。
- 单个代理攻击成功率高,干净性能损失小。
- 适合研究多智能体安全或防御的开发者参考。
近期研究表明,协作式多智能体深度强化学习(c-MADRL)面临后门攻击威胁。一旦检测到后门触发信号,系统将执行恶意行为导致失败或达成恶意目标。然而现有攻击存在诸多问题:触发模式缺乏隐蔽性、需额外网络训练或激活后门、所有智能体均被植入后门。为此,本文提出一种新型后门杠杆攻击——BLAST,仅在单个智能体中嵌入后门即可攻击整个多智能体团队。首先,引入对手时空行为模式作为后门触发器,而非固定视觉模式或瞬时状态,确保攻击的隐蔽性与实用性;其次,通过单边引导篡改后门智能体的原始奖励函数,实现“杠杆攻击效应”,即以单一后门代理撬动整个系统。我们在SMAC和Pursuit两个主流c-MADRL环境中,对三种经典算法(VDN、QMIX、MAPPO)及两种现有防御机制进行了评估。实验结果表明,BLAST在保持低干净性能方差率的同时,仍能实现高攻击成功率。
原文摘要 · Abstract (English)
Recent studies have shown that cooperative multi-agent deep reinforcement learning (c-MADRL) is under the threat of backdoor attacks. Once a backdoor trigger is observed, it will perform malicious actions leading to failures or malicious goals. However, existing backdoor attacks suffer from several issues, e.g., instant trigger patterns lack stealthiness, the backdoor is trained or activated by an additional network, or all agents are backdoored. To this end, in this paper, we propose a novel backdoor leverage attack against c-MADRL, BLAST, which attacks the entire multi-agent team by embedding the backdoor only in a single agent. Firstly, we introduce adversary spatiotemporal behavior patterns as the backdoor trigger rather than manual-injected fixed visual patterns or instant status and control the period to perform malicious actions. This method can guarantee the stealthiness and practicality of BLAST. Secondly, we hack the original reward function of the backdoor agent via unilateral guidance to inject BLAST, so as to achieve the \textit{leverage attack effect} that can pry open the entire multi-agent system via a single backdoor agent. We evaluate our BLAST against 3 classic c-MADRL algorithms (VDN, QMIX, and MAPPO) in 2 popular c-MADRL environments (SMAC and Pursuit), and 2 existing defense mechanisms. The experimental results demonstrate that BLAST can achieve a high attack success rate while maintaining a low clean performance variance rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。