在观测受限下,用少量样本攻击多智能体强化学习系统
Constrained Black-Box Attacks Against Cooperative Multi-Agent Reinforcement Learning
- 仅通过扰动部署后智能体的观测数据实施攻击
- 仅需1000次采样即生效,远低于以往方法的百万级需求
- 适用于真实场景中无权访问策略权重的隐蔽攻击
协作式多智能体强化学习虽已广泛应用于实际场景,但其对抗攻击脆弱性尚未被充分研究。现有工作多聚焦于训练阶段攻击或理想化假设(如可获取策略权重或训练替代模型)。本文在更严苛的约束条件下探索新漏洞:攻击者仅能收集并扰动已部署智能体的观测数据,甚至完全无观测、动作或权重访问权限。我们提出通过生成故意使目标智能体感知环境错乱的扰动来实现攻击,在三个基准和22个环境中验证了有效性,覆盖多种算法与环境。实验表明,该方法样本效率极高,仅需1,000次采样即可成功,而此前方法需数百万次。
原文摘要 · Abstract (English)
Collaborative multi-agent reinforcement learning has rapidly evolved, offering state-of-the-art algorithms for real-world applications, including sensitive domains. However, a key challenge to its widespread adoption is the lack of a thorough investigation into its vulnerabilities to adversarial attacks. Existing work predominantly focuses on training-time attacks or unrealistic scenarios, such as access to policy weights or the ability to train surrogate policies. In this paper, we investigate new vulnerabilities under more challenging and constrained conditions, assuming an adversary can only collect and perturb the observations of deployed agents. We also consider scenarios where the adversary has no access at all (no observations, actions, or weights). Our main approach is to generate perturbations that intentionally misalign how victim agents see their environment. Our approach is empirically validated on three benchmarks and 22 environments, demonstrating its effectiveness across diverse algorithms and environments. Furthermore, we show that our algorithm is sample-efficient, requiring only 1,000 samples compared to the millions needed by previous methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。