无需直接交互,用中立代理在多方系统中攻击强化学习模型。
Neutral Agent-based Adversarial Policy Learning against Deep Reinforcement Learning in Multi-party Open Systems
- 设计中立代理通过共享环境间接干扰目标智能体。
- 在SMAC和Highway-env上实现有效攻击,成功率超80%。
- 适合研究对抗性攻防或开放系统安全的学者。
强化学习(RL)是解决不确定性下长期序列决策问题的重要范式。将深度神经网络(DNN)融入RL框架后,深度强化学习(DRL)在多个领域取得显著成功。然而,DNN的引入也使其易受对抗攻击。现有攻击方法主要依赖对环境的完全控制或与目标智能体直接交互,难以应用于多方开放系统。为此,本文提出一种基于中立代理的对抗策略学习方法,在不直接交互、不控制环境的前提下,通过共享环境间接误导已训练好的目标智能体。该方法在StarCraft II的SMAC平台和自动驾驶仿真平台Highway-env上进行评估,实验表明其能在多方开放系统中实现通用且高效的对抗攻击,攻击成功率超过80%。
原文摘要 · Abstract (English)
Reinforcement learning (RL) has been an important machine learning paradigm for solving long-horizon sequential decision-making problems under uncertainty. By integrating deep neural networks (DNNs) into the RL framework, deep reinforcement learning (DRL) has emerged, which achieved significant success in various domains. However, the integration of DNNs also makes it vulnerable to adversarial attacks. Existing adversarial attack techniques mainly focus on either directly manipulating the environment with which a victim agent interacts or deploying an adversarial agent that interacts with the victim agent to induce abnormal behaviors. While these techniques achieve promising results, their adoption in multi-party open systems remains limited due to two major reasons: impractical assumption of full control over the environment and dependent on interactions with victim agents. To enable adversarial attacks in multi-party open systems, in this paper, we redesigned an adversarial policy learning approach that can mislead well-trained victim agents without requiring direct interactions with these agents or full control over their environments. Particularly, we propose a neutral agent-based approach across various task scenarios in multi-party open systems. While the neutral agents seemingly are detached from the victim agents, indirectly influence them through the shared environment. We evaluate our proposed method on the SMAC platform based on Starcraft II and the autonomous driving simulation platform Highway-env. The experimental results demonstrate that our method can launch general and effective adversarial attacks in multi-party open systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。