研究深度强化学习中隐蔽的后门攻击,发现自然数据触发更难防且真实威胁大。
Backdoors in DRL: Four Environments Focusing on In-distribution Triggers
- 在四种强化学习环境中植入自然分布内的后门触发器。
- 基础数据污染即可实现有效后门攻击,模型仍会响应恶意指令。
- 适合关注AI安全、模型鲁棒性的研究人员和开发者看。
后门攻击(或称木马攻击)通过在深度神经网络模型中隐藏不良行为构成安全风险。开源神经网络每日被下载,可能携带后门,第三方模型开发者普遍存在。为推进后门攻击缓解研究,我们开发了针对深度强化学习(DRL)代理的多种木马。重点研究分布内触发器(in-distribution triggers),因其在模型部署时易被攻击者激活,相比分布外触发器更具威胁。我们在四个强化学习环境(LavaWorld、Randomized LavaWorld、Colorful Memory、Modified Safety Gymnasium)中实现后门攻击,并训练多种干净与带毒模型以表征这些攻击。结果表明,分布内触发器实现难度更高、模型学习更困难,但仍可在基本数据投毒攻击下成功激活,构成对DRL的实际威胁。
原文摘要 · Abstract (English)
Backdoor attacks, or trojans, pose a security risk by concealing undesirable behavior in deep neural network models. Open-source neural networks are downloaded from the internet daily, possibly containing backdoors, and third-party model developers are common. To advance research on backdoor attack mitigation, we develop several trojans for deep reinforcement learning (DRL) agents. We focus on in-distribution triggers, which occur within the agent's natural data distribution, since they pose a more significant security threat than out-of-distribution triggers due to their ease of activation by the attacker during model deployment. We implement backdoor attacks in four reinforcement learning (RL) environments: LavaWorld, Randomized LavaWorld, Colorful Memory, and Modified Safety Gymnasium. We train various models, both clean and backdoored, to characterize these attacks. We find that in-distribution triggers can require additional effort to implement and be more challenging for models to learn, but are nevertheless viable threats in DRL even using basic data poisoning attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。