发现深度强化学习系统在组件级和训练后阶段的隐蔽后门漏洞
Beyond Training-time Poisoning: Component-level and Post-training Backdoors in Deep Reinforcement Learning
- 通过组件级漏洞植入可逃过重训练的持久后门
- 无需训练数据即可实现训练后攻击,成功率与已有方法相当
- 可绕过主流防御机制,适合安全研究者关注
深度强化学习(DRL)系统日益应用于安全关键场景,但其安全性仍严重缺乏研究。本文研究后门攻击——在观测空间中植入特定触发器,仅在特定输入出现时引发恶意行为。现有研究仅关注需访问训练管道的训练期攻击,而本文揭示了整个DRL供应链中的新漏洞:攻击者可在更低权限下植入后门。提出两种新攻击:(1) TrojanentRL,利用组件级缺陷植入可抵抗全模型重训练的持久后门;(2) InfrectroRL,一种无需训练、验证或测试数据的训练后攻击。在六个Atari环境上的实验表明,该攻击在更严格约束下达到与现有最先进训练期攻击相当的效果,且能绕过两种主流防御机制。这些发现挑战当前研究范式,凸显构建鲁棒防御的紧迫性。
原文摘要 · Abstract (English)
Deep Reinforcement Learning (DRL) systems are increasingly used in safety-critical applications, yet their security remains severely underexplored. This work investigates backdoor attacks, which implant hidden triggers that cause malicious actions only when specific inputs appear in the observation space. Existing DRL backdoor research focuses solely on training-time attacks requiring unrealistic access to the training pipeline. In contrast, we reveal critical vulnerabilities across the DRL supply chain where backdoors can be embedded with significantly reduced adversarial privileges. We introduce two novel attacks: (1) TrojanentRL, which exploits component-level flaws to implant a persistent backdoor that survives full model retraining; and (2) InfrectroRL, a post-training backdoor attack which requires no access to training, validation, nor test data. Empirical and analytical evaluations across six Atari environments show our attacks rival state-of-the-art training-time backdoor attacks while operating under much stricter adversarial constraints. We also demonstrate that InfrectroRL further evades two leading DRL backdoor defenses. These findings challenge the current research focus and highlight the urgent need for robust defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。