arXiv:2602.05089cs.CRcs.LG2026-02被引 3

不安全的模拟器可暗中植入强化学习后门,触发后操控智能体行为。

Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement Learning

  • 利用模拟器动态缺陷,在不修改奖励的情况下植入动作级后门。
  • 攻击在离散与连续动作空间任务中均成功激活目标动作,成功率高。
  • 首次实现后门从仿真到真实机器人硬件的迁移,警示安全风险。

模拟环境是强化学习成功的关键,使研究者可在无真实硬件成本下训练决策智能体。然而,模拟器仍存在安全盲点,恶意开发者可通过篡改其动态实现攻击。本文揭示一种新型威胁:通过操纵模拟器动态,隐蔽植入针对动作层面的后门。该后门可在检测到预设“触发信号”时,可靠激活特定动作,导致潜在危险后果。传统后门攻击依赖攻击者对训练流程的完全控制并可观察奖励,但在模拟器场景中难以实现。为此,本文提出新攻击方法 Daze,可在不修改或观测奖励的前提下,成功植入后门。我们提供了 Daze 在通用强化学习任务中保证攻击成功的形式化证明,并在离散与连续动作空间环境中进行了广泛实证评估。此外,首次展示了后门从仿真环境向真实机器人硬件的迁移实例。这些成果呼吁对强化学习训练全链路加强安全防护。

原文摘要 · Abstract (English)

Simulated environments are a key piece in the success of Reinforcement Learning (RL), allowing practitioners and researchers to train decision making agents without running expensive experiments on real hardware. Simulators remain a security blind spot, however, enabling adversarial developers to alter the dynamics of their released simulators for malicious purposes. Therefore, in this work we highlight a novel threat, demonstrating how simulator dynamics can be exploited to stealthily implant action-level backdoors into RL agents. The backdoor then allows an adversary to reliably activate targeted actions in an agent upon observing a predefined ``trigger'', leading to potentially dangerous consequences. Traditional backdoor attacks are limited in their strong threat models, assuming the adversary has near full control over an agent's training pipeline, enabling them to both alter and observe agent's rewards. As these assumptions are infeasible to implement within a simulator, we propose a new attack ``Daze'' which is able to reliably and stealthily implant backdoors into RL agents trained for real world tasks without altering or even observing their rewards. We provide formal proof of Daze's effectiveness in guaranteeing attack success across general RL tasks along with extensive empirical evaluations on both discrete and continuous action space domains. We additionally provide the first example of RL backdoor attacks transferring to real, robotic hardware. These developments motivate further research into securing all components of the RL training pipeline to prevent malicious attacks.

强化学习后门攻击模拟器安全鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。