arXiv:2506.12815cs.LG2025-06被引 1

首次实现对轨迹优化模型的动作级后门攻击,隐蔽性强且成本低。

TrojanTO: Action-Level Backdoor Attacks against Trajectory Optimization Models

  • 通过交替训练增强触发器与目标动作的关联性
  • 仅用0.3%轨迹即可成功植入后门,攻击效果显著
  • 适用于多种轨迹优化模型,适合安全研究者关注

轨迹优化(TO)模型在离线强化学习中取得显著进展,但其对后门攻击的脆弱性尚未被充分理解。现有强化学习中的后门攻击多依赖奖励操纵,因TO模型固有的序列建模特性而效果有限,高维动作空间也进一步增加了动作操控的难度。为此,我们提出TrojanTO,首个针对TO模型的动作级后门攻击方法。该方法采用交替训练提升触发器与目标动作间的关联性;为增强隐蔽性,结合轨迹过滤实现精准污染以维持正常性能,同时使用批量污染确保触发一致性。大量实验表明,TrojanTO可在多种任务和攻击目标下有效植入后门,攻击预算极低(仅需0.3%轨迹)。此外,该方法在DT、GDT和DC等多种TO模型架构上均表现良好,展现出强可扩展性。

原文摘要 · Abstract (English)

Recent advances in Trajectory Optimization (TO) models have achieved remarkable success in offline reinforcement learning. However, their vulnerabilities against backdoor attacks are poorly understood. We find that existing backdoor attacks in reinforcement learning are based on reward manipulation, which are largely ineffective against the TO model due to its inherent sequence modeling nature. Moreover, the complexities introduced by high-dimensional action spaces further compound the challenge of action manipulation. To address these gaps, we propose TrojanTO, the first action-level backdoor attack against TO models. TrojanTO employs alternating training to enhance the connection between triggers and target actions for attack effectiveness. To improve attack stealth, it utilizes precise poisoning via trajectory filtering for normal performance and batch poisoning for trigger consistency. Extensive evaluations demonstrate that TrojanTO effectively implants backdoor attacks across diverse tasks and attack objectives with a low attack budget (0.3\% of trajectories). Furthermore, TrojanTO exhibits broad applicability to DT, GDT, and DC, underscoring its scalability across diverse TO model architectures.

后门攻击轨迹优化强化学习安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。