arXiv:2412.10713cs.LGcs.AI2024-12AAAI被引 22

提出RAT方法,精准操控强化学习智能体行为。

RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors

  • 用意图策略对齐人类偏好,精准定义攻击目标
  • 动态调整经验回放状态分布,提升攻击成功率
  • 在机器人仿真中优于现有攻击方法,可提升模型鲁棒性

评估深度强化学习(DRL)智能体对定向行为攻击的鲁棒性至关重要。这类攻击旨在操纵目标智能体表现出特定行为以契合攻击者目的,常绕过基于奖励的传统防御机制。以往方法多聚焦于降低累积奖励,但奖励过于泛化,难以有效捕捉复杂安全需求,导致攻击策略次优,尤其在安全关键场景中。为此,我们提出RAT,一种面向通用、定向行为攻击的方法。RAT训练一个与人类偏好显式对齐的意图策略,作为攻击者的目标行为;同时,攻击者操纵受试智能体政策以遵循该目标。为增强攻击效果,RAT动态调整经验回放缓冲区中的状态占用度量,实现更可控、高效的操控。在机器人仿真任务上的实验表明,RAT在诱导特定行为方面优于现有对抗攻击算法。此外,RAT还展现出提升智能体鲁棒性的潜力,生成更具韧性策略。我们在多种MuJoCo任务中验证了RAT引导决策变换器智能体采纳人类偏好行为的有效性,证明其在多样化任务中的适用性。

原文摘要 · Abstract (English)

Evaluating deep reinforcement learning (DRL) agents against targeted behavior attacks is critical for assessing their robustness. These attacks aim to manipulate the victim into specific behaviors that align with the attacker's objectives, often bypassing traditional reward-based defenses. Prior methods have primarily focused on reducing cumulative rewards; however, rewards are typically too generic to capture complex safety requirements effectively. As a result, focusing solely on reward reduction can lead to suboptimal attack strategies, particularly in safety-critical scenarios where more precise behavior manipulation is needed. To address these challenges, we propose RAT, a method designed for universal, targeted behavior attacks. RAT trains an intention policy that is explicitly aligned with human preferences, serving as a precise behavioral target for the adversary. Concurrently, an adversary manipulates the victim's policy to follow this target behavior. To enhance the effectiveness of these attacks, RAT dynamically adjusts the state occupancy measure within the replay buffer, allowing for more controlled and effective behavior manipulation. Our empirical results on robotic simulation tasks demonstrate that RAT outperforms existing adversarial attack algorithms in inducing specific behaviors. Additionally, RAT shows promise in improving agent robustness, leading to more resilient policies. We further validate RAT by guiding Decision Transformer agents to adopt behaviors aligned with human preferences in various MuJoCo tasks, demonstrating its effectiveness across diverse tasks.

强化学习对抗攻击行为操控鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。