用扩散模型生成语义级扰动,突破现有攻击局限。
Diffusion Guided Adversarial State Perturbations in Reinforcement Learning
- 基于扩散模型生成语义差异大但视觉真实的干扰状态
- 在不被检测前提下使强化学习智能体行为严重偏离
- 适合研究鲁棒性、对抗攻防的学者和工程师
强化学习系统虽在多个领域取得显著成果,但在视觉输入环境下极易受到对抗攻击。当前主流防御方法在较大状态扰动下仍能保持鲁棒性,但经深入分析发现,其有效性源于现有 $l_p$ 范数约束攻击的根本缺陷:即使在较大的扰动预算下,也无法有效改变图像输入的语义。为此,本文提出 SHIFT——一种无需依赖策略的扩散模型驱动的状态扰动攻击方法。该攻击可生成与真实状态语义显著不同、但仍保持视觉真实性和历史一致性的扰动状态,从而规避检测。实验表明,该攻击能有效突破现有各类防御机制,包括最先进的防御方案,在性能上显著优于现有攻击,同时更具感知隐蔽性。结果揭示了强化学习智能体对语义感知型对抗扰动的高度脆弱性,凸显了开发更鲁棒策略的重要性。
原文摘要 · Abstract (English)
Reinforcement learning (RL) systems, while achieving remarkable success across various domains, are vulnerable to adversarial attacks. This is especially a concern in vision-based environments where minor manipulations of high-dimensional image inputs can easily mislead the agent's behavior. To this end, various defenses have been proposed recently, with state-of-the-art approaches achieving robust performance even under large state perturbations. However, after closer investigation, we found that the effectiveness of the current defenses is due to a fundamental weakness of the existing $l_p$ norm-constrained attacks, which can barely alter the semantics of image input even under a relatively large perturbation budget. In this work, we propose SHIFT, a novel policy-agnostic diffusion-based state perturbation attack to go beyond this limitation. Our attack is able to generate perturbed states that are semantically different from the true states while remaining realistic and history-aligned to avoid detection. Evaluations show that our attack effectively breaks existing defenses, including the most sophisticated ones, significantly outperforming existing attacks while being more perceptually stealthy. The results highlight the vulnerability of RL agents to semantics-aware adversarial perturbations, indicating the importance of developing more robust policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。