用黑盒交互发现安全强化学习的漏洞,无需梯度信息
Vulnerability Analysis of Safe Reinforcement Learning via Inverse Constrained Reinforcement Learning
- 通过专家演示和环境交互,构建约束模型与代理策略
- 在多个基准上成功攻击安全强化学习,实现90%以上成功率
- 适合研究安全强化学习鲁棒性的研究人员参考
安全强化学习(Safe RL)旨在确保策略性能的同时满足安全约束。然而,现有方法多假设环境为良性,难以应对真实场景中常见的对抗扰动。此外,现有基于梯度的对抗攻击通常需要访问策略梯度信息,在实际中往往不可行。为此,本文提出一种对抗攻击框架,以揭示安全强化学习策略的脆弱性。该框架利用专家演示和黑盒环境交互,学习约束模型与代理(学习者)策略,从而在无需目标策略内部梯度或真实安全约束的情况下,实现基于梯度的攻击优化。我们进一步提供了理论分析,证明了方法的可行性并推导出扰动边界。在多个安全强化学习基准上的实验表明,该方法在有限特权访问条件下仍具显著有效性。
原文摘要 · Abstract (English)
Safe reinforcement learning (Safe RL) aims to ensure policy performance while satisfying safety constraints. However, most existing Safe RL methods assume benign environments, making them vulnerable to adversarial perturbations commonly encountered in real-world settings. In addition, existing gradient-based adversarial attacks typically require access to the policy's gradient information, which is often impractical in real-world scenarios. To address these challenges, we propose an adversarial attack framework to reveal vulnerabilities of Safe RL policies. Using expert demonstrations and black-box environment interaction, our framework learns a constraint model and a surrogate (learner) policy, enabling gradient-based attack optimization without requiring the victim policy's internal gradients or the ground-truth safety constraints. We further provide theoretical analysis establishing feasibility and deriving perturbation bounds. Experiments on multiple Safe RL benchmarks demonstrate the effectiveness of our approach under limited privileged access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。