提升强化学习对动作扰动的鲁棒性,对抗最优攻击者。
Action Robust Reinforcement Learning via Optimal Adversary Aware Policy Optimization
- 通过评估最优对抗策略来优化策略,增强鲁棒性。
- 在多种环境中显著提升对动作扰动的抵抗能力。
- 可无缝集成主流算法,保持原有性能与效率。
强化学习在序列决策任务中取得显著进展,但近期研究揭示其对各类扰动存在脆弱性,影响实际应用中的有效性和安全性。本文聚焦于强化学习策略对动作扰动的鲁棒性,提出一种新框架——最优对抗感知策略迭代(OA-PI)。该框架通过评估并改进策略在对应最优对抗者下的表现,增强其在各种扰动下的鲁棒性。此外,该方法可集成至主流深度强化学习算法,如双延迟深度确定性策略梯度(TD3)和近端策略优化(PPO),在保持原始性能与样本效率的同时,有效提升动作鲁棒性。在多个环境中的实验结果表明,该方法能有效增强深度强化学习策略对不同动作对抗者的抵抗力。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has achieved remarkable success in sequential decision tasks. However, recent studies have revealed the vulnerability of RL policies to different perturbations, raising concerns about their effectiveness and safety in real-world applications. In this work, we focus on the robustness of RL policies against action perturbations and introduce a novel framework called Optimal Adversary-aware Policy Iteration (OA-PI). Our framework enhances action robustness under various perturbations by evaluating and improving policy performance against the corresponding optimal adversaries. Besides, our approach can be integrated into mainstream DRL algorithms such as Twin Delayed DDPG (TD3) and Proximal Policy Optimization (PPO), improving action robustness effectively while maintaining nominal performance and sample efficiency. Experimental results across various environments demonstrate that our method enhances robustness of DRL policies against different action adversaries effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。