arXiv:2502.16734cs.LG2025-02被引 2

提出新框架让强化学习在对抗攻击下仍保持最优性能。

Towards Optimal Adversarial Robust Reinforcement Learning with Infinity Measurement Error

  • 构建新马尔可夫决策模型,约束对抗扰动不改变状态本质。
  • 证明无穷测量误差是实现最优鲁棒策略的必要条件。
  • 设计统一框架,在多种算法上提升抗攻击能力且不牺牲正常性能。

确保深度强化学习(DRL)代理在对抗攻击下的鲁棒性对可信部署至关重要。现有研究指出,状态对抗鲁棒性难以实现,且最优鲁棒策略(ORP)未必存在,使严格鲁棒性约束难以落实。本文进一步探索ORP概念,提出内在状态对抗马尔可夫决策过程(ISA-MDP),其中对手无法根本改变状态观测的内在性质。基于实证与理论证据,ISA-MDP普遍刻画了状态对抗范式下的决策行为。我们严格证明,在ISA-MDP中存在确定性、平稳的最优鲁棒策略,其等同于贝尔曼最优策略。理论揭示:提升DRL鲁棒性不一定损害自然环境下的性能。此外,我们证明在$Q$函数和概率空间中,实现ORP需依赖无穷测量误差(IME),揭示了以往依赖1-测量误差的DRL算法的脆弱性。受此启发,我们提出一致对抗鲁棒强化学习(CAR-RL)框架,优化IME的代理目标。该框架适用于基于值和基于策略的DRL算法,显著提升性能并验证了理论分析。

原文摘要 · Abstract (English)

Ensuring the robustness of deep reinforcement learning (DRL) agents against adversarial attacks is critical for their trustworthy deployment. Recent research highlights the challenges of achieving state-adversarial robustness and suggests that an optimal robust policy (ORP) does not always exist, complicating the enforcement of strict robustness constraints. In this paper, we further explore the concept of ORP. We first introduce the Intrinsic State-adversarial Markov Decision Process (ISA-MDP), a novel formulation where adversaries cannot fundamentally alter the intrinsic nature of state observations. ISA-MDP, supported by empirical and theoretical evidence, universally characterizes decision-making under state-adversarial paradigms. We rigorously prove that within ISA-MDP, a deterministic and stationary ORP exists, aligning with the Bellman optimal policy. Our findings theoretically reveal that improving DRL robustness does not necessarily compromise performance in natural environments. Furthermore, we demonstrate the necessity of infinity measurement error (IME) in both $Q$-function and probability spaces to achieve ORP, unveiling vulnerabilities of previous DRL algorithms that rely on $1$-measurement errors. Motivated by these insights, we develop the Consistent Adversarial Robust Reinforcement Learning (CAR-RL) framework, which optimizes surrogates of IME. We apply CAR-RL to both value-based and policy-based DRL algorithms, achieving superior performance and validating our theoretical analysis.

强化学习对抗鲁棒最优策略决策建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。