arXiv:2409.00418cs.LGcs.AI2024-09被引 1

用分布约束提升强化学习抗干扰能力,支持高效离线训练。

Robust off-policy Reinforcement Learning via Soft Constrained Adversary

  • 基于先验分布的f散度约束,设计新型对抗攻击
  • 在多种环境上实现高样本效率的离线强化学习
  • 适合需要鲁棒性的实际部署场景

近期,针对输入观测扰动的鲁棒强化学习方法受到广泛关注,因强化学习存在潜在脆弱性。尽管现有方法取得一定成效,但在长时程规划中仍存在两大局限:其一,策略与其对应最优对抗者之间存在相互依赖,限制了离线强化学习算法的发展;虽最优对抗者应依赖当前策略,但此机制阻碍了离线应用。其二,现有方法通常仅基于$L_p$-范数假设扰动,未利用环境中已知的扰动分布先验。本文提出一种新视角:将对抗强化学习建模为带有先验分布的f散度约束问题,并由此推导出两种典型攻击及其对应的鲁棒学习框架。通过评估鲁棒性,结果表明所提方法在样本高效离线强化学习中表现优异。

原文摘要 · Abstract (English)

Recently, robust reinforcement learning (RL) methods against input observation have garnered significant attention and undergone rapid evolution due to RL's potential vulnerability. Although these advanced methods have achieved reasonable success, there have been two limitations when considering adversary in terms of long-term horizons. First, the mutual dependency between the policy and its corresponding optimal adversary limits the development of off-policy RL algorithms; although obtaining optimal adversary should depend on the current policy, this has restricted applications to off-policy RL. Second, these methods generally assume perturbations based only on the $L_p$-norm, even when prior knowledge of the perturbation distribution in the environment is available. We here introduce another perspective on adversarial RL: an f-divergence constrained problem with the prior knowledge distribution. From this, we derive two typical attacks and their corresponding robust learning frameworks. The evaluation of robustness is conducted and the results demonstrate that our proposed methods achieve excellent performance in sample-efficient off-policy RL.

强化学习对抗训练离线学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。