提出新算法让强化学习在不确定环境下更安全可靠
Rectified Robust Policy Optimization for Model-Uncertain Constrained Reinforcement Learning without Strong Duality
- 仅用原始优化问题,不依赖传统对偶方法
- 理论证明可收敛到近优可行策略,复杂度达最优下界
- 适合需要严格安全约束的机器人控制等场景
鲁棒约束强化学习的目标是在最坏情况模型不确定性下优化智能体性能,同时满足安全或资源约束。本文证明,鲁棒约束RL中通常不成立强对偶性,这意味着传统原-对偶方法可能无法找到最优可行策略。为此,我们提出一种新型纯原问题算法——修正鲁棒策略优化(RRPO),直接在原问题上求解,无需依赖对偶形式。在温和正则性假设下,我们提供了理论收敛保证,证明其可收敛至近似最优可行策略,且迭代复杂度在不确定性集直径受控时达到已知最优下界。网格世界环境中的实验验证了该方法的有效性:在模型不确定性下,RRPO能保持鲁棒且安全的表现,而非鲁棒方法则可能违反最坏情况下的安全约束。
原文摘要 · Abstract (English)
The goal of robust constrained reinforcement learning (RL) is to optimize an agent's performance under the worst-case model uncertainty while satisfying safety or resource constraints. In this paper, we demonstrate that strong duality does not generally hold in robust constrained RL, indicating that traditional primal-dual methods may fail to find optimal feasible policies. To overcome this limitation, we propose a novel primal-only algorithm called Rectified Robust Policy Optimization (RRPO), which operates directly on the primal problem without relying on dual formulations. We provide theoretical convergence guarantees under mild regularity assumptions, showing convergence to an approximately optimal feasible policy with iteration complexity matching the best-known lower bound when the uncertainty set diameter is controlled in a specific level. Empirical results in a grid-world environment validate the effectiveness of our approach, demonstrating that RRPO achieves robust and safe performance under model uncertainties while the non-robust method can violate the worst-case safety constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。