用模糊积分提升强化学习在不确定环境下的安全与鲁棒性
Fuz-RL: A Fuzzy-Guided Robust Framework for Safe Reinforcement Learning under Uncertainty
- 引入模糊贝尔曼算子,通过柯赫积分建模多重不确定性
- 理论证明可避免复杂极小极大优化,等价于分布鲁棒安全强化学习
- 在多种不确定性下显著提升控制性能与安全性,兼容现有方法
安全强化学习对实现真实场景中的高性能与高安全性至关重要。然而,现实环境中多重不确定性源的复杂交互给可解释的风险评估与鲁棒决策带来挑战。为此,我们提出Fuz-RL,一种基于模糊测度的鲁棒安全强化学习框架。该框架设计了一种新型模糊贝尔曼算子,利用柯赫积分估计鲁棒价值函数。理论上,我们证明求解Fuz-RL问题(以约束马尔可夫决策过程形式)等价于求解分布鲁棒安全强化学习问题(以鲁棒约束马尔可夫决策过程形式),有效避免了极小极大优化。在safe-control-gym和safety-gymnasium场景的实证分析表明,Fuz-RL能以无模型方式有效融合现有安全强化学习基线,在观测、动作和动态等多种不确定性条件下,显著提升安全性和控制性能。
原文摘要 · Abstract (English)
Safe Reinforcement Learning (RL) is crucial for achieving high performance while ensuring safety in real-world applications. However, the complex interplay of multiple uncertainty sources in real environments poses significant challenges for interpretable risk assessment and robust decision-making. To address these challenges, we propose Fuz-RL, a fuzzy measure-guided robust framework for safe RL. Specifically, our framework develops a novel fuzzy Bellman operator for estimating robust value functions using Choquet integrals. Theoretically, we prove that solving the Fuz-RL problem (in Constrained Markov Decision Process (CMDP) form) is equivalent to solving distributionally robust safe RL problems (in robust CMDP form), effectively avoiding min-max optimization. Empirical analyses on safe-control-gym and safety-gymnasium scenarios demonstrate that Fuz-RL effectively integrates with existing safe RL baselines in a model-free manner, significantly improving both safety and control performance under various types of uncertainties in observation, action, and dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。