arXiv:2510.15720cs.LGcs.AI2025-10

无需环境模型,用风险预算保障强化学习安全决策。

ProSh: Probabilistic Shielding for Model-free Reinforcement Learning

  • 在状态空间中引入风险预算,通过学习成本评价值控制动作安全性。
  • 训练时成本期望有严格上界,依赖于评价值函数的准确性。
  • 适合对安全性要求高的实时决策场景,如自动驾驶、工业控制。

安全性是强化学习中的关键挑战:我们希望开发不仅表现最优,且部署时具有形式化安全保障的RL系统。为此,本文提出无需环境模型的安全强化学习算法ProSh(Probabilistic Shielding via Risk Augmentation),在成本约束下实现安全学习。ProSh通过将风险预算引入约束马尔可夫决策过程(Constrained MDP)的状态空间,并利用学习到的成本评价值对智能体策略分布施加屏蔽,确保所有采样动作在期望上保持安全。当环境为确定性时,该方法保证最优性不受影响。由于ProSh为无模型方法,训练期间的安全性依赖于对环境的经验知识。我们给出了成本期望的紧致上界,仅取决于备份评价值的精度,且在训练过程中始终满足。在温和且实际可达成的假设下,实验表明ProSh可在训练阶段即提供安全保障。

原文摘要 · Abstract (English)

Safety is a major concern in reinforcement learning (RL): we aim at developing RL systems that not only perform optimally, but are also safe to deploy by providing formal guarantees about their safety. To this end, we introduce Probabilistic Shielding via Risk Augmentation (ProSh), a model-free algorithm for safe reinforcement learning under cost constraints. ProSh augments the Constrained MDP state space with a risk budget and enforces safety by applying a shield to the agent's policy distribution using a learned cost critic. The shield ensures that all sampled actions remain safe in expectation. We also show that optimality is preserved when the environment is deterministic. Since ProSh is model-free, safety during training depends on the knowledge we have acquired about the environment. We provide a tight upper-bound on the cost in expectation, depending only on the backup-critic accuracy, that is always satisfied during training. Under mild, practically achievable assumptions, ProSh guarantees safety even at training time, as shown in the experiments.

强化学习安全控制无模型风险预算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。