arXiv:2605.10293cs.LGcs.AI2026-05被引 1

为离线强化学习设计安全概率屏障,确保策略改进既高效又安全。

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning

  • 基于数据集和安全状态知识,动态限制动作空间以保障安全。
  • 在低数据场景下,平均与最差性能均显著优于无屏障方法。
  • 适合对安全性要求高的离线强化学习应用,如医疗或自动驾驶。

在离线强化学习中,我们从固定数据集学习策略而无需环境交互。主要挑战在于提供策略的(1)性能保证和(2)安全性保证。安全策略改进(SPI)技术可提供性能保证:以高概率确保新策略优于给定基线策略(假设其安全)。另一方面,在安全强化学习中,屏障通过限制动作空间至经证明安全的动作,提供安全保障。本文将两者结合,将屏障扩展至离线强化学习,仅依赖可用数据集及对安全与非安全状态的知识。通过在策略改进步骤中施加屏障,以高概率保证策略的安全性。实验结果表明,受屏障保护的SPI优于无屏障版本,尤其在低数据条件下,显著提升平均与最差性能。

原文摘要 · Abstract (English)

In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide guarantees on the (1) performance and (2) safety of the resulting policy. A technique called safe policy improvement (SPI) provides a performance guarantee: with high probability, the new policy outperforms a given baseline policy, which is assumed to be safe. Orthogonally, in the context of safe RL, a shield provides a safety guarantee by restricting the action space to those actions that are provably safe with respect to a given safety-relevant model. We integrate these paradigms by extending shielding to offline RL, relying solely on the available dataset and knowledge of safe and unsafe states. Then, we shield the policy improvement steps, guaranteeing, with high probability, a safe policy. Experimental results demonstrate that shielded SPI outperforms its unshielded counterpart, improving both average and worst-case performance, particularly in low-data regimes.

离线RL安全强化学习策略改进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。