arXiv:2505.12354cs.LGcs.AI2025-05被引 1

为强化学习设计可保证目标达成的通用安全包装器

A universal policy wrapper with guarantees

  • 用基线策略与保障策略交替,由价值函数控制切换
  • 保持基线性能的同时确保目标可达性,理论可证明
  • 无需额外系统知识,适配各类RL算法和任务

我们提出一种通用强化学习策略包装器,可提供形式化的目标达成保证。与通常表现优异但缺乏安全保证的标准RL算法不同,该包装器在高性能基线策略(来自任意现有RL方法)和具有已知收敛性质的备用策略之间进行选择性切换。基线策略的价值函数监督切换过程,决定何时由备用策略接管以维持系统稳定。分析表明,该包装器继承了备用策略的目标达成保证,同时保持或提升了基线策略的性能。值得注意的是,其无需额外系统知识或在线约束优化,可直接部署于多种RL架构和任务中。

原文摘要 · Abstract (English)

We introduce a universal policy wrapper for reinforcement learning agents that ensures formal goal-reaching guarantees. In contrast to standard reinforcement learning algorithms that excel in performance but lack rigorous safety assurances, our wrapper selectively switches between a high-performing base policy -- derived from any existing RL method -- and a fallback policy with known convergence properties. Base policy's value function supervises this switching process, determining when the fallback policy should override the base policy to ensure the system remains on a stable path. The analysis proves that our wrapper inherits the fallback policy's goal-reaching guarantees while preserving or improving upon the performance of the base policy. Notably, it operates without needing additional system knowledge or online constrained optimization, making it readily deployable across diverse reinforcement learning architectures and tasks.

强化学习安全保障策略包装可证明性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。