arXiv:2604.26836cs.LGcs.SY2026-04被引 2

用概率神经网络提升强化学习安全探索,防止模型欺骗。

Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics

论文配图:Uncertainty-Aware Predictive Safety Filters for Probabilistic Neural Network Dynamics
图 1 · 摘自论文原文
  • 将概率集成神经网络用于安全过滤,构建未来状态可达集
  • 在标准安全强化学习基准上显著提升探索安全性
  • 适合需要高安全性保障的实时控制场景

预测性安全过滤器(PSF)利用模型预测控制在深度强化学习探索中保证约束满足,但其依赖第一性原理模型或高斯过程,限制了可扩展性和通用性。而基于模型的强化学习(MBRL)方法通常使用概率集成(PE)神经网络,从数据中捕捉复杂高维动态,且对先验知识需求极少。然而,现有将PE引入PSF的方法缺乏严格的不确定性量化。本文提出不确定性感知的预测性安全过滤器(UPSi),通过将未来结果建模为可达集,实现基于PE动态模型的严格安全验证。UPSi引入显式的确定性约束,防止模型被恶意利用,并可无缝集成至常见MBRL框架。我们在基于Dyna风格的MBRL中,在标准安全强化学习基准上评估了UPSi,结果表明其在探索安全性上相比先前神经网络型PSF有显著提升,同时性能与标准MBRL相当。UPSi弥合了现代MBRL的可扩展性与通用性,以及预测性安全过滤器的安全保障之间的差距。

原文摘要 · Abstract (English)

Predictive safety filters (PSFs) leverage model predictive control to enforce constraint satisfaction during deep reinforcement learning (RL) exploration, yet their reliance on first-principles models or Gaussian processes limits scalability and broader applicability. Meanwhile, model-based RL (MBRL) methods routinely employ probabilistic ensemble (PE) neural networks to capture complex, high-dimensional dynamics from data with minimal prior knowledge. However, existing attempts to integrate PEs into PSFs lack rigorous uncertainty quantification. We introduce the Uncertainty-Aware Predictive Safety Filter (UPSi), a PSF that provides rigorous safety verification using PE dynamics models by formulating future outcomes as reachable sets. UPSi introduces an explicit certainty constraint that prevents model exploitation and integrates seamlessly into common MBRL frameworks. We evaluate UPSi under practical simplifications within Dyna-style MBRL on standard safe RL benchmarks and report substantial improvements in exploration safety over prior neural network PSFs while maintaining performance on par with standard MBRL. UPSi bridges the gap between the scalability and generality of modern MBRL and the safety guarantees of predictive safety filters.

强化学习安全控制神经网络不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。