arXiv:2601.05152cs.LGcs.AI2026-01综述

系统梳理非平稳环境下持续安全强化学习的前沿方法与挑战。

Safe Continual Reinforcement Learning Methods for Nonstationary Environments. Towards a Survey of the State of the Art

  • 按安全机制分类,归纳适应非平稳性的持续在线安全强化学习方法
  • 提出安全约束在分布偏移下的形式化框架,支持可靠持续学习
  • 适合关注安全强化学习、持续学习与鲁棒决策的研究者参考

本文系统综述了非平稳环境下持续安全在线强化学习(COSRL)的最新进展。探讨了构建持续在线安全强化学习算法中的理论基础、核心挑战与未解问题。基于安全学习机制对非平稳性适应能力的不同,提出方法分类体系,并对在线强化学习中安全约束的建模方式进行归纳。进一步分析了隐藏马尔可夫决策过程(HM-MDP)、非平稳马尔可夫决策过程(NSMDP)、部分可观测马尔可夫决策过程(POMDP)及其安全版本中的安全约束设计。最后讨论了未来构建可信、安全在线学习算法的发展方向。

原文摘要 · Abstract (English)

This work provides a state-of-the-art survey of continual safe online reinforcement learning (COSRL) methods. We discuss theoretical aspects, challenges, and open questions in building continual online safe reinforcement learning algorithms. We provide the taxonomy and the details of continual online safe reinforcement learning methods based on the type of safe learning mechanism that takes adaptation to nonstationarity into account. We categorize safety constraints formulation for online reinforcement learning algorithms, and finally, we discuss prospects for creating reliable, safe online learning algorithms. Keywords: safe RL in nonstationary environments, safe continual reinforcement learning under nonstationarity, HM-MDP, NSMDP, POMDP, safe POMDP, constraints for continual learning, safe continual reinforcement learning review, safe continual reinforcement learning survey, safe continual reinforcement learning, safe online learning under distribution shift, safe continual online adaptation, safe reinforcement learning, safe exploration, safe adaptation, constrained Markov decision processes, safe reinforcement learning, partially observable Markov decision process, safe reinforcement learning and hidden Markov decision processes, Safe Online Reinforcement Learning, safe online reinforcement learning, safe online reinforcement learning, safe meta-learning, safe meta-reinforcement learning, safe context-based reinforcement learning, formulating safety constraints for continual learning

安全强化学习持续学习非平稳环境在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。