提出安全约束下强化学习的凸正则化框架,保障高维系统的稳定训练。
Convex Regularization and Convergence of Policy Gradient Flows under Safety Constraints
- 用双重正则化构建可解的凸优化目标,融合奖励与参数正则
- 在足够正则化下实现策略梯度流指数收敛,理论保证安全稳定
- 适用于自动驾驶、金融等高风险场景,适合研究安全RL的学者
本文研究无限时域决策过程中的强化学习(RL),在几乎必然的安全约束下,对自动驾驶、金融和资源管理等应用至关重要。提出一种双重正则化的RL框架,结合奖励与参数正则化,以处理连续状态-动作空间中的安全约束。问题被形式化为均场范围内参数化策略的凸正则化目标。利用均场理论与Wasserstein梯度流,将策略建模于无穷维统计流形上,更新由参数分布梯度流驱动。主要贡献包括安全约束问题的可解性条件、梯度流的光滑有界逼近,以及在充分正则化下的指数收敛保证。通用正则化条件(如熵正则化)支持实际粒子方法实现。该框架为复杂高维环境中的安全强化学习提供了坚实的理论洞察与保障。
原文摘要 · Abstract (English)
This paper examines reinforcement learning (RL) in infinite-horizon decision processes with almost-sure safety constraints, crucial for applications like autonomous systems, finance, and resource management. We propose a doubly-regularized RL framework combining reward and parameter regularization to address safety constraints in continuous state-action spaces. The problem is formulated as a convex regularized objective with parametrized policies in the mean-field regime. Leveraging mean-field theory and Wasserstein gradient flows, policies are modeled on an infinite-dimensional statistical manifold, with updates governed by parameter distribution gradient flows. Key contributions include solvability conditions for safety-constrained problems, smooth bounded approximations for gradient flows, and exponential convergence guarantees under sufficient regularization. General regularization conditions, including entropy regularization, support practical particle method implementations. This framework provides robust theoretical insights and guarantees for safe RL in complex, high-dimensional settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。