提出新算法,让强化学习在未知参数下也能全局收敛。
Global Convergence of Policy Gradient for Entropy Regularized Linear-Quadratic Control with Multiplicative Noise
- 用熵正则化改进策略梯度,解决非凸优化难题。
- 新算法无需系统参数知识,仍保证全局收敛。
- 适合研究强化学习控制与理论保障的学者。
强化学习(RL)已成为动态环境中序列决策的强大框架,尤其适用于系统参数未知的情况。本文研究了具有乘性噪声的熵正则化线性-二次(LQ)控制问题在无限时域下的基于RL的控制方法。首先,将正则化策略梯度(RPG)算法适配至随机最优控制场景,并证明在梯度支配和近光滑条件下,尽管问题非凸,RPG仍可实现全局收敛。其次,基于零阶优化思想,提出一种新型无模型强化学习算法:基于样本的正则化策略梯度(SB-RPG)。该算法不依赖系统参数信息,但仍保持强理论保证的全局收敛性。模型通过熵正则化缓解了强化学习中探索与利用的权衡问题。数值模拟验证了理论结果,并展示了SB-RPG在未知参数环境中的高效性。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has emerged as a powerful framework for sequential decision-making in dynamic environments, particularly when system parameters are unknown. This paper investigates RL-based control for entropy-regularized linear-quadratic (LQ) control problems with multiplicative noise over an infinite time horizon. First, we adapt the regularized policy gradient (RPG) algorithm to stochastic optimal control settings, proving that despite the non-convexity of the problem, RPG converges globally under conditions of gradient domination and almost-smoothness. Second, based on zero-order optimization approach, we introduce a novel model free RL algorithm: Sample-based regularized policy gradient (SB-RPG). SB-RPG operates without knowledge of system parameters yet still retains strong theoretical guarantees of global convergence. Our model leverages entropy regularization to address the exploration versus exploitation trade-off inherent in RL. Numerical simulations validate the theoretical results and demonstrate the efficiency of SB-RPG in unknown-parameters environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。