首次证明连续时间强化学习中策略迁移的理论有效性。
Policy Transfer for Continuous-Time Reinforcement Learning: A (Rough) Differential Equation Approach
- 基于微分方程与粗糙路径理论,建立策略迁移的稳定性分析框架。
- 在非线性有界系统中仍保证收敛速度不下降,实现近优策略初始化。
- 适用于连续时间LQR与基于得分的扩散模型,适合强化学习研究者。
本文研究连续时间强化学习中的策略迁移问题。针对具有香农熵正则化的连续时间线性二次型系统,充分利用最优策略的高斯结构和相关里卡蒂方程的稳定性。在一般情形下,系统可能具有非线性且有界的动力学,关键技术在于通过粗糙路径理论证明扩散随机微分方程(SDE)的稳定性。本工作首次提供了连续时间强化学习中策略迁移的理论证明:一个任务上学到的最优策略可用于初始化另一个密切相关任务的搜索,且不降低原始算法的收敛速率。作为分析副产品,我们通过与线性二次型调节器(LQR)的关联,推导出一类具体连续时间得分驱动扩散模型的稳定性。为展示策略迁移的优势,我们提出一种新的连续时间LQR策略学习算法,实现了全局线性收敛和局部超线性收敛。
原文摘要 · Abstract (English)
This paper studies policy transfer, one of the well-known transfer learning techniques adopted in large language models, for continuous-time reinforcement learning problems. In the case of continuous-time linear-quadratic systems with Shannon's entropy regularization, we fully exploit the Gaussian structure of their optimal policy and the stability of their associated Riccati equations. In the general case where the system has possibly non-linear and bounded dynamics, the key technical component is the stability of diffusion SDEs which is established by invoking the rough path theory. Our work provides the first theoretical proof of policy transfer for continuous-time RL: an optimal policy learned for one RL problem can be used to initialize to search for a near-optimal policy for another closely related RL problem, while achieving (at least) the same rate of convergence for the original algorithm. As a byproduct of our analysis, we derive the stability of a concrete class of continuous-time score-based diffusion models via their connection with LQRs. To illustrate the benefit of policy transfer for RL, we propose a novel policy learning algorithm for continuous-time LQRs, which achieves global linear convergence and local super-linear convergence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。