arXiv:2505.04494math.OCcs.LG2025-05

提出新算法,用双时间尺度优化提升强化学习的离线数据利用效率。

A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance

  • 基于双时间尺度的投影梯度下降-上升框架,结合在线对偶变量引导。
  • 几乎必然收敛到正则化MDP最优值函数和策略,支持无模拟器运行。
  • 适合研究强化学习收敛性或需高效利用离线数据的工程场景。

我们通过结合正则化线性规划方法与随机逼近理论,研究强化学习。针对如何在利用离线数据的同时保持在线探索的问题,提出PGDA-RL——一种求解正则化马尔可夫决策过程的新型投影梯度下降-上升算法。该算法将基于经验回放的梯度估计与嵌套优化问题的双时间尺度分解相结合,异步运行,仅需单条相关数据轨迹与环境交互,并根据底层MDP状态占用量的对偶变量在线更新策略。我们证明了该算法几乎必然收敛至正则化MDP的最优值函数和策略。分析基于随机逼近工具,相较于现有方法,假设条件更弱,无需仿真器或固定行为策略。在加强遍历性假设下,建立了最后迭代的有限时间保证,均方收敛速率达$ ilde{O}(k^{-2/3})$,与马尔可夫采样和有偏梯度估计下的最优已知速率一致。

原文摘要 · Abstract (English)

We study reinforcement learning by combining recent advances in regularized linear programming formulations with the classical theory of stochastic approximation. Motivated by the challenge of designing algorithms that leverage off-policy data while maintaining on-policy exploration, we propose PGDA-RL, a novel primal-dual Projected Gradient Descent-Ascent algorithm for solving regularized Markov Decision Processes (MDPs). PGDA-RL integrates experience replay-based gradient estimation with a two-timescale decomposition of the underlying nested optimization problem. The algorithm operates asynchronously, interacts with the environment through a single trajectory of correlated data, and updates its policy online in response to the dual variable associated with the occupancy measure of the underlying MDP. We prove that PGDA-RL converges almost surely to the optimal value function and policy of the regularized MDP. Our convergence analysis relies on tools from stochastic approximation theory and holds under weaker assumptions than those required by existing primal-dual RL approaches, notably removing the need for a simulator or a fixed behavioral policy. Under a strengthened ergodicity assumption on the underlying Markov chain, we establish a last-iterate finite-time guarantee with $\tilde{O} (k^{-2/3})$ mean-square convergence, aligning with the best-known rates for two-timescale stochastic approximation methods under Markovian sampling and biased gradient estimates.

强化学习双时间尺度收敛性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。