arXiv:2603.17875cs.LGmath.OC2026-03被引 1

从算子理论出发,为广义强化学习提供新优化框架与高效算法。

Operator-Theoretic Foundations and Policy Gradient Methods for General MDPs with Unbounded Costs

  • 基于线性算子与摄动理论,建立通用策略差分关系。
  • 提出新型梯度算法MM-RKHS,计算与样本复杂度均低于PPO。
  • 适用于一般状态动作空间,可推广多种经典强化学习算法。

马尔可夫决策过程(MDPs)被视为在函数空间上对特定线性算子的优化问题。本文建立了广义MDPs中最优策略存在性的新结果,不同于已有文献。利用线性算子的摄动理论,推导出通用MDPs中的策略差分引理及目标函数关于策略算子的Gâteaux导数。通过积分概率度量理论对策略差分进行上界控制,提出一种新型的极大化-极小化型策略梯度算法。该方法将诸多经典强化学习算法推广至一般状态与动作空间。进一步地,当采用最大均值差异作为积分概率度量时,导出了针对有限MDPs的低复杂度策略梯度算法——MM-RKHS。实验表明,该算法在计算复杂度、样本复杂度和收敛速度上均优于PPO。

原文摘要 · Abstract (English)

Markov decision processes (MDPs) is viewed as an optimization of an objective function over certain linear operators over general function spaces. A new existence result is established for the existence of optimal policies in general MDPs, which differs from the existence result derived previously in the literature. Using the well-established perturbation theory of linear operators, policy difference lemma is established for general MDPs and the Gauteaux derivative of the objective function as a function of the policy operator is derived. By upper bounding the policy difference via the theory of integral probability metric, a new majorization-minimization type policy gradient algorithm for general MDPs is derived. This leads to generalization of many well-known algorithms in reinforcement learning to cases with general state and action spaces. Further, by taking the integral probability metric as maximum mean discrepancy, a low-complexity policy gradient algorithm is derived for finite MDPs. The new algorithm, called MM-RKHS, appears to be superior to PPO algorithm due to low computational complexity, low sample complexity, and faster convergence.

强化学习策略梯度算子理论优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。