用数学工具统一梳理强化学习的核心算法原理
Mathematical methods of reinforcement learning
- 基于泛函分析与优化理论,系统解析值迭代等算法的收敛性
- 建立有限样本与渐近结果的数学桥梁,涵盖函数逼近与离线评估
- 适合概率统计、优化方向研究者快速切入强化学习数学基础
强化学习日益依托概率论、优化和算子理论的工具。本文系统梳理现代强化学习算法设计与分析背后的数学结构。从马尔可夫决策过程(MDPs)与贝尔曼算子出发,强调压缩映射、单调性与不动点理论对值迭代、策略迭代及时序差分方法的收敛性与收敛速率保障。随后引入优化视角:随机逼近与鞅方法,凸对偶性与正则化在镜像/近端方法中的作用。函数逼近部分涵盖线性与非线性情形,通过依赖数据与混合过程的浓度不等式,分析稳定性、误差分解与样本复杂度。进一步讨论离线评估/学习、约束强化学习与约束马尔可夫决策过程(CMDPs)。全篇以共同的算子与变分视角统一算法模板,突出有限样本界与渐近结果。本综述旨在为概率、优化与统计领域的研究者提供进入强化学习的统一数学入口。
原文摘要 · Abstract (English)
Reinforcement learning (RL) is increasingly grounded in tools from probability, optimization, and operator theory. This survey organizes the mathematical structures that underpin the design and analysis of modern algorithms in RL. We begin from Markov decision processes (MDPs) and the Bellman operators, emphasizing contraction mappings, monotonicity, and fixed-point theory that yield convergence guarantees and rates for value and policy iteration, and temporal-difference schemes. We then develop the optimization perspective: stochastic approximation and martingale methods, convex duality and the role of regularization linking mirror/proximal methods. Function approximation is treated through linear and non-linear settings, covering stabilization, error decomposition, and sample-complexity via concentration inequalities for dependent data and mixing processes. We further cover off-policy evaluation/learning, constrained RL and constrained MDPs (CMDPs). Throughout we unify algorithmic templates under common operator and variational lenses, highlighting both finite-sample bounds and asymptotic results. Our presentation is intended to provide a unified mathematical entry point for researchers in probability, optimization, and statistics interested in reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。