arXiv:2502.03244cs.LG2025-02被引 1
用绝对概率序列分析价值迭代在L2范数下的收敛性
Analysis of Value Iteration Through Absolute Probability Sequences
- 引入绝对概率序列构建新分析框架
- 首次实现价值迭代在L2范数下的收敛性证明
- 为算法性能评估提供新视角,适合强化学习研究者
价值迭代是求解马尔可夫决策过程(MDPs)的常用算法。以往研究多聚焦于其在无穷范数下的收敛性,而本文通过绝对概率序列构建新的分析路径,首次实现了对价值迭代在L2范数下收敛性的严格分析,为理解该算法的行为与性能提供了全新视角。
原文摘要 · Abstract (English)
Value Iteration is a widely used algorithm for solving Markov Decision Processes (MDPs). While previous studies have extensively analyzed its convergence properties, they primarily focus on convergence with respect to the infinity norm. In this work, we use absolute probability sequences to develop a new line of analysis and examine the algorithm's convergence in terms of the $L^2$ norm, offering a new perspective on its behavior and performance.
强化学习价值迭代收敛分析
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。