用最大正线性近似改进离线强化学习的拟合Q迭代算法
Fitted Q-Iteration via Max-Plus-Linear Approximation
- 基于最大正线性近似设计新拟合Q迭代算法
- 算法每轮迭代复杂度与样本数无关,可高效处理大规模数据
- 理论保证收敛性,适合需要稳定训练的离线强化学习场景
本研究将最大正线性近似器应用于折扣马尔可夫决策过程的离线强化学习中。具体而言,我们将其融入拟合Q迭代(FQI)算法,提出具有可证明收敛性的新算法。利用贝尔曼算子与最大正运算的兼容性,我们发现所提算法每轮中的最大正线性回归可简化为简单的最大正矩阵-向量乘法。此外,我们还考虑了该算法的变分实现,使得每轮迭代的复杂度与样本数量无关。
原文摘要 · Abstract (English)
In this study, we consider the application of max-plus-linear approximators for Q-function in offline reinforcement learning of discounted Markov decision processes. In particular, we incorporate these approximators to propose novel fitted Q-iteration (FQI) algorithms with provable convergence. Exploiting the compatibility of the Bellman operator with max-plus operations, we show that the max-plus-linear regression within each iteration of the proposed FQI algorithm reduces to simple max-plus matrix-vector multiplications. We also consider the variational implementation of the proposed algorithm which leads to a per-iteration complexity that is independent of the number of samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。