提出可变环境下的强化学习框架,证明算法仍能收敛。
Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis
- 环境在多个MDP间切换,由隐马尔可夫链控制。
- 标准TD学习和Q-learning几乎必然收敛到最优解。
- 适用于通信网络等快速变化系统,适合工业应用。
我们引入切换非平稳马尔可夫决策过程(SNS-MDP)框架,其中环境在有限个MDP间由隐马尔可夫链驱动切换,而智能体仅观测外部状态。我们证明,这种切换的长期效应等价于由隐藏马尔可夫链平稳分布参数化的平稳动态。对于固定策略,我们推导出SNS价值函数的闭式表达,并证明尽管存在持续非平稳性,标准时差(TD)学习仍几乎必然收敛到该值。我们进一步证明策略迭代收敛至等效平均环境的最优策略,并证明表格型Q-learning几乎必然收敛到最优Q函数。该框架在具有马尔可夫信道噪声的无线通信网络上得到验证,展示了其在快速时变系统中决策的有效性。
原文摘要 · Abstract (English)
We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the external state. We show that the long-term effect of this switching is equivalent to stationary dynamics parameterized by the stationary distribution of the hidden Markov chain. For fixed policies, we derive a closed-form expression for the SNS value function and prove that standard temporal-difference (TD) learning converges to it almost surely despite persistent non-stationarity. We further establish that policy iteration converges to the optimal policy of the equivalent averaged environment, and prove that tabular Q-learning converges almost surely to the optimal Q-function. The framework is validated on a wireless communication network with Markovian channel noise, demonstrating its practical efficacy for decision-making in rapidly time-varying systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。