多组动作强化学习中,实时学控并满足弱耦合约束。
Multi-Action Restless Bandits with Weakly Coupled Constraints: Simultaneous Learning and Control
- 设计在线算法,边学边控,无需先验模型
- 证明时间与系统规模双维度收敛,性能偏差指数下降
- 适合资源受限的动态决策场景
我们研究一个包含有限组多动作老虎机过程的系统,每组均为具有有限状态和动作空间的马尔可夫决策过程(MDP),不同动作对应不同的转移矩阵。同一组内各过程共享相同的状态与动作空间,且执行相同动作时具有相同的转移矩阵。所有组间过程在状态和动作变量上受多个弱耦合约束。不同于以往仅关注离线情形的研究,本文考虑在线情形,不假设预先知晓转移矩阵和奖励函数,提出一种能实现同步学习与控制的有效方案。我们证明了相关过程在时间维度和过程数量维度上的收敛性,即时间与规模维度的收敛。此外,还证明了在规模维度上收敛速度呈指数级,导致所提在线算法与离线最优解之间的性能偏差指数衰减。
原文摘要 · Abstract (English)
We study a system with finitely many groups of multi-action bandit processes, each of which is a Markov decision process (MDP) with finite state and action spaces and potentially different transition matrices when taking different actions. The bandit processes of the same group share the same state and action spaces and, given the same action that is taken, the same transition matrix. All the bandit processes across various groups are subject to multiple weakly coupled constraints over their state and action variables. Unlike the past studies that focused on the offline case, we consider the online case without assuming full knowledge of transition matrices and reward functions a priori and propose an effective scheme that enables simultaneous learning and control. We prove the convergence of the relevant processes in both the timeline and the number of the bandit processes, referred to as the convergence in the time and the magnitude dimensions. Moreover, we prove that the relevant processes converge exponentially fast in the magnitude dimension, leading to exponentially diminishing performance deviation between the proposed online algorithms and offline optimality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。