多智能体协作学习中,通信成本极低仍能实现样本效率线性提升。
Variance-Reduced Q-Learning over Static and Time-Varying Networks

- 每轮迭代中,智能体本地估计贝尔曼算子并通过共识协议交换信息。
- 在静态与动态网络下,算法均实现高概率有限时间收敛,且收敛速度随智能体数线性加快。
- 突破以往通信开销瓶颈,仅需约常数级通信即可达成样本复杂度加速。
我们研究多个智能体共同与同一马尔可夫决策过程(MDP)交互的去中心化强化学习问题。智能体可通过网络交换信息,协同学习最优状态-动作值函数。针对该场景,我们提出一种基于轮次的分布式Q-learning算法VRDQ:在每个轮次中,智能体本地估计贝尔曼最优算子,并使用基于一致性的协议进行信息扩散。对于静态和时变网络,我们建立了VRDQ的高概率有限时间收敛速率,表明协作可带来线性加速。关键在于,我们证明了这种样本复杂度上的加速仅需 ilde{O}(1)级别的通信开销,显著优于已有工作。
原文摘要 · Abstract (English)
We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange information over a network to collectively learn the optimal state-action value function. For this setting, we introduce a novel epoch-based distributed $Q$-learning algorithm called VRDQ, where within each epoch, agents locally estimate the Bellman optimality operator and diffuse information using a consensus-based protocol. For both static and time-varying networks, we establish high-probability finite-time convergence rates for VRDQ that enjoy linear speedups from collaboration. Crucially, we prove that such speedups in sample-complexity require only $\tilde{O}(1)$ communication, substantially improving upon the communication costs in prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。